Since the advent of the Unimate in 1954, humanity has had to grapple with anticipating the behavior of a robotic system. The rate at which that behavior changed was significantly slower from the 1960s to today. We had robots that were programmed by hand, with no LLM to aid us, that used careful orchestration of models of the physical world to inform behavior. Because we had well-described models, we knew, more or less, what the robot would do in response to stimulus. We knew when it would be unsafe to be around a robot. We did this by leveraging our fantastic ability to model intelligent systems in the world. Even when we know better, our brains model physical systems as if they are intelligent like us. When Stefanie and I wrote about the physically grounded Turing test, we wrote about this exact phenomenon.
But now robots are being programmed with new behavior, live in front of us. Even on my locally hosted Qwen3.8-27B-fp8 model, I get 100 tok/s of code that can directly command and control a robot, or be reconfigured into a JEV-like configuration to directly command the robot. I can download the latest Physical Intelligence model or leverage Gemini Robotics’ behavior model. All of these different interfaces to command a robot are new entities we as humans need to model. The modeling now runs in both directions. Tavus recently previewed Griffin, a video model that watches your face and voice in real time and decides when to nod, interrupt, or wait. Forty-eight percent of participants in their one-minute test calls believed they were talking to a human. To model these things takes time, through empirical study of the system itself, and building an abstraction of the system we are engaged with. It is no surprise that we have generated new vernacular as a species to describe the “slop” that LLMs produce when we have observed poor textual output or personality in a model. Similarly, we are likely to see behavioral slop as robot models become more capable.
I see a future field we all engage with that focuses on robopsychology. Similar to Susan Calvin, Isaac Asimov’s archetypal robopsychologist, the field would be dedicated to an empirical study of input X resulting in behavior Y. Russ Tedrake in an interview with Brian Heater on the A3 Automated podcast said that we are more like behavioral scientists building things we don’t fully understand and then probing them to figure out what happened. We will build language to describe the environments that robots observe and what the resulting behavior will be. We’ll have language to describe the layout of the arms beyond bimanual and unimanual (might even look more like biblically accurate angels). We’ll have language to describe the types of manipulation interfaces we are using, such as UMI entering common robotics vocabulary since 2024. Due to the pre-paradigmatic condition we are in, we are building up our collective abstractions on how we describe our own world with robots in them. Calling them clankers will be considered trite relative to more nuanced terms we could use instead, such as Synth from Star Trek.
The more we can anticipate the systems we are building, the faster we will be at working on them and communicating between each other on what we are doing and what LLMs will code for us. In order to build better autonomous vehicles, the AV startups explained to drivers who were collecting data how the system would respond to certain stimuli. Because the automotive industry had spent a century on establishing language, the operator and technologist were able to align their thinking to respond to a pedestrian walking by, or obeying traffic signals, or defining a bike lane. This resulted in better data being collected by the drivers themselves. People that call agents rogue, such as what happened with the Hugging Face attack, are in a sense anthropomorphizing the system. I would characterize this form of anthropomorphization as productive because it helps us understand and build vocabulary around anomalous behavior.
I already have new symbols or words I use in my job all the time. When I talk about Data Factory 1 (DF1) at Tutor, the Koala grippers from RAI, UMI, VLAs from just about everyone, these are all words that are making it easier for me to communicate quickly with roboticists and LLMs alike to be more productive at producing better behavior of robots. These are shorthands or tricks that make it easier for all of us to anticipate what is going to happen when I mix these different interfaces together. I for one am excited to become fluent in robonese in the coming decade.




