We are building the instrument that records the whole sensory field of
physical work, and the first corpus that could give machines priors about the world.
Move through it
The premise
A person carries a lifetime of physical priors. A machine carries
pixels and joint angles, and nothing else.
Every embodied model on earth is trained on a thin visual slice of the world. Not because
vision is enough, but because vision is what anyone thought to record. Heat, slip,
resistance, chemistry, the sound a material makes as it changes state: none of it exists
in any dataset, anywhere.
This is the first time anyone has tried to record all of it at once, from
a person actually doing the work, in one coherent stream. Not a lab bench with instruments
pointed at an object. The complete sensory record of skilled physical action.
Where it goes
Anywhere skilled hands work.
The instrument is not built for one field. Wherever a person is doing something a machine
cannot yet do, the same senses are carrying the same information, and none of it is being
recorded. These are the places where the missing channels are not a nicety but the
entire basis of the judgement.
Senses that carry the judgement hereNot decisive in this sectorSeven marks per card, in channel order
What we need
People who want the answer.
Places to record, teams who would train on channels that have never existed, and anyone
working on sensing nobody has yet put on a working person.
A cook does not solve the kitchen. They recognise it.
A person walking into an unfamiliar kitchen already knows how things fall, how heat
travels through a pan handle, how much a surface gives before it slips, what a material
sounds like when it is about to break. One look is enough because
almost all of the work was done before they arrived.
That accumulated stock of expectation is what makes a human competent somewhere new. It
is called a prior, and every robot in the world is missing it.
Where a machine would get them
Only from data. And the data is wrong in three specific ways.
01
Wrong channels
Physical priors are not built by looking. They are built by contact: by burning
yourself, by feeling something slip, by smelling a change before you could see it.
Those are exactly the channels no dataset records. You cannot learn a prior about heat
from a corpus that has never measured temperature.
02
Wrong axis of scale
Robot data has been scaled the wrong way, and this is now measured rather than
argued. Generalisation scales as a power law in the number of environments and
objects, not in the number of demonstrations. The field has been collecting
more repetitions in the same few rooms.
03
Wrong source
Teleoperation data records what a robot's actuators did. It does not record what the
person controlling it perceived. The expertise, and therefore the prior, lives in the
human sensory stream, and that stream is thrown away at the moment of capture.
The field's three answers
Each one is real. None of them closes the gap.
These are the serious positions, stated at their strongest, with the numbers that
actually decide them.
World models
Learn to imagine the world, then plan inside it.
The strongest evidence for this position is also the strongest evidence against
relying on it. The flagship system is the only one in the field that genuinely
demonstrates the target property: dropped into a lab it has never seen, with no data
from that environment and no task-specific training, it works.
The catch16 seconds per action. A shipping humanoid's reflex tier
runs at 1,000 per second. That is four orders of magnitude too slow to be the control
loop, however good the representation is.
Scale
Stop engineering. Add compute and data.
The majority position, and it has seventy years of history behind it. General methods
that leverage computation win in the long run, and hand-designed structure gets
overtaken. This has happened repeatedly and it will happen again.
The catchRobotics has neither the data of vision nor the clean
symmetry of chemistry. The largest robot pretraining corpora are around 10,000 hours,
against more than a million hours of video for comparable vision systems.
Structural priors
Build the physics into the architecture.
Encode the symmetries of the physical world directly, so the model does not have to
discover them. In the low-data regime this wins decisively and the effect is
quantified, not hand-waved.
The catch200 demonstrations beat 1,000 without it, and success rose
21.9% at 100 demonstrations. But baked-in structure gets overtaken once the data
arrives, and whole subfields have been caught out this way.
Where we are unique
The field is arguing about how to give machines priors. We think the priors
were never recorded.
Every serious position in that debate is an argument about architecture: what shape the
model should be, how much structure to build in, how much to let scale discover. All
three take the input as given. But a human's physical priors are not architectural. They
are experiential. They were assembled from a lifetime of contact, heat and chemistry, and
if none of that was ever measured, no architecture recovers it and no amount of scale
recovers it either. Scale multiplies whatever you sensed. It cannot multiply what you
never sensed at all.
This is why we are not a model company arguing about architecture, and not a data company
collecting more of the same. We are changing what gets measured.
Everyone gathering robot data records actuators and cameras. Everyone gathering human
data records video and hand position. Nobody records what the person's body was actually
sensing while they worked. Slip, heat and chemistry sit at zero across every public
dataset in every language.
And the honest scope of the claim: this is an argument about capital efficiency,
not about the eventual ceiling. We are competing in the regime of thousands to
hundreds of thousands of trajectories, where the evidence for richer input is unambiguous
and measured. If someone eventually reaches the asymptote with vision alone and enough
compute, they may well get there. We think the cheaper road runs through sensing the
world properly first, and we have designed the experiment that decides it.
Next
What the instrument actually records.
Seven channels, one moment, and the reason making them describe the same instant is the
hard part.
The instrument is worn by a person doing real work. It records what they see, hear, touch,
feel as heat and detect as chemistry, at the moment they are doing it. Select a channel.
The hard part
These senses do not run at the same speed.
A microphone resolves microseconds. A chemical sensor takes seconds. Between them sit
contact, motion, slip and heat, spread across roughly seven orders of
magnitude. Recording them is not the difficulty. Making them describe the same
instant, provably, is.
Every existing multi-sensor system resolves this the same way, and that way destroys the
fast channels. Ours does not. How is not described on this page.
01
The instrument
Wearable capture built for this, because nothing you can buy records these channels
together on a person who is working rather than on a bench with instruments aimed at
an object.
02
The corpus
Real work in real places, not staged demonstrations. Deliberately small and dense. The
aim is not the most hours, it is the first hours that contain what nobody has recorded.
03
The models
Trained on channels no embodied model has ever seen, and judged on one question: does
it hold up somewhere it has never been. Deployed robots will not carry every sense the
instrument does, and the models are designed for that from the start.
Where this stands
What is established, and what is not.
Instrument companies are judged on candour about their own error bars.
Established by others
Extra senses raise performance
Adding non-visual channels within touch improved policy success by 63%.
Sparsh-X · arXiv:2506.14754
Touch and hearing help together
Fusing vision, touch and audio improved success by 20%.
FuSe · arXiv:2501.04693
Training senses need not survive to deployment
Policies trained with richer sensing than the robot will ever carry still deploy on cameras alone.
Scaffolder · arXiv:2405.14853
Open, and ours to answer
Whether it helps where it counts
Published gains are averages. Our question is whether these senses help
disproportionately more somewhere unfamiliar. Nobody has measured that. We are.
Whether chemistry carries usable signal
No dataset pairs it with physical work, so it cannot be settled with existing data.
Our strongest claim and our thinnest evidence, simultaneously.
What it costs per usable hour
Unknown until the first sessions run. We will publish it.
About
Aisthe
From koinē aísthēsis, Aristotle's
term for the faculty that binds the separate senses into a single perception. It is the
oldest name for the thing we are building, and it is still the correct one.
Machines are being asked to act in the physical world with a fraction of the sensory
access a person has. We think that is the constraint, not model scale, and that nobody
has tested it because the recording has never existed. So we are making the recording.
Position
We would rather spend a small amount finding out we are wrong than
a large amount finding out slowly.
The central question has a cheap answer. Before building anything, we are running it
against public data, with the falsification criterion written down in advance. If the
effect is not there, we will say so publicly and the instrument does not get built.
That discipline is the company. Everything else follows from it.
People
Who is behind it.
Mahmoud Omar
Chief Executive Officer
Origin of the sensory-priors thesis behind Aisthe. Owns the research programme, the
falsification discipline that governs it, and the case for why this is worth building
at all.
Noah Zerkin
Chief Technology Officer
Owns the instrument. The wearable capture hardware, the sensing stack that has to
survive contact with real working environments, and the timing architecture that makes
a multi-sensory recording mean one moment rather than seven.
Rulin Zhao
Product
Owns what the corpus becomes. Which sectors are recorded first, what a customer
actually receives, and how the data turns into something a robotics team can build
on rather than a pile of files.
Get in touch
If this is the problem you have been waiting for, say so.
Places to record, people to build with, and investors who fund questions rather than
certainties. A technical briefing is available under NDA.