This week, my X feed exploded with news that Anthropic had released hardware standards that allow AI agents to operate lab equipment directly. Their new “Model Hardware Standard” or MHS, built in collaboration with the Howard Hughes Medical Institute, allows an LLM to operate microscopes, liquid handlers, robotic arms and dozens of other instruments that scientists use to conduct experiments and to build the future.
Until now, there has been no standard way to connect an LLM to a physical instrument. Every time you want to connect two pieces of equipment together, you need to create a bespoke integration that requires some grad student to write custom software. And grad students are busy. So when a lab wants to increase throughput, the incentive structure is usually to just hire more grad students or pressure encourage everyone in the lab to work longer hours. And scientists already work much longer hours than other professions. The standard informal working schedule at many elite labs across the US and China is something close to 996, often for mediocre salaries in the range of $35-70k.
Of course, most people know that it’s silly for humanity to send our best and brightest to spend the better half of a decade getting PhDs in advanced bioengineering just to spend most of their time as manual labour meat sacks moving microliters of liquids between plastic tubes by hand.
The bottleneck in biology has always been the ability to perform rapid iteration experiments to collect real world data. Intellect has been abundant in science since the 1950s, and with the advent of super-intelligent models, it is already becoming commoditized. These days, what distinguishes Earth’s most impactful and highest throughput labs is largely prestige-driven meat sack talent acquisition, work ethic, pipetting technique and ability to fund expensive experiments.
This week’s episode with Lucas Mair is all about what the world will look like when we realise that robots exist and start building accessible, intuitive software to allow them to perform closed-loop laboratory experiments.
Lucas is a very German MD-PhD who took an unusual route into biology. Computer science first, then medicine, then neural engineering at TU Munich, and before any of that, building autonomous drone swarms for search and rescue. He thinks like a control theorist, and when a control theorist looks at how a modern biology lab actually runs, he doesn't see cutting-edge science. He sees a factory floor from 1975.
Lucas will tell you that most labs already have an Opentrons or something similar gathering dust in the corner. My own lab had one that I have never seen anyone use.
Grants pay for reagents and salaries, not for engineers to automate your protocols. Grad students are cheaper than robots, and publishing pressure means the three months you'd spend automating a workflow is three months you're not putting out papers, which is the equivalent of suicide by irrelevancy for many high-profile labs. Thus, the rational move, for every individual scientist and every individual lab, is to just keep pipetting.
And so that’s what everyone does.
Which is insane, because the robots already exist. We've had them for years. What we've been missing is the thing that turns "here's what I want to test" into machines actually doing it.
The Anthropic news is a bit of a preview into what that might look like. But the reality of the current tech is a mixed picture.
For example, scientists at Carnegie Mellon were able to use an AI agent with MHS to orchestrate a liquid handler, plate reader, robotic arm and monitoring cameras to get serial-dilution dose-response experiments done 3x faster.
Meanwhile at Genentech, humans had to intervene when Claude used MHS to coordinate a BCA protein assay because the model mistook errors caused by bubbles for a software failure, and then responded in a way that produced yet more bubbles.
Lucas Mair looks at biology as a control and engineering problem. You have a complex biological system, and you want to move it from State A to State B (e.g. diseased to healthy). But the problem is there is an unimaginably large search space of possibilities. For example, there are 2020 or 104,857,600,000,000,000,000,000,000 ways of making a 20 amino acid protein sequence.
Most of these sequences are useless but a few of them might help make humans immortal.
One more thing to understand before you listen to the full conversation…
If you’ve spent much time at all on the tech bro side of X, you’ll know about the scaling laws of large language models. Roughly speaking, intelligence scales linearly with the log of compute (or data). Thus to improve your model, you need more data.
But most of the recent improvement in LLM performance across disciplines hasn’t been in training larger language models. Rather, the most noticeable advances have been from “reinforcement learning from verifiable rewards” or RLVR. Basically a form of reinforcement learning where the reward you give the model can be checked objectively rather than judged subjectively by humans. After all, if you want to build super-intelligent systems, then at some point, even the smartest humans are going to be useless in teaching them anything. RLVR is really easy to do in fields like math and coding because the correct answer is already known and you can check if the program passed the tests.
But how do you do this in biology? You can’t just read all the papers because most biological data isn’t published and the results of many peer-reviewed papers can’t be reproduced by other labs (Cc: the reproducibility crisis in science).
You don’t need more text tokens. You need more ground truth data (i.e. the actual underlying molecular biology) aimed at the model’s blind spots. The jagged frontier of AI progress is more like a recursive Sierpiński triangle. Each field, subfield and sub-sub-field of human endeavor has their own jagged frontiers that would benefit from RLVR.
If you can generate data aimed at places where the model is uncertain or wrong, using real experiments, the efficiency gain can be enormous. Now imagine running that loop in biology, on its own. The model would propose the hypothesis, plan the experiment, dispatch agents to run the lab equipment and collect real measurements. The data then would get processed to generate the next hypothesis. This is a system that learns by doing, using verifiable, falsifiable predictions.
It’s the same technique that helped AI vibe coding go from being an unusable joke to the default way that 90% of today’s code is written, including at frontier labs. Once you collect enough data on the model’s biological blind spots, you cross some invisible threshold of reliability that makes your AI model generally useful in autonomously solving biological engineering problems, inventing drugs, running experiments, and, eventually, agentically building biotech companies.
Closed-loop automated labs could put real experimental biology in the hands of far more people than the credentialed few who have obtained their pipette licenses.
One day, you might be able to prompt a model to fix your body, cure your cold or build your dog a cancer vaccine.
Or maybe you already can, but just haven’t tried.
Watch on YouTube. Listen on Apple Podcasts or Spotify.
GUEST INFORMATION:
CONNECT WITH US:










