On August 27, 2026, Anthropic opened a research preview of the Model Hardware Standard (MHS), a shared specification that lets AI agents operate physical equipment such as microscopes and robotic arms. Integration work that used to take weeks or months now takes hours or minutes, and in laser tuning on a quantum computer the success rate jumped from 58 percent to 99.3 percent.

Every Instrument Speaks Its Own Dialect

Lab and factory instruments each come with their own control conventions. Detectors run on MATLAB, cameras on Python, electrophysiology rigs on C#. The languages do not match, and the devices never talk to each other directly. Every new connection means a specialist writing another bespoke bridge, and that is where the weeks and months go.

MHS standardizes that bridge. The design is almost disappointingly plain: driver commands are reduced to simple read and write primitives such as "get temperature" and "set temperature," and devices announce themselves on the network in a standard format. On top of that, users can attach natural-language tags describing things code cannot convey, such as the weight of a robotic arm, which matters directly for safe operation. Those tags can be written by hand or by talking to an agent that interviews you about the setup.

From the tags, MHS generates a reference file listing what the device can measure, what can be adjusted, and what safety limits will be enforced. Control comes through three paths, MCP, a command line interface, and code files, which together allow orchestration across multiple devices from a single line of code. The standard is model-agnostic and works in principle with any device that exposes a programmable interface.

The idea did not originate inside Anthropic alone. Arco Bast, a postdoctoral researcher at HHMI Janelia Research Campus, had already put the state of his custom microscope rig into a single shared-memory dictionary. Anthropic wired AI models into that interface, and his rig became the first to run on MHS.

From 150 Seconds at 58 Percent to 6 Seconds at 96 Percent Overnight

The most striking numbers come from QuEra Computing, which builds neutral-atom quantum computers. Its titanium-sapphire lasers must hold frequency to roughly one part in a trillion, and when they drift, a human takes 5 to 10 minutes to recover them. An automation script already existed, but the version four engineers spent months building succeeded 58 percent of the time and took about 150 seconds per attempt. Because it ran as a linear sequence, a change midway, such as someone opening the lab door and shifting temperature and pressure, sent it back to the start.

Anthropic dropped in a loop of four Claude instances with distinct roles: one proposes improvements, one writes them into the script, one runs them against the real laser and logs every step, and one reads the log to decide the next change. Left unattended, the loop ran several hundred iterations overnight. By morning it was down to about 6 seconds at 96 percent. A subsequent blind test succeeded on 695 of 700 attempts, or 99.3 percent, taking 10 to 14 seconds on the hardest disturbances and around a second on simple ones.

What made the difference was rewriting the procedure as a decision tree. If the frequency has barely moved, most adjustments are pointless, so the agent touches one or two things and stops. A human runs through everything in order just to confirm. The end product is a deterministic script that runs in production without any AI agent involved.

PID tuning of the servo loop produced a similar result. Where an expert's tune left 15.7 millivolts of residual error, Claude reached 1.55 millivolts after 363 experiments across 16 hours of unattended operation. A phase noise analyzer comparison showed the manual tune leaving vastly more noise at a resonance near 220 kilohertz. Over 19 hours of continuous operation, Claude's settings never lost lock.

Integration Moved from Weeks to Hours

At the Baker and Pinglay labs at the University of Washington, connecting six instruments through MHS took under a week including the time spent writing drivers. Wiring alone previously ran from months to years, with costs to match. The lab now watches qPCR amplification curves in real time, asks the researcher for permission to stop before the curve plateaus, and moves the reaction to a 4 degree hold on command. An open-source robotic arm built on LeRobot picks up a plate about 10 seconds after receiving a completion signal from the liquid handler, and across repeated trials the two never collided.

Carnegie Mellon University's case is more extreme. Its setup spans three computers, one running legacy Windows scripting and another with no programmatic interface at all, only a GUI. MHS unified them into a single manifest. Writing drivers from scratch and building the orchestration layer that lets a Claude Opus 4.8 agent run the whole protocol autonomously took about eight hours, against the several weeks a vendor-built setup normally requires. The process itself also ran roughly three times faster.

At Genentech, a liquid handler, a robotic arm, and a plate reader worked in concert, with the agent finding separate optimal aspiration speeds for water and for a viscous solution. The limits showed up here too. When bubbles formed, Claude retried in the same well with different parameters and produced more bubbles. It is a candid record of a system that does not actually understand the physics.

Stopping When Something Looks Risky

Safety rests on MHS enforcing device-level limits. On the Janelia microscopes, there is no longer a worry that an agent will push the power too high, bleach the fluorescent molecules, and damage the sample. At Carnegie Mellon, researchers deliberately created six hazardous conditions, a missing plate, a rotated plate, a busy reader, a disconnected camera, an unreachable device, and an active emergency stop. The system blocked all six before any device moved.

At QuEra the caution ran the other way. Claude halted for human confirmation whenever it sensed any risk, and on at least one occasion an experiment sat idle overnight waiting for approval. The team's verdict was that an overly cautious agent is preferable to one that is not cautious enough.

The list of backers is growing. Amazon Web Services supports MHS through Strands Robots, Tecan is adding support for Fluent, QIAGEN is experimenting on QIAsymphony Connect, and MBF Bioscience is building a driver for ScanImage. Universal Robots, Doosan Robotics, Automata, and Danaher are also on the list. Hugging Face is adding support to LeRobot, and Raspberry Pi is moving toward integration across several products after a successful camera driver test.

Summary

MHS strips instrument control down to read and write primitives and hands agents the safety limits along with the commands. Integration falls to hours or minutes, laser recovery reaches 99.3 percent, and PID tuning beat an expert. At the same time it does not yet reach devices without a programmable interface, and failures rooted in a lack of physical understanding are on record. The standard is currently a research preview accessed through a waitlist, and Anthropic has said it plans to open source it once the specification is further along.