Google DeepMind announced SynthID Bio on September 30 (US time), a technique that embeds an imperceptible signature into proteins designed by AI. The work was published in Nature, and the code, in vitro data and model weights have been released to the research community. It extends the idea of watermarking, long used for images and text, to biological design.
Signatures in both amino acid sequences and 3D structures
SynthID Bio is a family of methods for adding watermarks that are easy to detect yet preserve function in AI-generated biological sequences and structures. It targets two things.
The first is the amino acid sequence. When a ProteinMPNN-based sequence design model picks amino acids, a slight bias is applied to the probability distribution, leaving a statistically detectable pattern. Detection relies on a "g-value" that measures this statistical deviation: sequences without a watermark average about 0.5, while watermarked sequences score well above it.
The second is 3D structure. DeepMind fine-tuned the diffusion module of AlphaFold 3 so that the signature is embedded in the predicted atomic coordinates themselves. Prediction accuracy is maintained and detection is described as near-perfect. According to DeepMind, the signature can also be verified on the physical protein after synthesis.
Binding performance on par with unwatermarked designs
The evaluation used protein binders, proteins designed to bind a target, against three targets: VEGF-A, the receptor-binding domain of the SARS-CoV-2 spike protein, and PD-L1. Watermarked designs matched unwatermarked ones on hit rate, binding affinity and natural sequence diversity. In early tests on a modified bacteriophage, function was reportedly retained as well.
Aiming to close a biosecurity gap
The backdrop is safety management as AI protein design spreads. DNA synthesis providers screen the sequences they are asked to produce. With a watermark, it becomes easier to tell whether a sequence came from a trusted model or from an unknown source. DeepMind positions the watermark as a way to strengthen that screening.
Many challenges remain
For now, this is a proof of concept. DeepMind calls resistance to deliberate tampering a challenge still to be met. Reports also point to other limits.
- Short proteins may not contain enough amino acids for the signature to be read reliably
- The effect is uncertain for complex proteins such as enzymes, where small sequence changes alter function
- Compatible verification systems and shared standards need to be adopted across the industry
DeepMind therefore suggests pairing the watermark with provenance metadata or central repositories.
What has been released
On GitHub, the main software is released under the Apache License 2.0. The ProteinMPNN model parameters are under the MIT License, and the data files under CC BY 4.0. The alphabet consists of 21 amino acids plus a mask token, with a default n-gram length of 4 and a default context history size of 1,024. Researchers can try it on their own models.
Summary
SynthID Bio brings the idea of watermarking AI-generated content into protein design in the life sciences. Showing that a signature can be left without degrading design performance is a notable step forward. Tamper resistance and support for short sequences remain open problems on the road to practical use. How researchers verify the released results will be the next thing to watch.
References
[1] https://deepmind.google/blog/introducing-synthid-bio/
[2] https://github.com/google-deepmind/synthidbio
