Google launches AI protein watermarking technology
The AI protein design technology has rapidly gained popularity. Google has introduced a new protein watermarking technique that enables traceable and untraceable origin of AI-synthesized proteins, filling the regulatory gap in the industry.
The advantages and disadvantages of AI protein synthesis coexist, and the industry is facing a traceability dilemma.
With the help of AI tools, researchers can independently design completely new proteins, resulting in many beneficial achievements for society. For instance, specially designed enzymes can rapidly decompose plastic waste, and some proteins can block the harm caused by biological venom. The application prospects are extremely broad. However, this cutting-edge technology is a double-edged sword, and the risk of abuse can not be ignored. The same AI tools can also be used to modify virus proteins and synthesize toxic and harmful proteins, posing significant biological safety risks.
At present, the global biosafety screening system has obvious shortcomings. Conventional detection software can identify the DNA sequences of known viruses and toxins, but it is unable to distinguish the proteins newly designed by AI. These artificially synthesized proteins have no historical data for reference, carry unknown risks, and remain in a regulatory blind spot. Even though the industry discovered this problem nearly a year ago, there has been no feasible solution. The latest research by Google DeepMind has precisely filled this critical gap.
From digital to biological, the SynthID technology achieves cross-border implementation.
This protein traceability solution launched by Google is an upgraded version based on the mature SynthID watermarking technology. It has been newly named SynthIDBio. Previously, this technology was widely used in digital content such as images and texts, featuring seamless embedding, anti-tampering, and traceability. The advantages of traditional digital watermarking are very prominent. The watermark is evenly distributed throughout the content and is invisible to the naked eye. The watermark is not lost and can not be removed or cracked without an exclusive key, ensuring extremely high security.
Transplanting this technology to the protein field is an unprecedented challenge. The pixel base of digital content is huge, providing ample space to hide watermark signals. However, protein structures are extremely simple, consisting of only 20 types of amino acids. Small proteins have only a few hundred amino acids, and the operational space is extremely small. The protein fault tolerance rate is extremely low. How to embed the watermark without damaging the protein function becomes the biggest technical difficulty, and it is also the core reason why the industry has been unable to break through for a long time.
The principle of seamless embedding: Implanting identifiers based on the logic of protein generation.
The Google team, relying on the mainstream AI protein design tool ProteinMPNN, has successfully implemented the technology. This tool is also currently a core and commonly used software in the field of biological research. The conventional ProteinMPNN method for designing proteins involves two steps: first, building a stable protein backbone structure, then matching and adding amino acid side chains at each site, and finally forming a complete functional protein.
The watermark will be scattered and evenly distributed throughout the entire protein sequence, making it impossible to detect with the naked eye or through conventional detection methods. During the traceability process, the dedicated software will combine the key and use data analysis to determine whether the protein is artificially synthesized by AI, accurately completing the traceability process.
Experimental verification: Watermarked proteins fully retain their biological functions.
Many people may question whether fine-tuning of amino acids would damage the original functions of proteins. The Google team conducted multiple sets of control experiments to verify the feasibility of the technology. The experimental results showed that the proteins with watermarks and the ordinary AI-designed proteins performed identically, being able to precisely bind to target molecules, and their core biological functions were not lost at all.
The core application scenario of this technology focuses on the safety screening in the DNA synthesis industry. Currently, there are obvious loopholes in the global DNA order review process. When faced with completely new AI-synthesized unknown proteins, the detection personnel cannot quickly determine the risks and can only blindly investigate, consuming a lot of manpower and resources. SynthIDBio can reconfigure the screening logic. After the detection institution scans, they can quickly determine the compliance of the source and directly release it.
Staff do not need to repeatedly verify the protein samples from trusted institutions. They can focus on investigating unknown proteins without watermarks and unknown sources, significantly improving the efficiency of biological safety screening and precisely avoiding the risk of maliciously synthesizing harmful proteins. It should be clearly stated that this technology is not an absolute security guarantee tool. It can only optimize the screening process and cannot completely eliminate biological risks.