SynthID Bio is a family of methods developed by Google DeepMind for embedding highly detectable yet function-preserving watermarks directly into AI-generated biological sequences and structures.
For more details, please refer to our paper: "Function-preserving watermarking of AI-generated proteins".
SynthID Bio provides watermarking for autoregressive inverse folding models (such as ProteinMPNN) and structure prediction models (such as AlphaFold 3):
- SynthID Bio-structure: Fine-tuned AlphaFold 3 model that embeds imperceptible watermarks into generated 3D biomolecular structures.
- SynthID Bio-sequence: Watermarking protein sequences during autoregressive decoding with ProteinMPNN while preserving function and expressivity.
SynthID Bio-structure is a fine-tuned AlphaFold 3 (AF3) model that embeds an
imperceptible watermark into biomolecular structures during sampling.
For full instructions on downloading model weights and running watermarked structure predictions with AlphaFold 3, please refer to the AlphaFold 3 Repository.
SynthID Bio-sequence introduces watermarked sampling into ProteinMPNN.
Note
Drop-in Integration: This watermarked version of ProteinMPNN is designed to be a drop-in replacement for existing protein design pipelines. With minimal installation overhead, it can be integrated into your current workflows, enabling the generation of watermarked protein sequences by simply enabling the watermarking flags.
We recommend using uv or standard Python virtual environments (venv) to manage dependencies.
Follow these steps to set up the environment from scratch on a Linux system:
-
Create and activate a virtual environment:
uv venv synthid_bio --python 3.9 source synthid_bio/bin/activate(Or using standard
venv:python3.9 -m venv synthid_bio && source synthid_bio/bin/activate) -
Install dependencies:
cd synthidbio_sequence uv pip install -r requirements.txt(Or using
pip:pip install -r requirements.txt) -
Install SynthID Text package:
uv pip install -e third_party/synthid-text
(Or using
pip:pip install -e third_party/synthid-text)
Navigate to the ProteinMPNN directory before running generation scripts:
cd third_party/ProteinMPNNRun Example 1 to generate baseline unwatermarked sequences and pre-parsed inputs:
(cd examples && bash submit_example_1.sh)Outputs are stored in outputs/example_1_outputs/seqs/ and include baseline
g-values (~0.5) for detection:
>T=0.1, sample=1, score=0.8581, global_score=0.8581, seq_recovery=0.4151, mean_g_value=0.5037
SIDEDTQKALDFVKALEEANPELMKKVITPDTEMEVNGKKYKGEEIVEFVKELAAKGVK...
When comparing watermarked to unwatermarked sequences, the watermarking
arguments should remain identical across runs so that comparative g-values can
be evaluated. Without the --watermark flag, watermarks are not embedded into
generated sequences, but detection g-values are still computed so you can
measure the watermark's signal strength against the unwatermarked baseline.
[!TIP] If you have not run
submit_example_1.shabove, you can substitute--jsonl_path outputs/example_1_outputs/parsed_pdbs.jsonlwith the pre-parsed dataset at--jsonl_path test_data/helper_outputs/parsed_pdbs.jsonl.
Use repeating keys for a stronger watermark signal, best suited for low-temperature generation (e.g., 0.1):
python protein_mpnn_run.py \
--jsonl_path outputs/example_1_outputs/parsed_pdbs.jsonl \
--out_folder outputs/example_1_watermark_outputs \
--num_seq_per_target 2 \
--sampling_temp 0.1 \
--seed 37 \
--batch_size 1 \
--watermark \
--ngram_len 4 \
--watermark_keys 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0Output includes high g-values (>>0.5):
>T=0.1, sample=1, score=1.1907, global_score=1.1907, seq_recovery=0.3868, mean_g_value=0.8902
TIDDDTAIALKYVESLELADPALMSKVITPNTKMEYNGREFVGEEIVAYVEEVKKEGVK...
To run an equivalent unwatermarked sequence with distortionary g-values, remove
--watermark while keeping the same watermarking arguments:
python protein_mpnn_run.py \
--jsonl_path outputs/example_1_outputs/parsed_pdbs.jsonl \
--out_folder outputs/example_1_unwatermarked_outputs \
--num_seq_per_target 2 \
--sampling_temp 0.1 \
--seed 37 \
--batch_size 1 \
--ngram_len 4 \
--watermark_keys 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0Use non-repeating keys for minimal impact on sequence quality at medium-to-higher sampling temperatures (e.g., 0.5):
python protein_mpnn_run.py \
--jsonl_path outputs/example_1_outputs/parsed_pdbs.jsonl \
--out_folder outputs/example_3_watermark_outputs \
--num_seq_per_target 2 \
--sampling_temp 0.5 \
--seed 37 \
--batch_size 1 \
--watermark \
--ngram_len 4 \
--watermark_keys 0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24Output with moderate g-values:
>T=0.5, sample=1, score=1.2134, global_score=1.2134, seq_recovery=0.4150, mean_g_value=0.5892
SVDPDTKRAHDFVDALELADPALMREVITPDTQMEYNGQRFVGEEIVEFVKQIAAEGR...
Similarly, to run an equivalent unwatermarked sequence with non-distortionary
g-values, remove --watermark while keeping the same watermarking arguments:
python protein_mpnn_run.py \
--jsonl_path outputs/example_1_outputs/parsed_pdbs.jsonl \
--out_folder outputs/example_3_unwatermarked_outputs \
--num_seq_per_target 2 \
--sampling_temp 0.5 \
--seed 37 \
--batch_size 1 \
--ngram_len 4 \
--watermark_keys 0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24For existing FASTA files, use the standalone compute_g_values.py script to
compute g-values without regenerating sequences:
python compute_g_values.py \
--input_fasta test_data/test_sequences.fa \
--output_fasta outputs/sequences_with_gvalues.fa \
--ngram_len 4 \
--watermark_keys 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0Input FASTA format:
>test_seq_1
KMYEYKKIGDEYVVAIYNESEMMTALTTFCKDKNIKSGTITGIGQIKEITLKYFNPETKE...
>test_seq_2
KLYEYKKIGDEYVVNIYDNTEIVKSILEFCEEKNILSGTIQGIGQIKEIELQFFDPETKE...
Output FASTA format:
>test_seq_1, mean_g_value=0.7286
KMYEYKKIGDEYVVAIYNESEMMTALTTFCKDKNIKSGTITGIGQIKEITLKYFNPETKE...
>test_seq_2, mean_g_value=0.7714
KLYEYKKIGDEYVVNIYDNTEIVKSILEFCEEKNILSGTIQGIGQIKEIELQFFDPETKE...
Important: Use the same watermark parameters (e.g., --ngram_len,
--watermark_keys) as those used during sequence generation for detection.
| Parameter | Description | Default | Notes |
|---|---|---|---|
--watermark |
Enable watermarking | False | Must be set to embed watermarks |
--ngram_len |
N-gram length | 4 | Context window conditioning watermark (e.g., 2, 4) |
--watermark_keys |
Comma-separated keys | 0,1,2,...,24 | Sequence of integer keys, one per depth/tournament layer |
--context_history_size |
Context size | 1024 | SynthID internal buffer size to track seen contexts |
--watermark_temperature |
Watermark temp | 1.0 | Temperature scaling factor for watermark distortion strength |
--watermark_top_k |
Top-k value | 21 | Set to vocab size (21 amino acids + mask token) |
--skip_first_ngram_calls |
Skip first tokens | True | Disables watermarking for first (ngram_len - 1) tokens |
SynthID Bio-sequence is designed to be model-agnostic and can be integrated into other autoregressive protein sequence decoders with minimal changes.
To apply watermarking to a custom protein decoder, follow these steps:
-
Initialize the Processor: Create an instance of
SynthIDLogitsProcessorfromsynthid_text.logits_processing.- For protein sequences, set
apply_top_k=Falseto maintain the full amino acid alphabet, and settop_kto the vocabulary size (typically 21 for standard amino acids + mask token). - Ensure the
devicematches the model's device.
- For protein sequences, set
-
Enforce Sequential Decoding: SynthID Bio-sequence requires sequential (left-to-right) decoding because the watermark at position t depends on the history of generated residues from 0 to t - 1. If your model uses random decoding order (like vanilla ProteinMPNN), you must force it to decode sequentially when watermarking is enabled.
-
Intercept logits in the sampling loop:
- At each step t, before sampling, construct the context of already-sampled tokens (IDs) up to t - 1.
- Call
watermarked_callon the processor with the context and the raw logits. - Use the returned watermarked scores (which are log probabilities) for sampling.
Here is a conceptual example of integrating SynthID Bio-sequence into a generic autoregressive decoding loop:
from synthid_text.logits_processing import SynthIDLogitsProcessor
import torch
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
# 1. Initialize the processor
watermark_processor = SynthIDLogitsProcessor(
ngram_len=4,
# Watermark keys (one per depth layer)
keys=[
0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18,
19, 20, 21, 22, 23, 24,
],
context_history_size=1024,
temperature=1.0,
top_k=21, # Vocabulary size (e.g., 20 amino acids + mask token)
device=device,
apply_top_k=False, # Keep full alphabet for protein design
)
# 2. Sequential decoding loop
# generated_tokens: List of already sampled token IDs (shape: [batch_size, current_len])
generated_tokens = []
for t in range(sequence_length):
# Get raw logits from your model
logits = model.get_logits(..., generated_tokens) # [batch_size, vocab_size]
# Convert context to int64 tensor
input_ids = torch.tensor(
generated_tokens, dtype=torch.long, device=device
).unsqueeze(0) # [1, t]
# Get watermarked scores (log probabilities) and top_k indices
watermarked_scores, top_k_indices, _ = watermark_processor.watermarked_call(
input_ids, logits
)
# Convert log probs to probabilities
probs = torch.softmax(watermarked_scores, dim=-1)
# Sample next token
sampled_idx = torch.multinomial(probs, 1)
next_token = torch.gather(top_k_indices, 1, sampled_idx)
generated_tokens.append(next_token.item())This repository includes a test suite to verify the installation and ensure watermarking works correctly.
test_example_1.py: Verifies deterministic non-watermarked and watermarked generation (validating sequences, scores, recovery rates, and exact g-values).test_example_*_watermark.py: Tests end-to-end watermarked generation across examples 2–8.test_g_values.py: Tests g-value detection accuracy for both watermarked and non-watermarked sequences.test_compute_g_values.py: Tests the standalone FASTA g-value calculator.
To run the tests, navigate to the synthidbio_sequence/third_party/ProteinMPNN directory and use pytest:
cd synthidbio_sequence/third_party/ProteinMPNN
# Run core tests
pytest test_example_1.py test_g_values.py -v
# Run the standalone calculator test
pytest test_compute_g_values.py -v
# Run all tests
pytest -v==================== test session starts ====================
test_example_1.py::test_example_1_non_watermarked PASSED [ 25%]
test_example_1.py::test_example_1_watermarked PASSED [ 50%]
test_g_values.py::test_g_values_non_watermarked PASSED [ 75%]
test_g_values.py::test_g_values_watermarked PASSED [100%]
==================== 4 passed in 45.23s =====================
-
ProteinMPNN (
synthidbio_sequence/third_party/ProteinMPNN/protein_mpnn_run.py,synthidbio_sequence/third_party/ProteinMPNN/protein_mpnn_utils.py)- Protein sequence generation model
- Modified to support watermarking during decoding
- Forces left-to-right decoding when watermarking enabled
-
SynthID Text (
synthidbio_sequence/third_party/synthid-text/)- Watermarking logic processor
- G-value computation for detection
- Tournament-based token selection
-
Detector (integrated in
protein_mpnn_run.pyandcompute_g_values.py)- Always-on g-value computation
- Uses same SynthID processor for detection
- Independent of watermarking mode
-
ProteinMPNN (
synthidbio_sequence/third_party/ProteinMPNN/): MIT License- Copyright (c) 2022 Justas Dauparas
- Original publication: Dauparas et al., "Robust deep learning-based protein sequence design using ProteinMPNN" (2022)
-
SynthID-text (
synthidbio_sequence/third_party/synthid-text/): Apache License 2.0- Copyright Google DeepMind
- Watermarking technology for text and sequences
To integrate SynthID watermarking and ensure determinism, the following files in
the synthidbio_sequence/third_party/ProteinMPNN directory were modified,
added, or removed:
synthidbio_sequence/third_party/ProteinMPNN/protein_mpnn_run.py:- Added command-line arguments to configure SynthID watermarking (e.g.,
--watermark,--ngram_len,--watermark_keys,--num_leaves). - Imports
SynthIDLogitsProcessorandmean_scorefromsynthid_text. - Initializes the
SynthIDLogitsProcessorand passes it to the model'ssamplefunction. - Enforces
batch_size=1when watermarking is active. - Computes g-values for all generated sequences and appends the
mean_g_valueto the output FASTA headers.
- Added command-line arguments to configure SynthID watermarking (e.g.,
synthidbio_sequence/third_party/ProteinMPNN/protein_mpnn_utils.py:- Modified the
samplemethod ofProteinMPNNclass to acceptwatermark_processor. - Forces left-to-right sequential decoding when
watermark_processoris present, overriding the default random decoding order. - Enforces
batch_size=1check insidesamplewhenwatermark_processoris present. - Intercepts logits at each step of the decoding loop and applies
watermark_processor.watermarked_callto compute watermarked scores before sampling. - Modified
tied_samplesignature to acceptwatermark_processorfor compatibility.
- Modified the
synthidbio_sequence/third_party/ProteinMPNN/helper_scripts/parse_multiple_chains.py:- Sorted the output of
glob.globto ensure deterministic parsing order of PDB files for testing.
- Sorted the output of
synthidbio_sequence/third_party/ProteinMPNN/compute_g_values.py:- Standalone script to compute g-values for existing sequences in FASTA
format, using
SynthIDLogitsProcessorfor detection.
- Standalone script to compute g-values for existing sequences in FASTA
format, using
synthidbio_sequence/third_party/ProteinMPNN/test_*.py:- Pytest suite (10 files) for validating various ProteinMPNN examples with and without watermarking, ensuring exact score and g-value matches.
synthidbio_sequence/third_party/ProteinMPNN/test_data/:- Directory containing golden outputs and helper inputs used by the test suite to ensure determinism.
synthidbio_sequence/third_party/ProteinMPNN/examples/*_watermark.sh:- Example bash scripts (for examples 2, 3, 4, 5, 6, 8) demonstrating how to run ProteinMPNN with watermarking parameters.
synthidbio_sequence/third_party/ProteinMPNN/training/:- Removed training code, data, and local training weights, as they are not required for inference or validation.
Derived g-values and in vitro validation data supporting the publication are
available in the synthidbio_sequence/data/ directory. This includes binding
data CSVs and kinetics plots for each target:
synthidbio_sequence/data/all_g_values.csv- G-values for all sequencessynthidbio_sequence/data/SC2RBD/- Binding data (sc2rbd_data.csv) and kinetics plots (sc2rbd_kinetics/) for SARS-CoV-2 RBDsynthidbio_sequence/data/VEGF-A/- Binding data (vegfa_data.csv) and kinetics plots (vegfa_kinetics/) for VEGF-Asynthidbio_sequence/data/PD_L1/- Binding data (pdl1_data.csv) and kinetics plots (pdl1_kinetics/) for PD-L1
If you use SynthID Bio in your research, please cite:
@article{synthidbio2026,
title={Function-preserving watermarking of AI-generated proteins},
author={David Stutz and Alexander I. Cowen-Rivers and Guillermo Ortiz-Jimenez and Jeremy Ratcliff and Vinicius Zambaldi and Lindsay Willmore and Josh Abramson and Harshnira Patani and Christina Kouridi and Florian Stimberg and Mel Vecerik and Alex Chu and Sukhdeep Singh and Sumanth Dathathri and Eliseo Papa and Valentin De Bortoli and Arnaud Doucet and Demis Hassabis and Jue Wang and Sven Gowal and Pushmeet Kohli},
journal={Nature},
volume={?},
number={?},
pages={?},
year={2026},
publisher={Nature Publishing Group UK London}
}For questions surrounding SynthID Bio, please contact synthidbio@google.com.
Copyright 2026 Google LLC
All software is licensed under the Apache License, Version 2.0 (Apache 2.0); you may not use this file except in compliance with the Apache 2.0 license. You may obtain a copy of the Apache 2.0 license at: https://www.apache.org/licenses/LICENSE-2.0
The ProteinMPNN model parameters were originally made available at https://github.com/dauparas/ProteinMPNN#proteinmpnn under the MIT License. For ease of access, the ProteinMPNN model parameters are also available in this repository under the same license.
The AlphaFold 3 model parameters are made available under the AlphaFold 3 Model Parameters Terms of Use (the "Terms"); you may not use these except in compliance with the Terms. You may obtain a copy of the Terms at https://github.com/google-deepmind/alphafold3/blob/main/WEIGHTS_TERMS_OF_USE.md and in the file WEIGHTS_TERM_USE.
Data in the synthidbio_squence/data directory is licensed under the Creative Commons Attribution 4.0 International License (CC-BY). You may obtain a copy of the CC-BY license at: https://creativecommons.org/licenses/by/4.0/legalcode
Unless required by applicable law or agreed to in writing, all software and materials distributed here under the Apache 2.0, the Terms or CC-BY licenses are distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the licenses for the specific language governing permissions and limitations under those licenses.
You are solely responsible for determining the appropriateness of using the software, model parameters, materials or using or distributing outputs, and assume any and all risks associated with such use or distribution and your exercise of rights and obligations under these Terms. You and anyone you share output with are solely responsible for these and their subsequent uses.
Output are predictions with varying levels of confidence and should be interpreted carefully. Use discretion before relying on, publishing, downloading or otherwise using AlphaFold 3.
All software, model parameters, materials and any outputs you create are for theoretical modeling only. They are not intended, validated, or approved for clinical use. You should not use the software, model parameters, materials or outputs for clinical purposes or rely on them for medical or other professional advice. Any content regarding those topics is provided for informational purposes only and is not a substitute for advice from a qualified professional. This is not an official Google product.