Official ICML 2026 Implementation
- Pretrained checkpoints: HG38 and Rice models are available on Hugging Face.
- VisualDNA pipeline: installation and data-rendering instructions are public.
- Inference APIs: visual features, Decoder features, prompt-conditioned generation, and multi-page inputs are supported.
- π’ News
- π§ͺ 1. Summary
- π€ 2. Pretrained Models and Usage
- π Repository Structure
- βοΈ 3. Environment
- 𧬠4. Install VisualDNA
- ποΈ 5. Data Preparation
- π 6. Pre-training
- π 7. Notes for Release Users
- β 8. Testing and CI
- π 9. License
- [2026/08/27] π€ Released the pretrained OpticalDNA-HG38-2048 and OpticalDNA-Rice-2048 checkpoints on Hugging Face.
- [2026/05/09] Repository initialized!
- [2026/04/30] π Paper accepted by ICML 2026.
OpticalDNA reformulates genomic sequence modeling as a document-understanding problem. DNA sequences are rendered into structured visual pages, encoded by a vision-language backbone, and trained with prompt-conditioned genomic objectives including reading, grounding, ROI transcription, masked completion, subsequence localization, and chromosome-level recognition.
- DNA as visual documents. Long genomic sequences are converted into multi-page DNA documents with pixel-level coordinate annotations.
- Prompt-conditioned genomic pretraining. Six OCR-style tasks cover recognition, grounding, retrieval, and completion.
- Efficient long-context representation. Visual tokens provide compact representations for downstream long-range genomic prediction.
- Lightweight downstream adaptation. The pretrained visual encoder can be reused with linear or shallow MLP heads.
π Citation
# ICML @inproceedings{xiang2026rethinking, title = {Rethinking Genomic Modeling Through Optical Character Recognition}, author = {Xiang, Hongxin and Ma, Pengsen and Cao, Yunkang and Yu, Di and Chen, Haowen and Yang, Xinyu and Zeng, Xiangxiang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, year = {2026}, url = {https://openreview.net/forum?id=nggzekChuU} } # arXiv @article{xiang2026rethinking_arxiv, title = {Rethinking Genomic Modeling Through Optical Character Recognition}, author = {Xiang, Hongxin and Ma, Pengsen and Cao, Yunkang and Yu, Di and Chen, Haowen and Yang, Xinyu and Zeng, Xiangxiang}, journal = {arXiv preprint arXiv:2602.02014}, year = {2026}, url = {https://arxiv.org/abs/2602.02014} }
Most users do not need to pre-train OpticalDNA from scratch. The released checkpoints can be downloaded automatically from Hugging Face and used directly for genomic feature extraction, prompt-conditioned Decoder representations, and OCR-style genomic inference.
| Model | Pretraining corpus | Checkpoint | Feature width | Hugging Face |
|---|---|---|---|---|
| OpticalDNA-HG38-2048 | Human reference genome (HG38) | step 190,000 | 1,280 | π€ hxxiang/opticaldna-hg38-2048 |
| OpticalDNA-Rice-2048 | Rice NIP-T2T (w2048, o1920) |
step 150,000 | 1,280 | π€ hxxiang/opticaldna-rice-2048 |
Use the HG38 checkpoint for human-genome applications and the Rice checkpoint for rice/plant genomic applications. Both checkpoints expose the same public OpticalDNA API.
Input format. OpticalDNA operates on rendered DNA document images. Raw DNA sequences can first be converted into OpticalDNA-compatible pages with VisualDNA.
The simplest downstream use is to extract a compact visual genomic representation. This path runs the visual encoder, projector, and page-fusion module, and does not execute the language Decoder.
from opticaldna import OpticalDNA
model = OpticalDNA(
"hxxiang/opticaldna-hg38-2048",
device="cuda",
)
features = model.extract_features(
"assets/640x640.png",
pooling="mean",
to_cpu=True,
)
print(type(features))
print(features.shape)Expected output:
<class 'torch.Tensor'>
torch.Size([1280])
extract_features(...) is an alias of extract_visual_features(...). Both return visual representations before the language Decoder.
For a multi-page DNA document, pass the pages in reading order:
features = model.extract_features(
["page1.png", "page2.png", "page3.png"],
pooling="mean",
to_cpu=True,
)
print(features.shape)
# torch.Size([1280])The list represents one multi-page document, not a batch. OpticalDNA fuses page-level representations before returning the document features.
Both released models use a 1,280-dimensional OpticalDNA representation space. The final tensor shape depends on the pooling strategy.
pooling |
Output shape | Description | Typical use |
|---|---|---|---|
"mean" |
[1280] |
Mean over all visual tokens | Recommended compact document embedding for linear probing / MLP heads |
"max" |
[1280] |
Dimension-wise maximum over visual tokens | Emphasizes strongly activated visual/genomic features |
"none" |
[N_visual_tokens, 1280] |
Keeps every fused visual token | Custom attention, token-level analysis, or user-defined pooling |
N_visual_tokens is not fixed: it depends on the number of pages and image/crop configuration. The feature width is always 1,280 for the released HG38 and Rice checkpoints.
Examples:
# Mean-pooled document representation.
feat_mean = model.extract_features(
"assets/640x640.png",
pooling="mean",
to_cpu=True,
)
print(feat_mean.shape)
# torch.Size([1280])
# Max-pooled document representation.
feat_max = model.extract_features(
"assets/640x640.png",
pooling="max",
to_cpu=True,
)
print(feat_max.shape)
# torch.Size([1280])
# Keep all visual tokens.
feat_tokens = model.extract_features(
"assets/640x640.png",
pooling="none",
to_cpu=True,
)
print(feat_tokens.shape)
# torch.Size([N_visual_tokens, 1280])Useful feature-extraction arguments:
| Argument | Default | Meaning |
|---|---|---|
pooling |
"mean" |
"mean", "max", or "none" |
to_cpu |
False |
If True, returns a detached CPU tensor; otherwise features remain on the model device |
base_size |
640 |
Global image size used by the released inference pipeline |
image_size |
640 |
Local/crop image size |
crop_mode |
False |
Enables the local-crop visual path when needed |
For most downstream genomic benchmarks, pooling="mean" with to_cpu=True is the simplest starting point.
OpticalDNA can also expose prompt-conditioned Decoder hidden states. Unlike pure visual features, this path executes the language Decoder.
from opticaldna import OpticalDNA, PromptGenerator, PromptLength, TaskType
model = OpticalDNA(
"hxxiang/opticaldna-hg38-2048",
device="cuda",
)
prompts = PromptGenerator()
prompt = prompts.build(
TaskType.T1_FULL_OCR,
PromptLength.SHORT,
sample={},
)
decoder_features = model.extract_decoder_features(
"assets/640x640.png",
prompt=prompt,
layer=-1,
pooling="mean",
image_tokens_only=True,
to_cpu=True,
)
print(prompt)
print(decoder_features.shape)Expected:
Free OCR.
torch.Size([1280])
Decoder feature options:
| Argument | Default | Meaning |
|---|---|---|
layer |
-1 |
Decoder hidden-state layer; -1 selects the final hidden state |
pooling |
"mean" |
"mean", "max", or "none" |
image_tokens_only |
True |
Keep only hidden states aligned with OpticalDNA image tokens |
to_cpu |
False |
Move the returned detached tensor to CPU |
With pooling="none" and image_tokens_only=True, the output has shape:
[N_image_tokens, 1280]
With image_tokens_only=False, the token dimension can also include prompt/text positions:
[N_sequence_tokens, 1280]
If prompt is omitted, OpticalDNA uses the short T1 prompt:
Free OCR.
Generation returns a Python str. The simplest generation example is:
text = model.generate(
"assets/640x640.png",
max_new_tokens=256, # Increase this value if a longer generated sequence is expected.
)
print(text)A typical Free-OCR output is a DNA sequence string:
AAGCCAAGAGTCTTCTAATATTTTACATTCACTAAGCAATATGAAAATT...
OpticalDNA exposes the six prompt families used during pretraining:
| Task | TaskType |
Purpose | Possible output |
|---|---|---|---|
| T1 | T1_FULL_OCR |
Read the full DNA document | DNA sequence text |
| T2 | T2_FULL_OCR_GROUNDING |
Read DNA and ground lines/regions | Sequence plus bounding boxes |
| T3 | T3_ROI_OCR |
OCR specified DNA regions | Sequence for each requested box |
| T4 | T4_MASK_COMPLETION |
Recover masked/occluded DNA | Predicted DNA plus region boxes |
| T5 | T5_SUBSEQ_LOCATE |
Locate a query subsequence | Matching bounding boxes, or [] |
| T6 | T6_CHR_CLASSIFICATION |
Chromosome classification | Chromosome label; primarily HG38/human-oriented |
Prompt lengths:
PromptLength.SHORT
PromptLength.MEDIUM
PromptLength.LONG
Examples:
from opticaldna import PromptGenerator, PromptLength, TaskType
prompts = PromptGenerator()
# T1: Free OCR
free_ocr_prompt = prompts.build(
TaskType.T1_FULL_OCR,
PromptLength.SHORT,
sample={},
)
# T5: subsequence localization
locate_prompt = prompts.build(
TaskType.T5_SUBSEQ_LOCATE,
PromptLength.MEDIUM,
sample={"query": "ACGTACGT"},
)
locate_output = model.generate(
"assets/640x640.png",
prompt=locate_prompt,
max_new_tokens=256,
)
print(locate_output)For T3 ROI OCR and T4 masked completion, provide bounding boxes in:
sample = {
"boxes": [
[img_id, x1, y1, x2, y2],
# ...
]
}For multi-page documents, img_id is the 0-based page index.
The Hugging Face repositories also expose the OpticalDNA methods through the standard Transformers custom-model interface.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "hxxiang/opticaldna-hg38-2048"
tokenizer = AutoTokenizer.from_pretrained(
model_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype="auto",
).cuda().eval()
# Pure visual features: language Decoder is not executed.
features = model.extract_features(
"assets/640x640.png",
pooling="mean",
to_cpu=True,
)
print(features.shape)
# torch.Size([1280])
# Build a prompt directly from the remote OpticalDNA model.
prompt = model.build_prompt(
"t1_full_ocr",
length="short",
)
# Prompt-conditioned Decoder features.
decoder_features = model.extract_decoder_features(
tokenizer,
"assets/640x640.png",
prompt=prompt,
layer=-1,
pooling="mean",
to_cpu=True,
)
# Decoder output.
text = model.generate_document(
tokenizer,
"assets/640x640.png",
prompt=prompt,
max_new_tokens=256,
)
print(text)OpticalDNA uses fail-fast checkpoint loading: trained checkpoint parameters are validated during loading rather than silently falling back to partially initialized weights.
OpticalDNA/
βββ opticaldna/ # Model, tokenizer, and public inference API
βββ src/ # Training and data-loading source code
β βββ pretrain_opticaldna.py # Pre-training entry point
β βββ opticaldna_data/ # Dataset, conversation builder, and data collator
βββ scripts/ # Standalone utility scripts
β βββ data/ # Data preparation utilities
βββ tests/ # Fast CI tests + released-checkpoint smoke tests
βββ .github/workflows/ # GitHub Actions CI
βββ assets/ # README figures and example DNA page
βββ environment.yml # Conda environment specification
βββ LICENSE
βββ README.md
The main directories are:
opticaldna/: model configuration, tokenizer files, and OpticalDNA model wrappers.src/: training entry point and data-loading modules used during pre-training.scripts/data/: standalone data preparation scripts for generating processed VisualDNA data.assets/: figures used in this README.
- CUDA Toolkit >= 11.8
- CUDA >= 7.5
- Triton: 3.4.0
- Unsloth 2025.10.12
- Transformers: 4.56.2
- Torch >= 2.6.0+cu118
Install the environment from environment.yml:
conda env create -f environment.yml
conda activate opticaldnaAlternatively, you can create the environment manually:
conda create -n opticaldna python=3.12 -y
conda activate opticaldna
pip install PyMuPDF img2pdf einops easydict addict Pillow
pip install flash-attn==2.7.3 --no-build-isolation --use-pep517 --no-deps
pip install pytabix pandas pyfaidx
pip install kipoiseq==0.5.2
pip install tensorflow==2.16.1
pip install selene-sdk==0.5.3 # Requires torch <= 2.3.1
pip install natsort==8.4.0
pip install pyBigWig==0.3.23
pip install decorator
pip install rdkit
pip install dictionary omegaconf safetensors
pip install timm==1.0.22 --no-deps
pip install tf-keras
pip install datasets==4.3.0
pip install trl==0.24.0 --no-deps
pip install transformers==4.56.2 --no-deps
pip install peft==0.17.1
pip install tokenizers==0.22.1
pip install "unsloth[cu118-torch260]"
# Optional: install PyTorch and related local wheels if needed
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu118
# download from https://pytorch-geometric.com/whl/torch-2.6.0%2Bcu118.html
pip install torch_cluster-1.6.3+pt26cu118-cp312-cp312-linux_x86_64.whl
pip install torch_scatter-2.1.2+pt26cu118-cp312-cp312-linux_x86_64.whl
pip install torch_sparse-0.6.18+pt26cu118-cp312-cp312-linux_x86_64.whl
pip install torch_spline_conv-1.2.2+pt26cu118-cp312-cp312-linux_x86_64.whl
pip install numpy==1.26VisualDNA is the rendering and dataset-building toolkit used by OpticalDNA to convert DNA sequences into images.
Install the public package from PyPI:
pip install visualdnaVerify the installation:
visualdna --versionVisualDNA is used in the following sections to render raw genomic sequences into the image-based inputs required by OpticalDNA.
OpticalDNA uses data generated by the VisualDNA rendering pipeline. Each dataset contains a raw/ directory for source files and a processed/ directory for rendered DNA document images, bounding boxes, and metadata.
A typical dataset directory is organized as follows:
/path/to/opticaldna_dataset/
βββ <dataset_name>/
βββ raw/
β βββ <dataset_name>.parquet # Source genome files, annotations, or metadata
β βββ ...
βββ processed/
βββ <render_config_name>/
βββ index.csv # Metadata index used by OpticalDNA
βββ images/ # Rendered DNA document images
βββ bbox/ # Bounding-box annotations
In the training scripts, set --dataroot to the parent data directory and --dataset to the dataset subdirectory name.
For example, if the dataset is stored as:
/path/to/opticaldna_dataset/hg38-2048/
then the corresponding arguments should be:
--dataroot /path/to/opticaldna_dataset
--dataset hg38-2048We provide the raw data links used to construct the OpticalDNA pre-training datasets. After downloading the raw files, place them under the corresponding raw/ directory before running the VisualDNA rendering pipeline.
| Dataset | Description | Raw Data Link | Expected Raw Directory |
|---|---|---|---|
| HG38 | Human reference genome data used for OpticalDNA pre-training | Download | /path/to/opticaldna_dataset/hg38-2048/raw/ |
| Rice | Rice genome data used for OpticalDNA pre-training | Download | /path/to/opticaldna_dataset/rice/raw/ |
The data processing utilities are placed under:
scripts/
βββ data/
βββ generate_processed.py
βββ add_raw_columns_to_processed_index.py
The processing workflow contains two steps:
- Generate the
processed/directory from raw genome files. - Add selected metadata columns from the raw file to the generated
processed/index.csv.
Step 1: Generate processed data
Run scripts/data/generate_processed.py to convert raw genome sequences into rendered DNA document images and bounding-box annotations.
python scripts/data/generate_processed.py \
--dataroot /path/to/opticaldna_dataset \
--dataset hg38-2048 \
--raw-format parquet \
--seq-columns seq \
--img-width 640 \
--img-height 640 \
--font-size 14 \
--line-spacing 1.6 \
--merge-pages \
--save-bbox \
--shard-size autopython scripts/data/generate_processed.py \
--dataroot /path/to/opticaldna_dataset \
--dataset rice \
--raw-format parquet \
--seq-columns seq \
--img-width 640 \
--img-height 640 \
--font-size 14 \
--line-spacing 1.6 \
--merge-pages \
--save-bbox \
--shard-size autoThe command above corresponds to the following VisualDNA configuration:
from visualdna.data import ShardedBuilder
from visualdna.render import BaseRenderConfig
config = BaseRenderConfig(
img_width=640,
img_height=640,
font_size=14,
line_spacing=1.6,
merge_pages=True,
save_bbox=True,
)
builder = ShardedBuilder(
dataroot="/path/to/opticaldna_dataset",
dataset="hg38-2048",
render_config=config,
seq_columns=["seq"],
raw_csv_url=None,
force_generate=False,
shard_size="auto",
raw_format="parquet",
)After this step, the processed files should be generated under:
/path/to/opticaldna_dataset/
βββ hg38-2048/
βββ raw/
βββ processed/
βββ <render_config_name>/
βββ index.csv
βββ images/
βββ bbox/
Step 2: Add metadata columns to processed/index.csv
Some downstream tasks require extra metadata columns, such as chromosome names. These columns can be copied from the raw file into the generated processed/index.csv.
Run scripts/data/add_raw_columns_to_processed_index.py after Step 1.
Add chr_name for HG38
python scripts/data/add_raw_columns_to_processed_index.py \
--dataroot /path/to/opticaldna_dataset \
--processed-dataset hg38-2048 \
--raw-dataset hg38-2048 \
--render-id <render_config_name> \
--raw-format parquet \
--key index \
--columns chr_nameAdd multiple columns
python scripts/data/add_raw_columns_to_processed_index.py \
--dataroot /path/to/opticaldna_dataset \
--processed-dataset hg38-2048 \
--raw-dataset hg38-2048 \
--render-id <render_config_name> \
--raw-format parquet \
--key index \
--columns chr_name,split,speciesHere:
--processed-datasetis the dataset whoseprocessed/index.csvwill be updated.--raw-datasetis the dataset where the source raw file is stored.--render-idis the generated render configuration directory name underprocessed/.--keyis the column used to match rows between raw data and processed data.--columnsspecifies one or more raw metadata columns to add toprocessed/index.csv.
The expected processed index path is:
/path/to/opticaldna_dataset/<processed-dataset>/processed/<render-id>/index.csv
The expected raw file path is:
/path/to/opticaldna_dataset/<raw-dataset>/raw/<raw-dataset>.parquet
Most pre-training workflows are expected to run on Linux servers. Windows cmd examples are also provided for users who prepare data locally.
Linux/macOS examples
Generate processed data
python scripts/data/generate_processed.py \
--dataroot /path/to/opticaldna_dataset \
--dataset hg38-2048 \
--raw-format parquet \
--seq-columns seq \
--img-width 640 \
--img-height 640 \
--font-size 14 \
--line-spacing 1.6 \
--merge-pages \
--save-bbox \
--shard-size autoAdd metadata columns
python scripts/data/add_raw_columns_to_processed_index.py \
--dataroot /path/to/opticaldna_dataset \
--processed-dataset hg38-2048 \
--raw-dataset hg38-2048 \
--render-id render_w640_h640_fs14_ls1.6_hash_f345fcfc \
--raw-format parquet \
--key index \
--columns chr_nameWindows CMD examples
If you run the commands in Windows cmd, use ^ for line continuation.
Generate processed data
python scripts\data\generate_processed.py ^
--dataroot F:\path\to\opticaldna_dataset ^
--dataset hg38-2048 ^
--raw-format parquet ^
--seq-columns seq ^
--img-width 640 ^
--img-height 640 ^
--font-size 14 ^
--line-spacing 1.6 ^
--merge-pages ^
--save-bbox ^
--shard-size autoAdd metadata columns
python scripts\data\add_raw_columns_to_processed_index.py ^
--dataroot F:\path\to\opticaldna_dataset ^
--processed-dataset hg38-2048 ^
--raw-dataset hg38-2048 ^
--render-id render_w640_h640_fs14_ls1.6_hash_f345fcfc ^
--raw-format parquet ^
--key index ^
--columns chr_nameraw/{dataset}.parquetshould exist before runninggenerate_processed.py.- The
processed/directory is generated automatically by VisualDNA. - The
render-idshould match the directory name generated underprocessed/. - The same scripts can be used for HG38, rice, or any other dataset by changing parser arguments.
- Keep these scripts under
scripts/data/because they are command-line data preparation utilities rather than model or training modules.
OpticalDNA is initialized from the DeepSeek-OCR checkpoint. Download the checkpoint from this link and place the extracted checkpoint directory under opticaldna/.
conda activate opticaldna
model_name=./opticaldna
output_dir=./outputs/pretrain_opticaldna/hg38
dataroot=/path/to/opticaldna_data
dataset=hg38-2048
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
OMP_NUM_THREADS=8 MKL_NUM_THREADS=8 \
PYTHONUNBUFFERED=1 \
python -m torch.distributed.run --standalone --nproc_per_node=8 --master_port=29501 \
src/pretrain_opticaldna.py --backend nccl \
--dataloader_num_workers 12 --dataloader_prefetch_factor 4 \
--per_device_train_batch_size 8 --save_total_limit 10 --save_steps 1000 \
--warmup_steps 20000 --output_dir ${output_dir} \
--dataroot ${dataroot} --dataset ${dataset} \
--model_name ${model_name} \
--task_sampling "p_t1=0.17,p_t2=0.17,p_t3=0.17,p_t4=0.17,p_t5=0.17,p_t6=0.15" \
--tail_truncation "enabled=true,base_delete_ratio=0,max_delete_ratio=0.98" \
--line_span_cfg "min_n_base=1,max_n_base=8,min_n_sample=1,max_n_sample=3,unique_lines=true" \
--subseq_locate_cfg "min_len=6,max_len=32,allow_overlap=true" \
--lora_r 128 --trainable_modules_to_save page_fusion_layer \
--learning_rate 5e-4 --max_steps 281250 --annealed_sampler_total_steps 1 \
--lora_projectorconda activate opticaldna
model_name=./opticaldna
output_dir=./outputs/pretrain_opticaldna/rice
dataroot=/path/to/opticaldna_data
dataset=rice-2048 # we use NIP-T2T_w2048_o1920_SeqCase.UPPER
max_steps=235405
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
OMP_NUM_THREADS=8 MKL_NUM_THREADS=8 \
PYTHONUNBUFFERED=1 \
python -m torch.distributed.run --standalone --nproc_per_node=8 --master_port=29501 \
src/pretrain_opticaldna.py --backend nccl \
--dataloader_num_workers 12 --dataloader_prefetch_factor 4 \
--per_device_train_batch_size 8 --save_total_limit 10 \
--warmup_steps 20000 --output_dir ${output_dir} \
--dataroot ${dataroot} --dataset ${dataset} \
--model_name ${model_name} \
--task_sampling "p_t1=0.17,p_t2=0.17,p_t3=0.17,p_t4=0.17,p_t5=0.17,p_t6=0.15" \
--tail_truncation "enabled=true,base_delete_ratio=0,max_delete_ratio=0.9" \
--line_span_cfg "min_n_base=1,max_n_base=8,min_n_sample=1,max_n_sample=3,unique_lines=true" \
--subseq_locate_cfg "min_len=6,max_len=32,allow_overlap=true" \
--lora_r 128 --trainable_modules_to_save page_fusion_layer \
--learning_rate 5e-4 --max_steps ${max_steps} --lora_projector --lora_decoder- Replace example paths with local paths when needed.
- Large model weights and generated datasets are hosted externally and are not committed to GitHub.
- Loading custom OpticalDNA checkpoints through Transformers requires
trust_remote_code=True.
Fast tests run automatically on every GitHub push and pull request. They check syntax, the Hugging Face configuration/public API contract, and fail-fast checkpoint validation without downloading model weights. Pull requests and main also run a lightweight CPU runtime-import check.
python -m pytest tests/unitBefore a model release, run the real-checkpoint smoke test:
python tests/smoke/test_checkpoint_loading.py \
--model hxxiang/opticaldna-hg38-2048 \
--device cuda \
--image assets/640x640.pngSee tests/README.md for the rice checkpoint and optional generation test.
This repository is released under the MIT License. Parts of the model implementation are adapted from third-party open-source projects; please also follow the corresponding upstream licenses and notices where applicable.


