Nemotron Tinker is an experimental Tinker-style API service for training and serving multiple LoRA adapters over one resident base model with NeMo AutoModel. It includes:
- A FastAPI service for adapter creation, SFT, RL-style LoRA updates, sampling, checkpointing, async jobs, tenant scoping, and worker metadata.
- A small Python SDK for experiment code.
- Named workload recipes for repeatable SFT and RL tests.
- A standalone Nemotron-themed adapter flywheel demo.
- An operator UI at
/uiwhen the service is running.
This repository is the product wrapper. AutoModel remains the training engine
and model integration layer for now, so keep an editable AutoModel checkout on
PYTHONPATH until the dependency is packaged.
The current focus is single-node V1 readiness. The service has been validated on Qwen and Nemotron Nano 30B A3B BF16, including two-adapter SFT, RL LoRA workloads, inference, save, restore, and UI/API smoke paths.
- Architecture: service shape, worker model, storage, distributed scope, and kernel scope.
- Setup: local development requirements, AutoModel linkage, GPU deployment prerequisites, and sanity checks.
- SFT Workflows: cross-entropy LoRA training, recipes, validated workloads, and sampling expectations.
- RL LoRA Workflows: rollout collection, RL losses, NeMo Gym bridge, and NeMo-RL bridge boundaries.
- Python SDK: client objects, server-owned training, sampling, OpenAI/Gym calls, and recipes.
src/nemotron_tinker/server.py: HTTP API and orchestration.src/nemotron_tinker/mixed_client.py: resident base-model and mixed-adapter LoRA execution.src/nemotron_tinker/sdk.py: Python SDK.src/nemotron_tinker/operator_ui.html: service UI.scripts/run_mixed_lora_server.py: service entry point.scripts/run_recipe.py: named workload runner.recipes/: SFT and RL workload configs.clients/: runnable API and workload clients.tools/: converters and benchmark helpers.demos/async_lora_demo.html: standalone animated demo.prototypes/: direct full-model smoke prototypes kept out of the main path.
Local development requirements:
uv sync --extra dev
export PYTHONPATH="$(pwd)/src:/path/to/Automodel"Deploy the live single-node GPU service and open a local UI tunnel:
GPU_HOST=<ssh-host> \
REMOTE_HOME_SCRATCH=/home/scratch.<user> \
scripts/deploy_gpu.sh startBy default the script looks for the Nemotron Nano 30B A3B BF16 model under
REMOTE_HOME_SCRATCH and starts the resident-only operator UI at:
http://127.0.0.1:18081/ui
Useful variants:
GPU_HOST=<ssh-host> scripts/deploy_gpu.sh status
GPU_HOST=<ssh-host> scripts/deploy_gpu.sh tunnel
GPU_HOST=<ssh-host> scripts/deploy_gpu.sh stop
GPU_HOST=<other-host> LOCAL_PORT=18082 scripts/deploy_gpu.sh startStart the service with a small model:
python scripts/run_mixed_lora_server.py \
--base-model Qwen/Qwen3-0.6B \
--scratch-dir /tmp/nemotron_tinker \
--cache-dir /tmp/nemotron_tinker_hf \
--host 127.0.0.1 \
--port 18080Run a quick SFT recipe:
python scripts/run_recipe.py qwen_sft_quick \
--base-url http://127.0.0.1:18080Open the operator UI:
http://127.0.0.1:18080/ui
Open the standalone flywheel demo directly in a browser:
demos/async_lora_demo.html
- Real model operations still run in the API process, not fully inside worker subprocesses.
- The production worker fleet, multi-node orchestration, and restart rehydrate path are not complete.
grouped_tritonis correct in tests but not the preferred performance path.- RL LoRA support is useful for service-level experiments but is not a full production GRPO/PPO training stack.
- Use NeMo-RL or Megatron Bridge for large dedicated distributed training jobs.
Future agent sessions should use the repo-local nemotron-tinker skill for
work on this prototype. It points Codex at the right docs, recipes, validation
commands, and GPU run conventions without loading this README as a giant
runbook.