VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
-
Updated
Aug 12, 2026 - Python
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
Firefly: 大模型训练工具,支持训练Qwen2.5、Qwen2、Yi1.5、Phi-3、Llama3、Gemma、MiniCPM、Yi、Deepseek、Orion、Xverse、Mixtral-8x7B、Zephyr、Mistral、Baichuan2、Llma2、Llama、Qwen、Baichuan、ChatGLM2、InternLM、Ziya2、Vicuna、Bloom等大模型
Windows 本地多模态 AI 桌面桌宠,支持屏幕与音频感知。
Community model zoo for Apple Core AI (iOS/macOS 27): 62 models — LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting — each gated against its source model and shipped with the recipe that produced it. Downloadable from Hugging Face, runnable in one line of Swift via CoreAIKit. Plus benchmarks, Metal kernels, knowledge base.
Inferflow is an efficient and highly configurable inference engine for large language models (LLMs).
Explore LLM model deployment based on AXera's AI chips, It provides OpenAI‑compatible APIs and supports AX620E, AX650 and AX637 series chips.
A custom ComfyUI node for MiniCPM vision-language models, supporting v4, v4.5, and v4 GGUF formats, enabling high-quality image captioning and visual analysis.
Get more intelligence from every bit. Better quantization formats and smarter calibration let larger, stronger models run smoothly on the hardware you already own.
Advanced Android photo editor powered by on-device AI models (LaMa, EdgeSAM, MobileViT) with GPU-accelerated adjustments, voice commands, generative fill, and professional editing tools. Built with Jetpack Compose and ONNX Runtime for fast, offline-first image processing.
A minimal MiniCPM5-1B model inference engine built purely in Rust
CRNN/LPRNet/STNet + CTC + CCPD
为 DeepSeek V4.0 等单模态大模型装上眼睛和耳朵 —— 本地视觉+音频 MCP 服务器,基于 Ollama + MiniCPM-V 4.6 + faster-whisper,图片描述·视频分析·语音转文字
MaskClaw项目构建了一套基于端侧多模态大模型的自进化隐私防护架构,通过视觉脱敏与Agentic 规则演进,实现了 Agent 行为偏好与隐私防护策略的动态协同生长,确保了高危隐私数据在云端交互中的物理级安全与个体化适配。
PicQ: Demo for MiniCPM-o 4.5 to answer questions about images using natural language.
Generate realistic speech and clone voices in ComfyUI with VoxCPM custom nodes, featuring token-free TTS, style guidance, and LoRA training support.
Add a description, image, and links to the minicpm topic page so that developers can more easily learn about it.
To associate your repository with the minicpm topic, visit your repo's landing page and select "manage topics."