-
Notifications
You must be signed in to change notification settings - Fork 101
Pull requests: InfiniTensor/InfiniLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
perf(nvidia): reduce refactored inference overhead
#565
opened Sep 8, 2026 by
voltjia
Collaborator
Loading…
24 of 40 tasks
feat(hygon): enable modern Infini stack inference
#564
opened Sep 8, 2026 by
gongchensu
Collaborator
Loading…
2 of 49 tasks
feat(nvidia): support MiniMax-Text-01 (hybrid Lightning/full-attention MoE, MXFP4, TP/PP)
#563
opened Sep 4, 2026 by
Kritace
Loading…
29 of 34 tasks
feat(nvidia): add reusable GGUF Route B support for Qwen3.5
#559
opened Sep 3, 2026 by
xindongliu594
Loading…
32 of 48 tasks
feat(ascend): enable InfiniOps flash attention
#558
opened Sep 3, 2026 by
baominghelly
Contributor
Loading…
21 of 32 tasks
feat(server): add agent support with tool-call and reasoning parsing
#554
opened Sep 1, 2026 by
rubik-hua
Contributor
Loading…
feat: support Ktransformers, CPU-GPU MoE offload via FusedMoE layer
#548
opened Aug 21, 2026 by
whjthu
Contributor
Loading…
37 of 49 tasks
fix(cuda-graph): keep replay metadata dynamic across tensor-parallel ranks
#540
opened Aug 15, 2026 by
junjiewang253-ctrl
Loading…
feat: add aclnnMatmulAllReduce fusion in InfiniLM for Ascend RowParallelLinear
#533
opened Aug 11, 2026 by
ShaneWoof
Contributor
Loading…
feat(hygon): add Qwen3-235B-A3B BF16/W8A8 inference support
#532
opened Aug 7, 2026 by
qinyiqun
Contributor
Loading…
49 tasks
feat(engine): overlap decode steps with asynchronous token handoff
#524
opened Aug 3, 2026 by
qinyiqun
Contributor
Loading…
feat: add Qwen3.6 MoE model support
#521
opened Jul 31, 2026 by
qinyiqun
Contributor
Loading…
49 tasks
perf(server): coalesce streaming SSE output
#517
opened Jul 28, 2026 by
wooway777
Collaborator
Loading…
49 tasks
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-08-08.