Default Branch

41abbfd599 · qwen4exp: enable rms_norm + mul fusion (#28896) · Updated 2026-09-14 16:16:30 +02:00

Branches

187276d8d1 · ui : model download pipeline · Updated 2026-09-07 22:10:09 +02:00

117
7

06b6ea69dd · ui : Hugging Face Hub data layer · Updated 2026-09-07 22:10:09 +02:00

117
5

7f59290fd0 · ui : optional sanitized raw HTML in markdown · Updated 2026-09-07 22:10:09 +02:00

117
8

e2e504df5d · ui : model memory-fit estimation · Updated 2026-09-07 22:10:09 +02:00

117
6

0809c495ba · ui : model id grammar for sidecars, quants and capability parsing · Updated 2026-09-07 22:10:09 +02:00

117
4

ca5c3be7f6 · ui : shared model display primitives · Updated 2026-09-07 22:10:09 +02:00

117
9

cdc2053bf7 · server : fix deadlock when removing a finished download · Updated 2026-09-07 22:10:08 +02:00

117
2

a56cfe1bcf · common : resolve <quant>-<sidecar> download tags and list cached sidecars · Updated 2026-09-07 22:10:08 +02:00

117
1

936adc32b9 · ui : type-safe API types, fetch helpers and download-ready models store plumbing · Updated 2026-09-07 22:10:08 +02:00

117
3

f01a75498e · CUDA: pick MMQ tile size against ncols_opt set on the host side · Updated 2026-09-07 16:39:17 +02:00

119
2

2265b7fa41 · Revert "CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (#24546)" · Updated 2026-09-07 16:19:36 +02:00

120
1

eee2fa9c79 · server: keep a queued model out of the victim pool until its waiters leave · Updated 2026-09-07 13:39:02 +02:00

131
2

3f6205741d · llama: properly handle KV on training · Updated 2026-09-07 01:00:51 +02:00

138
1

921c3e1aac · llama-context: add warning if sm tensor is used with one device · Updated 2026-09-06 14:21:21 +02:00

160
5

f0f9902d4c · use ggml_cuda_syncwarp · Updated 2026-09-06 10:08:55 +02:00

150
2

1cd349c00d · refactor : address review remarks · Updated 2026-09-06 03:29:20 +02:00

148
15

5150132b9a · metal : add remaining fa-vec tunings for M2 Max · Updated 2026-09-05 23:16:07 +02:00

148
1

4c31a6aac2 · rename unknown to none · Updated 2026-09-05 21:52:33 +02:00

149
2

5b1a8c2364 · cuda: top-k MoE should always fire · Updated 2026-09-05 10:20:32 +02:00

150
1

2e270db86a · Use fp8 intrinsics · Updated 2026-09-04 21:56:11 +02:00

156
5