GitHub / ggml-org/llama.cpp issues and pull requests
#29192 - Feature Request: Add opt-in CLI flags to llama-server
Issue -
State: open - Opened by muvo243 16 days ago
Labels: enhancement
#29191 - streams experts from qwen 3.8 moe with gpu prefill on vulkan and hip
Pull Request -
State: open - Opened by Atomic-Germ 16 days ago
- 2 comments
Labels: Vulkan, server, ggml
#29190 - ggml-sycl: remove dpct (SYCLomatic) emulation layer, switch to out-of-order queues with native sycl::event dependencies
Pull Request -
State: closed - Opened by Asahi-Prv 16 days ago
- 2 comments
Labels: devops, ggml, SYCL
#29189 - Optimizing Gemma4 graph for prompt processing
Pull Request -
State: open - Opened by Meet91721 16 days ago
Labels: model
#29188 - Eval bug: SIGSEGV in token-counting routes when the request arrives while the server is sleeping (stale vocab/mctx captured before the wake barrier)
Issue -
State: open - Opened by dreamsqk 16 days ago
- 1 comment
#29187 - CUDA: Fuse GDN alpha/beta projections and MMVQ small-batch epilogues
Pull Request -
State: open - Opened by ynankani 16 days ago
Labels: testing, ggml, CUDA
#29186 - SYCL: Q8_0 DMMV ESIMD and MMVQ wide load
Pull Request -
State: open - Opened by cwriter 16 days ago
Labels: ggml, SYCL
#29185 - [OPENVINO] Q1, Q2 quantization plus shape fixes for Bonsai 8B
Pull Request -
State: open - Opened by PiotrKrzem 16 days ago
Labels: ggml, OpenVINO
#29184 - CUDA: fuse shared experts into MMVQ
Pull Request -
State: open - Opened by am17an 16 days ago
Labels: testing, ggml, CUDA
#29183 - [SYCL] DeltaNet kernel produces token soup on long prompts (Qwen3.6, Qwen3.8, Nemotron)
Issue -
State: open - Opened by baustin27 16 days ago
#29182 - vulkan: MOE aware mul_mat_id tile selection
Pull Request -
State: open - Opened by Ankk98 16 days ago
Labels: Vulkan, ggml
#29181 - cuda : add grouped experts top-k fusion
Pull Request -
State: open - Opened by CISC 16 days ago
- 9 comments
Labels: testing, ggml, CUDA
#29180 - llama: enable shape-aware sampler probing
Pull Request -
State: open - Opened by aparmp-quic 16 days ago
- 1 comment
Labels: testing
#29179 - qwen3tts : guard speaker_encoder_config patch for CustomVoice variant
Pull Request -
State: open - Opened by SIDDARTHAREDDY8 16 days ago
Labels: testing, conversion
#29178 - common: validate --spec-draft-n-min argument bounds
Pull Request -
State: open - Opened by MirkoKiwi 16 days ago
#29177 - pyproject : add linux platform marker to uv torch source (#29176)
Pull Request -
State: open - Opened by Yezat 16 days ago
#29176 - Misc. bug: pyproject.toml: uv sync fails on macOS due to unconditional PyTorch CPU index
Issue -
State: open - Opened by Yezat 16 days ago
Labels: bug-unconfirmed
#29175 - llama-server: one prefill chunk blocks decode for all sequences — co-tenant stalls = ceil(prompt/n_batch), so `-b = n_parallel x -ub` maximises them
Issue -
State: open - Opened by oscarmherrera 16 days ago
#29174 - Eval bug: Qwen3.8-Flash-Next MTP draft GGUFs fail to load on b11058 (tensor not found) — no MTP path on Pascal
Issue -
State: open - Opened by podolchakagency 16 days ago
- 1 comment
Labels: bug-unconfirmed
#29173 - CUDA: Handle compute type for NVFP4 on cublass path
Pull Request -
State: open - Opened by ynankani 16 days ago
Labels: ggml, CUDA
#29172 - Eval bug: CUDA decode+prefill collapse at deep KV position on qwen35 hybrid (~5.5x tg, ~21x pp at d154855)
Issue -
State: open - Opened by hot-YUser 16 days ago
#29171 - sycl: accelerate GLM MLA prefill with MKL flash attention
Pull Request -
State: open - Opened by anantshri 16 days ago
- 1 comment
Labels: testing, ggml, SYCL
#29170 - ggml-cpu: probe _pdep_u64 for BMI2 on MSVC
Pull Request -
State: open - Opened by dontgitit 16 days ago
- 1 comment
Labels: ggml
#29169 - metal : support arbitrary hc in dsv4_hc_pre
Pull Request -
State: closed - Opened by ggerganov 16 days ago
Labels: testing, ggml, Apple Metal
#29167 - [SYCL] Gated DeltaNet produces token soup on long prompts (greedy fails, dense models fine to 131K+)
Issue -
State: open - Opened by baustin27 16 days ago
- 1 comment
#29166 - qwen4exp: fix per-block bias indexing when a unified cache holds several sequences
Pull Request -
State: open - Opened by akionux 17 days ago
- 1 comment
#29164 - Misc. bug: Model loads into mobile GPU CUDA0, but inference happens on iGPU
Issue -
State: open - Opened by kaosmaja 17 days ago
- 4 comments
Labels: bug-unconfirmed
#29161 - common/peg : handle invalid utf-8 sequences in the AST
Pull Request -
State: closed - Opened by aldehir 17 days ago
- 4 comments
Labels: testing
#29155 - ggml-cuda : convert contiguous tensors four elements at a time
Pull Request -
State: open - Opened by pwilkin 17 days ago
Labels: ggml, CUDA
#29152 - CUDA: tune FA for Gemma 4 on Ampere or newer
Pull Request -
State: closed - Opened by JohannesGaessler 17 days ago
- 5 comments
Labels: ggml, merge ready, CUDA
#29151 - model : add Ling 3.0 VL (BailingMoeV3VL) support
Pull Request -
State: open - Opened by aetherbird 17 days ago
Labels: model, testing, mtmd, conversion
#29146 - [ENH] server: tools: add "apptainer:" runtime support
Pull Request -
State: open - Opened by farhi 17 days ago
- 1 comment
Labels: server
#29139 - vulkan: hide internal symbols to prevent duplicate-dlopen state destr…
Pull Request -
State: open - Opened by ewintr 17 days ago
- 2 comments
Labels: Vulkan, ggml
#29136 - metal : fix deprecation warnings from macOS 27 SDK
Pull Request -
State: open - Opened by nikwen 17 days ago
Labels: ggml, merge ready, Apple Metal
#29133 - test-llama-archs : make tensor data stdev configurable and improve help
Pull Request -
State: open - Opened by ggerganov 17 days ago
Labels: testing
#29132 - sycl : support gated DSV4_HC_PRE and optional HC_POST comb matrix
Pull Request -
State: open - Opened by cwriter 17 days ago
Labels: ggml, merge ready, SYCL
#29130 - Feature Request: Slow prompt processing when MoE experts are read from disk via mmap - an optimization using O_DIRECT
Issue -
State: open - Opened by A233S 17 days ago
- 2 comments
Labels: enhancement
#29119 - Eval bug: can not run models
Issue -
State: open - Opened by mynameisken 17 days ago
- 1 comment
Labels: bug-unconfirmed
#29108 - ui: Fix mobile breakpoint + content overflow issues
Pull Request -
State: closed - Opened by allozaur 18 days ago
- 1 comment
Labels: merge ready, server/ui
#29107 - sycl: IQ3 code reorder
Pull Request -
State: open - Opened by anantshri 18 days ago
- 1 comment
Labels: testing, ggml, merge ready, SYCL
#29102 - vocab : match longest leftmost added token in added tokens parser
Pull Request -
State: open - Opened by fairydreaming 18 days ago
- 2 comments
#29097 - server: serialize vocab_type as integer in /v1/models (#29091)
Pull Request -
State: open - Opened by NolanMasc17 18 days ago
- 6 comments
Labels: server
#29094 - metal: add F16 input to the FWHT
Pull Request -
State: closed - Opened by bri-prism 18 days ago
- 4 comments
Labels: testing, ggml, merge ready, Apple Metal
#29092 - HIP/ROCm — fused Gated Delta Net op carries recurrent state across requests on a reused server slot (qwen35 / qwen35moe); earlier prompts' text is emitted verbatim in later completions
Issue -
State: open - Opened by jgoellermaximus 18 days ago
- 7 comments
Labels: bug-unconfirmed
#29091 - Misc. bug: [llama-server] /v1/models metadata serializes "vocab_type" as boolean true instead of integer enum
Issue -
State: open - Opened by stonybeach 18 days ago
- 4 comments
Labels: bug-unconfirmed
#29085 - llama: reserve graph nodes for LFM2 recurrent rollback
Pull Request -
State: open - Opened by cjl4hd 18 days ago
- 3 comments
Labels: testing
#29075 - metal : key the fa-vec tuned table by family instead of SKU
Pull Request -
State: open - Opened by forforever73 18 days ago
- 1 comment
Labels: documentation, examples, ggml, Apple Metal
#29073 - Feature Request: Passing environment variables to child processes in router mode
Issue -
State: open - Opened by dfriehs 18 days ago
- 1 comment
Labels: enhancement
#29060 - Update embeddings server: return HTTP 400 for invalid embedding requests
Pull Request -
State: open - Opened by SamMalayek 19 days ago
- 5 comments
Labels: server
#29054 - Vulkan: deterministic GPU hang + SIGABRT on Intel Gen9.5 iGPU (UHD 630) with q8_0 KV cache — vk::Error escapes as std::terminate
Issue -
State: open - Opened by NaustudentX18 19 days ago
- 1 comment
#29029 - metal : gate mul_mm_id src1 rescale behind ggml_prec
Pull Request -
State: open - Opened by mdegans 19 days ago
- 6 comments
Labels: testing, ggml, Apple Metal
#29027 - fit : optimize for dense models
Pull Request -
State: open - Opened by John-194 19 days ago
- 4 comments
#29022 - Feature Request: Fast Tool Gating & Single-Pass Selection via Prefill Logit Slicing
Issue -
State: open - Opened by mattepiu 19 days ago
- 4 comments
Labels: enhancement
#28990 - Feature Request: Performance Improvements on SYCL
Issue -
State: open - Opened by anantshri 20 days ago
- 4 comments
Labels: enhancement
#28985 - sycl : do not use slow oneDNN reference matmul and fattn
Pull Request -
State: open - Opened by lslusarczyk 20 days ago
- 2 comments
Labels: ggml, SYCL
#28976 - ggml-webgpu: add fused gated_delta_net + cpy
Pull Request -
State: open - Opened by yomaytk 20 days ago
- 1 comment
Labels: ggml, merge ready, WebGPU
#28968 - llama-bench : add --repack switch option
Pull Request -
State: open - Opened by truecoder34 21 days ago
- 6 comments
Labels: documentation, examples
#28967 - Add nccl support for multi-node tensor parallelism
Pull Request -
State: open - Opened by zyang-dev 21 days ago
- 8 comments
Labels: ggml, CUDA
#28938 - server : do not forward --api-key-file to router-spawned child instances
Pull Request -
State: open - Opened by nandan2003 21 days ago
- 4 comments
Labels: server
#28931 - sycl: extend MMVQ GLU fusion to mixed quant types; add rms_norm+scale and ssm_conv+silu fusions
Pull Request -
State: open - Opened by anantshri 21 days ago
- 7 comments
Labels: ggml, SYCL
#28913 - server: fix model eviction race in router(#28698)
Pull Request -
State: open - Opened by ac-mmi 22 days ago
- 9 comments
Labels: server
#28895 - sycl : pinned memory uses right device context instead of 0
Pull Request -
State: open - Opened by lslusarczyk 22 days ago
- 1 comment
Labels: ggml, merge ready, SYCL
#28845 - vocab : add pre-tokenizer for fraunhofer-iis/elmod-2.7b-it model
Pull Request -
State: open - Opened by fairydreaming 23 days ago
- 22 comments
Labels: conversion
#28840 - Eval bug: Vulkan device lost (crash) when processing images — CLIP graph flash-attention auto-enabled on unsupported backend
Issue -
State: open - Opened by jacycatlin 23 days ago
- 1 comment
Labels: bug-unconfirmed
#28832 - fix(mamba) : make time-step projection input contiguous
Pull Request -
State: closed - Opened by abetlen 23 days ago
Labels: model
#28770 - CUDA: enable sparse fa for qwen4
Pull Request -
State: closed - Opened by am17an 25 days ago
- 3 comments
Labels: model, testing, ggml, merge ready, CUDA
#28752 - Misc. bug: Severe drop in prompt processing speed after b10780 on Vulkan, RDNA3
Issue -
State: open - Opened by ComputerGuy4157 25 days ago
- 9 comments
Labels: bug-unconfirmed
#28734 - Eval bug: qwen4exp (Qwen3.8-Flash-Next) CUDA: decode slows linearly with context
Issue -
State: open - Opened by lukolszewski 25 days ago
- 5 comments
Labels: bug-unconfirmed
#28682 - chat: add dedicated Ling 3.0 (Bailing V3) parser
Pull Request -
State: closed - Opened by aetherbird 26 days ago
- 14 comments
Labels: testing
#28613 - HIP: tune MMVQ batch thresholds on RDNA3.5
Pull Request -
State: open - Opened by SimonTeixidor 28 days ago
- 4 comments
Labels: ggml, CUDA
#28536 - CUDA: Follow up of #25635, refactoring FA shared smem swizzle
Pull Request -
State: open - Opened by ynankani 29 days ago
- 23 comments
Labels: testing, ggml, CUDA
#28518 - Fixed json enum handling
Pull Request -
State: open - Opened by Silverside 30 days ago
- 6 comments
Labels: testing
#28498 - kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata
Pull Request -
State: open - Opened by eapache 30 days ago
- 1 comment
Labels: testing
#28439 - metal : wide query tile for kernel_flash_attn_ext
Pull Request -
State: open - Opened by masterFoad about 1 month ago
- 9 comments
Labels: documentation, testing, examples, ggml, Apple Metal
#28368 - cache_prompt reuse changes computed logprobs on a plain (non-hybrid) transformer — reproducible by toggling cache_prompt alone on an otherwise-identical warmed server
Issue -
State: open - Opened by jmpmachado about 1 month ago
- 1 comment
#28305 - graph : keep the backend sampling subgraph static across ubatches also helps fix CI issue
Pull Request -
State: open - Opened by ynankani about 1 month ago
#28261 - docs: document streaming tool calls and the Jinja requirement
Pull Request -
State: open - Opened by apollo-2006 about 1 month ago
- 1 comment
Labels: documentation
#28260 - Misc. bug: --ui-config-file requires to click "Reset to default" in settings to apply values
Issue -
State: open - Opened by MuAlphaOmegaEpsilon about 1 month ago
Labels: bug-unconfirmed
#28259 - webui: recognize audio/ogg MIME type in file uploads
Pull Request -
State: open - Opened by iambhuvan about 1 month ago
- 1 comment
Labels: server/ui
#28258 - ci : enable hf-jobs on self-hosted server-cuda
Pull Request -
State: closed - Opened by CISC about 1 month ago
- 3 comments
Labels: devops
#28257 - [WIP] Fix failing GitHub Actions job gpu-cuda
Pull Request -
State: closed - Opened by Copilot about 1 month ago
- 1 comment
#28256 - Pathological reads on N-gram embedding model
Issue -
State: open - Opened by IMbackK about 1 month ago
#28255 - Eval bug: Missing documentation howto setup a draft model via ENVvar
Issue -
State: open - Opened by bluemoehre about 1 month ago
Labels: bug-unconfirmed
#28254 - sycl : fix test-backend-ops CI break && restore Kronecker product FWHT support (#28016)
Pull Request -
State: open - Opened by philip-jingxin about 1 month ago
Labels: testing, ggml, SYCL
#28253 - vulkan: support type-aligned GET_ROWS
Pull Request -
State: open - Opened by jeffbolznv about 1 month ago
Labels: testing, Vulkan, ggml
#28252 - Whole-host hard lock during MTP draft catch-up prefill on multi-GPU tensor-split (b9745+)
Issue -
State: open - Opened by taylorsatula about 1 month ago
#28251 - Eval bug: MoE models crashes llama with CUDA Error on first or second prompt.
Issue -
State: open - Opened by LucasFur about 1 month ago
Labels: bug-unconfirmed
#28250 - mtmd: add mtmd_tokenize_from_parts()
Pull Request -
State: open - Opened by ngxson about 1 month ago
- 3 comments
Labels: documentation, testing, mtmd
#28249 - wiring up jinja input marking
Issue -
State: open - Opened by ngxson about 1 month ago
#28247 - Eval bug: [Vulkan] GGML_ASSERT(wg0 <= ctx->device->properties.limits.maxComputeWorkGroupCount on Intel Arc A770 when running Qwen 3.8 flash next
Issue -
State: open - Opened by danielstevens277-debug about 1 month ago
Labels: bug-unconfirmed
#28246 - ui: add close button to UI toasts
Pull Request -
State: open - Opened by agustinmista about 1 month ago
Labels: server/ui
#28245 - kv-cells: keep the used-cell set as a bitmap
Pull Request -
State: closed - Opened by ServeurpersoCom about 1 month ago
- 1 comment
#28244 - qwen4exp: attend the selected cells instead of masking the window
Pull Request -
State: open - Opened by ServeurpersoCom about 1 month ago
Labels: model
#28243 - models: Qwen3.8-Flash-Next MTP
Pull Request -
State: open - Opened by danielhanchen about 1 month ago
- 1 comment
Labels: model, ggml, CUDA, conversion
#28242 - server: fail initialization on async compute errors (#27309)
Pull Request -
State: open - Opened by ac-mmi about 1 month ago
Labels: server
#28241 - Eval bug: CUDA "illegal memory access" with -cmoe on Turing (sm_75) at exactly 94 prompt tokens
Issue -
State: open - Opened by Vyvrnc about 1 month ago
Labels: bug-unconfirmed
#28240 - MTMD: Fix qwen3-tts on Vulkan (get_rows index rows)
Pull Request -
State: open - Opened by ServeurpersoCom about 1 month ago
- 3 comments
Labels: mtmd
#28239 - [SYCL]Sysman free-memory query may be unavailable
Issue -
State: open - Opened by DDXDB about 1 month ago
- 1 comment
Labels: bug-unconfirmed
#28238 - quantize : run the imatrix fitter for Q4_1/Q5_1 without an imatrix
Pull Request -
State: open - Opened by sanmai about 1 month ago
Labels: ggml