mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-17 20:31:47 +02:00
The server caches the most recent compute graph per device so that GRAPH_RECOMPUTE can re-execute it without resending tensor data. The cached graph nodes hold direct pointers to backend buffers that were live at graph_compute() time. If any of those buffers is later released via FREE_BUFFER, the next GRAPH_RECOMPUTE re-executes the cached graph through the dangling pointers (use-after-free). The bug is reachable by an unauthenticated remote client. The dangling pointers point into chunks an attacker can reshape via subsequent ALLOC_BUFFER/SET_TENSOR commands, and the resulting read/write through the cached graph is sufficient to leak libc addresses and hijack the buffer iface vtable used by BUFFER_CLEAR, yielding remote code execution. Discard all cached graphs in free_buffer(). The existing null-check in graph_recompute() then rejects the request and the client falls back to GRAPH_COMPUTE on the next call. No protocol or API change.