The critical CVE-2026-105192 vulnerability in LMCache, used to accelerate large language model servers such as vLLM, allows an unauthenticated remote attacker to execute arbitrary code on the cache server when multiprocess mode is enabled; no patch is available yet. Organizations using LMCache with network access to the ZeroMQ socket must immediately move it to a local or trusted segment and review their Kubernetes deployment configurations.
Technical details of the vulnerability
The CVE-2026-105192 vulnerability affects LMCache starting from version 0.3.9 up to and including stable 0.5.5, as well as 0.5.6 release candidates and the current development branch. The CVE description is published in the CVE database (CVE-2026-105192 record), a formal entry in NVD may be available under the NVD CVE-2026-105192 template, and the technical analysis is presented in the JFrog advisory (JFrog JFSA-2026-001694382).
Key technical points:
- The vulnerability is triggered in LMCache’s multiprocess mode, where the cache runs as a separate server and LLM worker processes connect to it via the ZeroMQ library. The mode is described in the official LMCache documentation: multiprocess mode.
- By default, the LMCache server listens only on localhost. The risk appears when the operator explicitly specifies a routable address (for example, in a multi-node cluster), making the ZeroMQ socket accessible from the network.
- In multiprocess mode there is a message type that on the server side is deserialized via Python’s pickle mechanism before the message type is checked. The corresponding code can be seen in the LMCache repository (ipc_wrapper.py).
- Because:
- there is no authentication on the ZeroMQ socket;
- the pickle format allows arbitrary code execution during deserialization;
- the message type is checked after unpacking the data, —
any remote client capable of establishing a connection to this socket can send a specially crafted message and achieve execution of its code with the privileges of the LMCache process.
- According to JFrog, in the official LMCache container images the cache process runs with root privileges, which turns the deserialization problem into a full host takeover.
This configuration nuance exacerbates the risk: in the official Kubernetes deployment example for the LMCache multiprocess server (lmcache-daemonset.yaml) the server is configured to listen on all network interfaces, not just localhost. This example is often copied into production clusters without changes, making exposure more likely.
The LMCache developers have not yet published an official security advisory in the security advisories section, and no fixed release is available. The original material does not report exploitation of the vulnerability “in the wild”, but the attack vector is trivial, requires no authentication, and is not tied to a specific version of Python or ZeroMQ, which makes practical exploitation highly realistic.
Threat context: recurring ShadowMQ pattern
JFrog researchers note that the logical error matches a previously discovered series of vulnerabilities in AI inference infrastructure, informally dubbed ShadowMQ: data from an unauthenticated network socket is passed directly to pickle deserialization. This is not a one-off mistake in a single project, but a persistent anti-pattern in the design of high-performance IPC and network components in the AI stack.
It is notable that LMCache is not the only component in the chain. In the vLLM ecosystem, a separate denial-of-service vulnerability CVE-2026-105756 has already been fixed, where a single request with an invalid cache_salt value could crash the engine when using the LMCache multiprocess connector. Details are available in the vLLM project advisory (GHSA-2823-qmq8-rwvj). That issue only concerns process crashes, not code execution, but taken together it underscores that the combination of LLM engine + cache + interprocess communication has become a new attack surface.
Additional tickets in the LMCache GitHub repository (for example, issue #5507 and issue #5508) describe presumed problems with unauthenticated access to cache data and commands of various network services, but they have not yet been confirmed by maintainers, do not have CVEs, and have no fixes. For practical risk assessment they should be treated as an indicator of the product’s overall trust model rather than as formally verified vulnerabilities.
Impact assessment for organizations
The highest risk is borne by environments where the following factors coincide:
- use of LMCache in multiprocess mode with the ZeroMQ socket network-accessible from other hosts (including multi-node clusters);
- a microservices or Kubernetes architecture deployed following the official daemonset example, with the server bound to “all interfaces”;
- the LMCache container runs as root or has access to shared volumes containing access tokens, LLM engine configurations, and user request logs.
Potential consequences of successful exploitation:
- Full compromising access to the node running LMCache: installation of backdoors, cryptomining, using the node as a foothold for further lateral movement.
- Compromise of confidential data:
- access to cache contents (fragments of prompts, model responses, intermediate representations);
- access to the container/node file system with possible disclosure of orchestrator and LLM platform secrets.
- Disruption of AI service availability: arbitrary commands can stop processes, consume resources, change configuration, or use the cache as a channel for subsequent attacks.
The greatest threat is to:
- LLM service providers and internal AI platforms with a multi-tenant model;
- clouds and large clusters where the cache is shared across nodes and trust domains may overlap;
- any environments where LMCache is deployed without strict network segmentation (shared cluster networks, access from auxiliary services, development namespaces, etc.).
Practical security recommendations
1. Immediate prioritization and change of network exposure
- Identify all LMCache deployments:
- via dependencies of LLM platforms (for example, vLLM),
- via images in the container registry that contain the name
lmcache, - via Kubernetes configurations that use the daemonset example from the official repository.
- For all multiprocess server instances:
- restrict the listening address to localhost or strictly trusted cluster subnets;
- exclude the possibility of connections from user segments, general-purpose VPN zones, and the internet.
- Use Kubernetes network policies, dedicated security groups, or firewalls to minimize the set of hosts that have access to the ZeroMQ socket. At the same time, you should assume that any host with network access is capable of performing RCE.
2. Privilege reduction and strict runtime control
- Rebuild or override LMCache containers so that the process does not run as root:
- set an unprivileged user in the Dockerfile or in the PodSpec;
- disable unnecessary kernel capabilities and access to hostPath volumes.
- Enable container behavior controls (seccomp, AppArmor or similar mechanisms) to restrict system calls and complicate post-exploitation.
3. Patches and versions of related components
- Update vLLM to version 0.30.0 or newer to eliminate the DoS vulnerability CVE-2026-105756 in the LMCache connector (vLLM advisory).
- Monitor for a fixed LMCache release in the security advisories section and in the multiprocess mode documentation; plan an accelerated update as soon as a patch becomes available.
4. Monitoring and indicators of possible exploitation
JFrog does not publish explicit indicators of compromise (IP addresses, hashes of malicious images, etc.), but you can take the following steps:
- Analyze network and Kubernetes logs for:
- connections to the LMCache multiprocess server port from atypical subnets or namespaces;
- abnormal growth in the number of worker registrations or unexpected restarts of pods with LMCache.
- Check nodes running LMCache for:
- unknown processes and cron jobs;
- changes in container images, configuration volumes, and binaries.
- Enable additional auditing of commands and system calls in LMCache containers until an official fix is released.
The critical takeaway: until a patch for CVE-2026-105192 is available, the only effective protection remains strict limitation of network access to the LMCache multiprocess server and avoiding running it with root privileges. The first step should be an inventory of all LMCache deployments and immediate reattachment of the ZeroMQ socket to local or strictly controlled network segments.