CacheSolidarity Artifact for Middleware '26
139
Artifact repository for the paper:
CacheSolidarity: Preventing Prefix Caching Side Channels in Multi-tenant LLM Serving Systems Panagiotis Georgios Pennas (IMDEA Software Institute, Universidad PolitΓ©cnica de Madrid), Konstantinos Papaioannou (Universidad PolitΓ©cnica de Madrid & IMDEA Software Institute), Marco Guarnieri (IMDEA Software Institute), Thaleia Dimitra Doudali (IMDEA Software Institute)
Proceedings of the ACM International Middleware 2026 Conference.
This repository/image contains the source code, scripts, and supporting materials required to reproduce the experimental evaluation presented in the paper. The artifact was created specifically to support the paper's evaluation and is intended to facilitate artifact evaluation, reproducibility, and reuse.
This artifact provides a complete, containerized framework to reproduce CacheSolidarity's security and performance evaluation. It is designed so that the experiments can be executed end-to-end via a single Docker image, without requiring manual modification of the source code.
In particular, the artifact provides:
run.sh).plots.ipynb) that parses the results and generates every figure/table reported in the paper.This artifact is submitted for evaluation under the following criteria:
For artifact evaluators: The recommended evaluation path is described in Getting Startedβ .
| Paper component | Artifact component | Reproduction instructions |
|---|---|---|
| Table 1 | src/pii_masking.py, src/bert_pii_rank.py | [Experiment 1] |
| Figure 6 | src/safeKV_privacy_eval.py | [Experiment 1] |
| Figure 3 | run.sh (Figure-3 block) | [Experiment 2] |
| Figure 7 | run.sh (Fig-7 block) | [Experiment 3] |
| Figure 8 | run.sh (Fig-8 block) | [Experiment 4] |
| Figure 9 | run.sh (Fig-9 block) | [Experiment 5] |
| Figure 10 | run.sh (Fig-10 block) | [Experiment 6] |
Disclaimer: Full-scale results across all ten evaluated LLM models (up to Llama3-70B on 4ΓH100 GPUs) require substantial multi-GPU resources and multi-day runtimes. Reviewers without access to comparable hardware are encouraged to reproduce the artifact on a subset of models/workloads (e.g., the smaller models and workload_avg), which demonstrates the same qualitative trends reported in the paper.
The main directories and files are organized as follows:
.
βββ run.sh
βββ plots.ipynb
βββ models/
βββ data/
βββ temp_results/
βββ plots/
βββ src/
| File | Description |
|---|---|
run.sh | Master script orchestrating the full evaluation (Phase 1: security pipeline, Phase 2: performance sweeps). |
plots.ipynb | Parses the results under temp_results/ and produces the final figures/tables to directly compare with the paper. |
models/ | Mount point for downloaded model weights (not bundled in the image; see Hardware Requirementsβ ). |
src/Contains the implementation of the security-evaluation pipeline and the performance-evaluation driver.
| File | Description |
|---|---|
build_datasets.py | Builds the privacy-evaluation datasets (Flan, UltraChat, ShareGPT, OASST1). |
patch_ai4privacy.py | Patches the ai4privacy PII-detection library used to identify secrets in prompts. |
pii_masking.py | Masks and replaces sensitive words in each dataset (Table 1). |
bert_pii_rank.py | Matches masked entries against candidate reconstructions (Table 1). |
safeKV_privacy_eval.py | Runs the PII-detection/recovery evaluation behind Figure 6, comparing CacheSolidarity against SafeKV. |
performance_evaluation_with_timestamps.py | Main experiment driver; run per workload/model/request-rate combination with --sys_mode set to APC, CacheSolidarity, Isolation, or SafeKV. |
temp_results/Raw results are written here, one subdirectory per experiment (e.g. fig7_workloads/, fig8_models/, fig9_RPS/, fig10_theta_performance/, Figure_3_exploitability/, pii_results/).
plots/Contains the figures generated from the experimental results by plots.ipynb. These are expected to match the trends of Figures 3, 6β10 and Table 1 of the paper.
This artifact requires an NVIDIA GPU and has been tested on the following configuration:
nvidia-container-toolkit installed to run the container with --gpus all.The experiments are expected to take approximately:
| Experiment |
|---|
| Experiment 1 (Security / Table 1, Figure 6) |
| Experiment 2 (Figure 3, exploitability) |
| Experiment 3 (Figure 7, cross-workload) |
| Experiment 4 (Figure 8, cross-model) |
| Experiment 5 (Figure 9, request-rate sensitivity) |
| Experiment 6 (Figure 10, threshold sensitivity) |
Runtime may vary depending on GPU model, model sizes evaluated, and system load.
All software dependencies are pre-installed in the Docker image; no host-level installation is required beyond Docker and the NVIDIA container runtime. Model weights are not bundled in the image due to size/licensing and must be downloaded separately (see below).
The following instructions provide the shortest path for an evaluator to verify that the artifact has been installed correctly.
docker pull panapenn/cachesolidarity:latest
Create the local directory and copy run.sh and plots.ipynb from the Docker image:
mkdir -p ~/cachesolidarity_run/models
docker create --name temp_container panapenn/cachesolidarity:latest
docker cp temp_container:/app/run.sh ~/cachesolidarity_run/run.sh
docker cp temp_container:/app/plots.ipynb ~/cachesolidarity_run/plots.ipynb
docker rm temp_container
Create a free Hugging Face account:
Accept the license for each gated model used in the evaluation:
Click "Agree and access repository" on each repository.
Install and authenticate the Hugging Face CLI:
pip install huggingface_hub
huggingface-cli login
Download all required models directly into the local models/ directory:
huggingface-cli download meta-llama/Llama-2-13b-chat-hf \
--local-dir ~/cachesolidarity_run/models/Llama-2-13b-chat-hf
huggingface-cli download meta-llama/Llama-2-7b-chat-hf \
--local-dir ~/cachesolidarity_run/models/Llama-2-7b-chat-hf
huggingface-cli download meta-llama/Llama-3.2-1B-Instruct \
--local-dir ~/cachesolidarity_run/models/Llama-3.2-1B-Instruct
huggingface-cli download google/gemma-3-4b-it \
--local-dir ~/cachesolidarity_run/models/gemma-3-4b-it
huggingface-cli download llava-hf/llava-onevision-qwen2-0.5b-ov-hf \
--local-dir ~/cachesolidarity_run/models/llava-onevision-qwen2-0.5b-ov-hf
huggingface-cli download Qwen/Qwen2.5-VL-3B-Instruct \
--local-dir ~/cachesolidarity_run/models/Qwen2.5-VL-3B-Instruct
After completing the setup, the directory should contain:
~/cachesolidarity_run/
βββ run.sh
βββ plots.ipynb
βββ models/
βββ Llama-2-13b-chat-hf/
βββ Llama-2-7b-chat-hf/
βββ Llama-3.2-1B-Instruct/
βββ gemma-3-4b-it/
βββ llava-onevision-qwen2-0.5b-ov-hf/
βββ Qwen2.5-VL-3B-Instruct/
Once the required models have been downloaded into ~/cachesolidarity_run/models, run:
docker run --gpus all \
--shm-size=16g \
-v ~/cachesolidarity_run/models:/app/models \
-v ~/cachesolidarity_run/temp_results:/app/temp_results \
-v ~/cachesolidarity_run/run.sh:/app/run.sh \
panapenn/cachesolidarity:latest
The local models/ directory is mounted to /app/models inside the container, so the container can access all downloaded models.
run.sh performs pre-flight verification on startup: it checks that every required model is present under ./cachesolidarity_run/models/<name>/config.json and aborts with a clear error naming any missing model.
The generated figures will be in ./plots. By comparing those with the paper figures as described in Relationship between the artifact and the paperβ , you can complete the evaluation of this artifact.
Pull image & mount models
β
Run run.sh (Phase 1: security pipeline)
β
Run run.sh (Phase 2: performance sweeps)
β
Generate figures and tables (plots.ipynb)
β
Compare with paper
β
Done!
This section provides a direct mapping between the main claims of the paper and the experiments available in the artifact. Experiments can also be run in batch via run.sh, as explained above, and directly compared to the figures of the paper without running each block separately.
Paper claim: CacheSolidarity secures the large majority of real-world prompts containing secrets (96.12%) across four datasets, closely matching the semantic-aware SafeKV baseline (96.95%), without relying on semantic analysis.
Artifact support: Evaluated using the Table-1/Table-2 and PII-detection blocks of run.sh (pii_masking.py β bert_pii_rank.py β safeKV_privacy_eval.py).
Expected result: Up to ~76% of secrets detected per dataset by the rule-based/LLM-based pipeline; CacheSolidarity secures β96% of prompts, matching SafeKV, with failures concentrated in first-entry/first-try edge cases.
Corresponding paper result: Table 1, Figure 6, Section 4.1/5.2.5.
Paper claim: The distinguishability of cache-hit/miss timing depends on prefix length, model size, request rate, and GPU hardware; the side channel weakens under high load and strengthens for larger models and longer prefixes.
Artifact support: Evaluated using the Figure-3 block of run.sh.
Expected result: KDE overlap decreases (side channel strengthens) with longer prefixes and larger models, and increases (side channel weakens) with higher request rate.
Corresponding paper result: Figure 3, Section 2.2.
Paper claim: CacheSolidarity stays within 5β10% of the insecure Prefix Caching baseline's TTFT and hit rate across workloads, while significantly outperforming User Cache Isolation and SafeKV.
Artifact support: Evaluated using the Fig-7 block of run.sh.
Expected result: CacheSolidarity tracks Prefix Caching closely (within ~6% TTFT) while User Cache Isolation and SafeKV show markedly lower hit rate / higher TTFT.
Corresponding paper result: Figure 7, Section 5.2.1.
Paper claim: CacheSolidarity's performance advantage over the baselines holds consistently across LLM families and sizes.
Artifact support: Evaluated using the Fig-8 block of run.sh.
Expected result: Consistent hit-rate/TTFT ordering (Prefix Caching β CacheSolidarity > SafeKV > User Cache Isolation) across all evaluated models.
Corresponding paper result: Figure 8, Section 5.2.2.
Paper claim: CacheSolidarity's hit rate remains stable and close to Prefix Caching as request rate increases, unlike SafeKV, whose hit rate degrades under load.
Artifact support: Evaluated using the Fig-9 block of run.sh.
Expected result: CacheSolidarity's hit rate stays flat as RPS increases; SafeKV's hit rate degrades and its TTFT rises more sharply.
Corresponding paper result: Figure 9, Section 5.2.3.
Paper claim: Increasing the Activator's threshold ΞΈ monotonically reduces attack success rate, with adjusting ΞΈ sufficient to fully prevent the evaluated prompt-stealing attack, at the cost of reduced hit rate/increased TTFT.
Artifact support: Evaluated using the Fig-10 block of run.sh.
Expected result: Attack success rate falls to 0%, with a corresponding gradual increase in TTFT and decrease in hit rate as ΞΈ grows.
Corresponding paper result: Figure 10, Section 5.3.
Results are saved under: ./temp_results/<experiment_name>/
After completing all experiments, this directory should include (among others):
| Output | Description |
|---|---|
privacy_evaluation/eval_data/ | Masked and ranked PII-evaluation datasets. |
pii_results/ | Per-dataset PII-detection/recovery results (Table 1, Figure 6). |
Figure_3_exploitability/ | Raw TTFT logs across models/request rates (Figure 3). |
fig7_workloads/ | Hit rate and TTFT logs across workloads and baselines (Figure 7). |
fig8_models/ | Hit rate and TTFT logs across models (Figure 8). |
fig9_RPS/ | Hit rate and TTFT logs across request rates (Figure 9). |
fig10_theta_performance/ | KDE overlap, attack success rate, hit rate, and TTFT logs across thresholds (Figure 10). |
The resulting figures will be generated at ./plots.
Expected result: The generated plots should reproduce the qualitative and quantitative trends shown in Figures 3, 6β10 and Table 1 of the paper. Small numerical differences may occur due to hardware differences and workload sampling randomness.
This artifact uses real user prompts from public datasets for academic research, reference, and evaluation purposes only.
ai4privacy and a BERT-based classifier to identify and mask personally identifiable information.Disclaimer: All datasets are used solely to evaluate the exploitability and mitigation of timing side channels in shared prefix caches. No secrets or personally identifiable information from these datasets are redistributed as part of this artifact beyond what is necessary to reproduce the paper's aggregate statistics.
This artifact is distributed under the Apache-2.0 license.
If you use this artifact, please cite the paper: CacheSolidarity: Preventing Prefix Caching Side Channels in Multi-tenant LLM Serving Systems.
The work by the authors at the IMDEA Software Institute was partially funded by the Madrid Regional Government through the CΓ©sar Nombela grant (2024-T1/COM-31302) and by the Comunidad de Madrid through the DATIA project, co-funded by the European Union's FEDER funds. Their work was also supported by grant PID2022-142290OB-I00, funded by MCIN/AEI/10.13039/501100011033 and FEDER, UE, and by grant CEX2024-001471-M, funded by MICIU/AEI/10.13039/ 501100011033. The work by J.O.I is partially supported by the EU EDGELESS project funded by the EU programme (agreement No. 101092950) and by the Smart Networks and Services Joint Undertaking (SNS JU) under the EU HE programme (agreement No. 101293102).
TODOTODOhttps://hub.docker.com/r/panapenn/cachesolidarity


Content type
Image
Digest
sha256:68f6b62b0β¦
Size
31.8 GB
Last updated
11 days ago
docker pull panapenn/cachesolidarity