This repository contains the analysis artifact for the paper A Measurement Study of LLM Inference Trade-offs Across Edge-Continuum Hardware.
The paper studies how model choice, quantization, execution platform, latency, memory footprint, energy, and streamed-token delivery overhead affect LLM deployment decisions across the edge continuum. The analysis uses two result files produced by the benchmarking pipeline and recreates the main post-processing steps used for the paper figures.
.
├── data
│ ├── results.csv
│ └── results_chatgpt.csv
├── figures
├── Analysis.ipynb
├── README.md
└── requirements.txt
The notebook focuses on four parts of the paper analysis.
- Loading and validating the self-hosted and cloud-reference measurements.
- Comparing accuracy, model size, latency, and energy across models and devices.
- Computing accuracy-latency Pareto frontiers for the self-hosted deployments.
- Recomputing the Pareto frontier after adding effective streamed-token delivery overhead to server-side deployments.
The main notebook is available at Analysis.ipynb.
Create a Python environment and install the required packages.
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtOpen the documented notebook.
jupyter lab Analysis.ipynbThe code is written to work when launched either from the repository root.
data/results.csv contains self-hosted measurements for Orin, Server GPU, and Server CPU deployments. data/results_chatgpt.csv contains the GPT-4o cloud-reference run. The cloud reference is used for accuracy and latency comparison, but it is excluded from energy analysis because the API does not expose hardware-level power or utilization metrics.
If you use this repository, its datasets, or its analysis code, please cite the following paper:
@inproceedings{khatib2026WIMS,
title = {A Measurement Study of LLM Inference Trade-offs Across Edge-Continuum Hardware},
author = {Khatib, Maysam and Symeonides, Moysis and Trihinas, Demetris and Pallis, George and Dikaiakos, Marios D.},
booktitle = {Proceedings of the 16th International Conference on Web Intelligence, Mining and Semantics (WIMS)},
year = {2026}
}This repository is licensed under the Apache License, Version 2.0. See the LICENSE file for details.