Skip to content

About

Measurement study of LLM inference trade-offs across edge-continuum hardware, including accuracy, latency, model footprint, energy, and streaming-delay sensitivity analysis.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

Repository files navigation

A Measurement Study of LLM Inference Trade-offs Across Edge-Continuum Hardware

This repository contains the analysis artifact for the paper A Measurement Study of LLM Inference Trade-offs Across Edge-Continuum Hardware.

The paper studies how model choice, quantization, execution platform, latency, memory footprint, energy, and streamed-token delivery overhead affect LLM deployment decisions across the edge continuum. The analysis uses two result files produced by the benchmarking pipeline and recreates the main post-processing steps used for the paper figures.

What is included

.
├── data
│   ├── results.csv
│   └── results_chatgpt.csv
├── figures
├── Analysis.ipynb
├── README.md
└── requirements.txt

Analysis overview

The notebook focuses on four parts of the paper analysis.

  1. Loading and validating the self-hosted and cloud-reference measurements.
  2. Comparing accuracy, model size, latency, and energy across models and devices.
  3. Computing accuracy-latency Pareto frontiers for the self-hosted deployments.
  4. Recomputing the Pareto frontier after adding effective streamed-token delivery overhead to server-side deployments.

The main notebook is available at Analysis.ipynb.

Quick start

Create a Python environment and install the required packages.

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Open the documented notebook.

jupyter lab Analysis.ipynb

The code is written to work when launched either from the repository root.

Data files

data/results.csv contains self-hosted measurements for Orin, Server GPU, and Server CPU deployments. data/results_chatgpt.csv contains the GPT-4o cloud-reference run. The cloud reference is used for accuracy and latency comparison, but it is excluded from energy analysis because the API does not expose hardware-level power or utilization metrics.

Citation

If you use this repository, its datasets, or its analysis code, please cite the following paper:

@inproceedings{khatib2026WIMS,
  title     = {A Measurement Study of LLM Inference Trade-offs Across Edge-Continuum Hardware},
  author    = {Khatib, Maysam and Symeonides, Moysis and Trihinas, Demetris and Pallis, George and Dikaiakos, Marios D.},
  booktitle = {Proceedings of the 16th International Conference on Web Intelligence, Mining and Semantics (WIMS)},
  year      = {2026}
}

License

This repository is licensed under the Apache License, Version 2.0. See the LICENSE file for details.

About

Measurement study of LLM inference trade-offs across edge-continuum hardware, including accuracy, latency, model footprint, energy, and streaming-delay sensitivity analysis.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages