Machine Learning Systems: Foundations, Scaling, Agentic AI, and Physical AI (Vols I–IV) • Harvard CS249r | https://mlsysbook.ai
-
Updated
Oct 3, 2026 - Python
Machine Learning Systems: Foundations, Scaling, Agentic AI, and Physical AI (Vols I–IV) • Harvard CS249r | https://mlsysbook.ai
TinyML & Edge AI: On-device inference, model quantization, embedded ML, ultra-low-power AI for microcontrollers and IoT devices.
Ahead-of-time compiler that fits ML models on embedded devices with hard memory budgets.
Python ML for training a custom on-device cry model (knowledge-distilled from YAMNet, INT8, deployed on ESP32-S3)
Hardware-aware face detection on Samsung GT-S7392 (ARM Cortex-A9)
Static INT8 AOT compiler that turns fixed PyTorch MCU models into standalone C11
ChatTLM: a 10.9M-parameter language model that runs on a TI-Nspire CX II CAS graphing calculator, with the code, data and measurements behind the paper "Letting the Tools Do the Math"
Six-class activity classifier with compact softmax modeling, int8 quantized C++ inference, reproducible benchmarking, and an ESP32 replay sketch using UCI HAR data.
Compress PyTorch models for edge devices — CPU-only, no GPU, no retraining. One function call.
Early-exit, full-model and static-baseline inference on RubikPi 3, with live-camera, HTP, latency and thermal evidence.
Don't Think It Twice: Exploit Shift Invariance for Efficient Online Streaming Inference of CNNs
Wearable fall-detection device with double-layer LSTM (97.8% accuracy), companion to published research
Deploy and manage ML models at the edge — OPC-UA integration, PLC connectivity, real-time inference on embedded hardware for sub-millisecond decisions
Cloud-to-edge image classification: ResNet-18 fine-tuned in Colab, exported to ONNX, deployed to an NVIDIA Jetson Nano for on-device inference. NVIDIA DLI Jetson AI Specialist final project.
A lightweight benchmark for tiny ML models on ESP32-class devices.
RIN Engine Runtime - Universal Inference Engine
Offline Sri Lankan Sign Language recognition on Raspberry Pi 5 — MediaPipe hand landmarks + RandomForest, streamed to a Flask dashboard. No cloud, no internet.
Multiposition heart sound analysis
Satellite-image CNN compressed 7.5x to a 0.83 MB INT8 ONNX model: 95.7% top-1, 0.76 ms p95 on one CPU core, and 95.9% retained under simulated radiation bit-flips that reduce FP32 to random guessing.
To associate your repository with the embedded-ml topic, visit your repo's landing page and select "manage topics."