Tiny AI Inference

AI within a smaller footprint

Capable models.
Smaller GPUs.

Ideas for efficient AI on small, memory-constrained GPUs. Model distillation, quantization, and inference optimization.

01 / DISTILL

Transfer capabilities

Investigate how smaller models can learn useful capabilities from larger teacher models.

02 / OPTIMIZE

Reduce the footprint

Explore ways to reduce inference memory requirements and improve execution efficiency.

03 / EVALUATE

Measure the tradeoffs

Compare model quality, latency, throughput, and GPU memory use on constrained hardware.