Transfer capabilities
Investigate how smaller models can learn useful capabilities from larger teacher models.
AI within a smaller footprint
Ideas for efficient AI on small, memory-constrained GPUs. Model distillation, quantization, and inference optimization.
Investigate how smaller models can learn useful capabilities from larger teacher models.
Explore ways to reduce inference memory requirements and improve execution efficiency.
Compare model quality, latency, throughput, and GPU memory use on constrained hardware.