Artificial Intelligence Research Intern
Qualcomm · Internship
• Engineered graph-level ONNX quantization using AIMET-ONNX for Large Vision Models (SSD-1B, LCM, LoRA), cutting base-model pipeline runtime by 37%.
• Preserved float precision in 8/16-bit conversion across UNet, VAE, and Text Encoders while maintaining >99.3% precision retention.
• Optimized multi-adapter LoRA quantization, boosting accuracy by 9.8% and reducing per-adapter latency by 10.4%.
• Built an ONNX Runtime (ORT) execution wrapper and test suite to seamlessly execute and evaluate QDQ models, driving a 1.6× pipeline speedup.