Open to AI Engineer roles — available immediately, no notice period to serve
Back to work experience

Vision-Language Model Fine-Tuning for Food Recognition

TapHealth · AI Engineer · Jul 2025 – Oct 2025

Shipped to production

Fine-tuned and compressed a vision-language model for on-device food recognition.

Led end-to-end fine-tuning of a vision-language model for food recognition: curated 300,000+ training samples with quantity-aware supervision, then quantized and packaged the model for edge deployment — compressing it roughly 68% (6GB to 1.9GB via INT4 GGUF) while keeping end-to-end inference latency at 6 seconds or less on-device.

300K+ samplesDataset
6GB → 1.9GB (~68%, INT4 GGUF)Compression
≤6sEdge latency
Qwen3-VL-2BLoRA / QLoRAUnslothGGUF / INT4KV-Cachellama.cpp