Project Koru • Deep Dive Racecraft: Fine-Tuning Gemma 4 with DPO for Sub-150ms On-Device Coaching at 130 MPH How we distilled expert human motorsport coaching into a 12MB Gemma 4 LoRA adapter using Direct Preference Optimization (DPO)—cutting NPU TTFT to 149ms and locking sub-5-word root-cause cues at high speed. By Rabimba Karanjai • August 2026 • 10 min read Figure 1: Empirical DPO Fine-Tuning Loss convergence curve (0.693 down to 0.141), reward margin expansion (+1.768), and NPU performance benchmarks on Google Tensor G5 / NVIDIA GB10. P icturing a car barreling into Turn 3 at Sonoma Raceway at 130 MPH gives you an immediate sense of scale. At that speed, you are traveling 190.6 feet every single second (58.1 meters/sec). Your brain is processing visual apex marks, tire slip audio, steering effort, and lateral G-forces in split-second 200–300ms reaction windows. If an AI coach speaks a typical convers...
Racecraft · Technical Deep Dive · Sensor Fusion & Gemma Generalization Fusing 100+ On-Device Sensors at 130 MPH , And Why Gemma Didn't Need Fine-Tuning Inside Racecraft's real-time telemetry engine: multiplexing AiM CAN bus, RaceBox BLE, OBDLink, and 6-DOF IMU streams on Android, zero-shot corner doctrine generalization, and benchmarking Gemma 4 E2B on Tensor G5 silicon. A few days ago after sharing early results from Racecraft (Project Koru) , a question popped up in one of the comments that immediately caught my attention: "Curious how your team handled real-time fusion of 100+ sensors on-device. Did Gemma need race-specific fine-tuning, or did general coaching logic generalize?" It's the exact right question to ask. When people hear "AI race coach running on a phone in real time at Sonoma Raceway," they usually picture one of two extremes: either a brittle set of hardcoded if (speed < 40) statements, o...