top of page


What 13 Milliseconds Buys You
We changed one setting, the floating-point format a detection model runs in; on a service already deployed to a Jetson board, and cut its latency in half. Here's the full accounting: what FP16 actually costs, what it doesn't, and exactly where the savings run out. The service in question detects a set of product-condition classes on a live line, using a YOLOv8-style model exported to ONNX and served behind a FastAPI endpoint. It runs on an NVIDIA Jetson (JetPack 6.2, Ampere-c

Rajeev Gadgil
Aug 243 min read
Â
Â
Â
bottom of page

