Back
This is Part 2 of our Reverse Pickup QC series. Read Part 1: Solving Reverse Pickup Verification with Vision AI.
In Part 1, we walked through how we arrived at a vision-language model for reverse-pickup QC we walked through how we arrived at a vision-language model for reverse-pickup QC in Indian e-commerce logistics reasoning through a doorstep photo comparison rather than scoring image similarity. Getting the model right was one problem. Serving it to tens of thousands of doorsteps a day, reliably and cheaply, was a separate one. This post is about that second problem.
We self-host the model rather than call an external API for data residency, for unit economics at volume, and because latency behaviour under load becomes something we control rather than something we discover.
At tens of thousands of validations a day, the deployment question stops being incidental. Three choices carried most of the weight, and they are worth stating as principles rather than as a parts list.
The model is quantised so that it fits comfortably on a single GPU. The instinctive reading of that is "the model got smaller," which misses the point entirely.
Model weights and the inference cache share the same GPU memory, and that cache is the working memory for every request in flight. Vision requests are unusually expensive there, because an image expands into far more context than a comparable amount of text.
Compression does not shrink the cache. It makes room for more of it, and cache headroom is what concurrency actually is. The right way to size a deployment is therefore backwards from how many simultaneous image-bearing requests you need to serve, not forwards from how large the model file is.
We began on a runtime that is excellent for local experimentation and wrong for production: it processes requests one at a time and allocates memory in fixed blocks with real waste.
The production engine batches requests in parallel and allocates memory in small blocks on demand, the same idea an operating system uses for virtual memory. Moving between them changed throughput by an order of magnitude on identical hardware and an identical model. The experimentation runtime also proved unstable under sustained load; the production one did not.
Its second contribution is continuous batching. As soon as one request finishes, its slot is filled immediately rather than waiting for the whole batch to drain. The GPU is never idle, which raises throughput and lowers average latency at the same time a rare optimisation that does not trade one against the other.
The last choice is the least glamorous, and the one most often skipped. The model runs as a containerised service under standard orchestration, with the same health checks, autoscaling, self-healing and dashboards as any other production system. Latency, throughput and cache pressure are monitored continuously, because a model that quietly degrades under load is indistinguishable from one that is working until someone looks.
The system has been running PAN India across both routes (vision embeddings and the vision-language model; see the companion post for how that split is decided) for months, at a scale of tens of thousands of validations per day, with daily monitoring of outcome rates and of the flag rate that surfaces potential mismatches while the rider is still at the door.
Two things are worth reporting from live operation.
Stability: The flag rate holds steady in a narrow band week over week, with no drift or degradation at full volume. A model that quietly gets noisier under production traffic is a model that will eventually be switched off, and this one has not.
Impact:
There is also a measurable second layer of value still on the table. In a share of cases, the system flags a likely mismatch and the item is collected regardless. Converting that behaviour through better in-flow prompting and confidence-tiered actions rather than a single binary flag represents a further meaningful increment on what has already been realised, without any change to the model itself.
None of the infrastructure work here is specific to reverse-pickup QC quantisation for cache headroom, a batching-aware serving engine, and treating the model as just another production service are the same three decisions we'd make for any self-hosted vision-language model at this volume. What made it work was tuning them against a real, measured trade-off, the same discipline that shaped the prompt itself: measure, don't assume, and let the operating constraint (latency a doorstep interaction can absorb) drive the engineering choice rather than the other way around.
New to the series? Start with Part 1: Solving Reverse Pickup Verification with Vision AI, which covers the four experiments behind the model.
