Shadowfax Logo
back-arrow
Shadowfax Logo
Services
image
Company
image
Resources
image
Partners
image
backArrow

Serving an Open-Source Vision-Language Model in Production at National Scale

Published by Shadowfax
Reverse Pickup
Serving an Open-Source Vision-Language Model in Production at National Scale
Shadowfax
facebookbluexbluelinkedinblue
Posted on:October 05, 2026

This is Part 2 of our Reverse Pickup QC series. Read Part 1: Solving Reverse Pickup Verification with Vision AI.

In Part 1, we walked through how we arrived at a vision-language model for reverse-pickup QC we walked through how we arrived at a vision-language model for reverse-pickup QC in Indian e-commerce logistics reasoning through a doorstep photo comparison rather than scoring image similarity. Getting the model right was one problem. Serving it to tens of thousands of doorsteps a day, reliably and cheaply, was a separate one. This post is about that second problem.

Why self-host at all

We self-host the model rather than call an external API for data residency, for unit economics at volume, and because latency behaviour under load becomes something we control rather than something we discover.

At tens of thousands of validations a day, the deployment question stops being incidental. Three choices carried most of the weight, and they are worth stating as principles rather than as a parts list.

Compression is a concurrency decision, not a storage one

The model is quantised so that it fits comfortably on a single GPU. The instinctive reading of that is "the model got smaller," which misses the point entirely.

Model weights and the inference cache share the same GPU memory, and that cache is the working memory for every request in flight. Vision requests are unusually expensive there, because an image expands into far more context than a comparable amount of text.

Compression does not shrink the cache. It makes room for more of it, and cache headroom is what concurrency actually is. The right way to size a deployment is therefore backwards from how many simultaneous image-bearing requests you need to serve, not forwards from how large the model file is.

The serving engine matters as much as the model

We began on a runtime that is excellent for local experimentation and wrong for production: it processes requests one at a time and allocates memory in fixed blocks with real waste.

The production engine batches requests in parallel and allocates memory in small blocks on demand, the same idea an operating system uses for virtual memory. Moving between them changed throughput by an order of magnitude on identical hardware and an identical model. The experimentation runtime also proved unstable under sustained load; the production one did not.

Its second contribution is continuous batching. As soon as one request finishes, its slot is filled immediately rather than waiting for the whole batch to drain. The GPU is never idle, which raises throughput and lowers average latency at the same time a rare optimisation that does not trade one against the other.

Treat the model as ordinary infrastructure

The last choice is the least glamorous, and the one most often skipped. The model runs as a containerised service under standard orchestration, with the same health checks, autoscaling, self-healing and dashboards as any other production system. Latency, throughput and cache pressure are monitored continuously, because a model that quietly degrades under load is indistinguishable from one that is working until someone looks.

What it changed

The system has been running PAN India across both routes (vision embeddings and the vision-language model; see the companion post for how that split is decided) for months, at a scale of tens of thousands of validations per day, with daily monitoring of outcome rates and of the flag rate that surfaces potential mismatches while the rider is still at the door.

Two things are worth reporting from live operation.

Stability: The flag rate holds steady in a narrow band week over week, with no drift or degradation at full volume. A model that quietly gets noisier under production traffic is a model that will eventually be switched off, and this one has not.

Impact:

  • Material reduction in loss: Loss accepted per order on reverse pickups fell substantially in the first week after go-live, measured against the week before at constant order volume.
  • Larger still in direct scope: Restricted to the design, pattern and colour mismatches the system is actually built to catch, the reduction is substantially larger.
  • Escalation rate flat: Unchanged post go-live, confirming the reduction came from catching mismatches at the door rather than from displacing the problem downstream.

There is also a measurable second layer of value still on the table. In a share of cases, the system flags a likely mismatch and the item is collected regardless. Converting that behaviour through better in-flow prompting and confidence-tiered actions rather than a single binary flag represents a further meaningful increment on what has already been realised, without any change to the model itself.

Closing thought

None of the infrastructure work here is specific to reverse-pickup QC quantisation for cache headroom, a batching-aware serving engine, and treating the model as just another production service are the same three decisions we'd make for any self-hosted vision-language model at this volume. What made it work was tuning them against a real, measured trade-off, the same discipline that shaped the prompt itself: measure, don't assume, and let the operating constraint (latency a doorstep interaction can absorb) drive the engineering choice rather than the other way around.

New to the series? Start with Part 1: Solving Reverse Pickup Verification with Vision AI, which covers the four experiments behind the model.

Related Blogs

image
img

Get in Touch

Subscribe now for the latest updates delivered straight to your inbox!
Delivery Partner App
img
Shadowfax Courier App
imgimage
© 2026 All rights reserved. Shadowfax Technologies Limited