

THE TAKEAWAYThe cheapest option on the comparison sheet was the one we rejected — it only spoke English, and the client's customers did not. The second-cheapest still cut the bill by 80%.
A five-agent lead-qualification calling system — and the hosting decision that took voice from ₹1.7 lakh a month of vendor billing to ₹33,680 of self-hosted compute.
Mockingbird is a lead-qualification calling agent with a CRM attached. A lead comes in from an ad; within minutes it rings them, qualifies them, follows up on WhatsApp if they do not pick up, and hands the sales team a warm buyer with the whole conversation attached. Leads stop going cold in an inbox. I built the architecture, the front end and the back end. What follows is the part that actually decided whether the product could exist.
The system: five agents, scoped on purpose. A core runtime answering as an OpenAI-compatible chat brain, a FastAPI backend routing between it and a fleet of five agents over a Docker network — the caller, WhatsApp, memory, follow-up, and an analyst. A gateway orchestrator exposed eighteen CRM tools to them, scoped per agent rather than shared: memory and analyst got read-only access and could not write to a lead record no matter what they were asked to do. That scoping is not ceremony. An agent that can both summarise a conversation and edit the record it summarised will eventually edit the record to match its summary.
Then the hosting question, which was the real one. The agent was built first on a managed voice API, because that is the fastest way to a working call. Testing put the average AI call at five to six minutes — longer than a human call, because the agent is patient and does not cut people off. The client needed two hundred to two hundred and fifty calls a day, on working days only. So I wrote the comparison properly before anyone committed to anything, and costed three routes against the same requirement.
Route one — the managed voice platform already in use. Its top tier was ₹87,120 a month for 12,375 included minutes. That works out at ₹7.04 a minute, and the plan is billed for the whole month whether the campaign runs that long or not. 200 calls × 22 working days × 5.5 minutes = 24,200 minutes a month. The top plan does not cover it — it covers roughly half, and the rest bills as overage at a rate the vendor would not confirm. Taken at the plan's own effective rate: ₹1,70,368 a month, ₹38.72 a call, for voice alone. Concurrency was also capped by tier, which matters when calls go out in parallel batches rather than one at a time.
Route two — a pay-per-minute voice agent platform. No subscription, billed on usage, with telephony and number provisioning included, which the self-hosted route does not have. But the headline per-minute rate only covers the voice engine; add the language model and the telephony leg and a realistic production rate lands around ₹12.35 a minute. At 24,200 minutes that is roughly ₹2,98,870 a month — more expensive than the platform it was meant to undercut, at exactly our volume.
Route three — self-hosted, and the one I recommended. No subscription, no per-minute fee, no concurrency ceiling, no call-duration cap. Pay for the server by the hour and nothing else. The honest caveat went into the document in the same breath: a compute figure is compute only. It carries no phone number and no telephony, which both managed routes bundle in, so it was never an apples-to-apples number and I did not present it as one. It also needed two weeks of build and training before it could place a call at all.
And then the cheapest option on the sheet got rejected — by me, on the one ground that mattered. The self-hosted voice stack we had costed spoke English. The client's customers speak Marathi, Hindi and English, often inside the same sentence, and bending an English-only model into a trilingual one is not a configuration change — it is a different project with a worse result. So the architecture stayed self-hosted and the model was swapped: an open-source multilingual voice model instead, fine-tuned, with its API wired into the calling agent as the audio layer. A model that is cheaper and speaks the wrong language is not cheaper. It just fails for free.
What actually shipped, across two machines. An NVIDIA A100 at roughly ₹144 an hour carried the voice model and the language model, serving several concurrent calls off one card. The LLM was Qwen 3.5 at 7 billion parameters, pre-trained by us on real sales material rather than prompted into behaving — small enough to self-host, specific enough not to need to be large. Everything else sat on a separate VM at around ₹2,000 a month: the five agents, the orchestrator, the CRM and the front end. None of that needs a GPU, and pinning it to the same box would have meant paying GPU rates to run a web app.
The last saving was a schedule. The GPU billed by the hour. Left running it is a machine idle two-thirds of its life, because nobody answers a sales call at four in the morning. So the dialling window became infrastructure: I gave the Mockingbird worker scoped CLI access to the GPU provider and had it bring the server up at 10:00 and take it down at 20:00 — ten hours, the only ten hours the agent is allowed to dial anyway, on the twenty-two days a month the desk actually works.
The numbers, end to end: 10 hours × 22 working days = 220 GPU hours a month, at ₹144 an hour — ₹31,680. With the application VM on top, about ₹33,680 a month to run, or roughly ₹7.65 a call. Against ₹1,70,368 on the managed plan, that is ₹1,36,688 saved a month — close to ₹16.4 lakh a year, an 80% cut, and about five times cheaper per call. The schedule alone accounts for ₹72,000 of it. Running that same GPU around the clock would have cost ₹1,03,680 a month instead of ₹31,680 — the difference between renting a machine and renting the hours you actually use it.
The behaviour came from studying one practitioner rather than writing prompts — Saad Khaja, who posts as @saadsells. Not his slogans; his mechanics. How an objection gets acknowledged before it gets answered, what a second follow-up says that a first one cannot, when a call should be ended rather than saved. His published material was transcribed and used as the behavioural reference the model was trained against. Credit where it belongs: the agent sounds like it knows what it is doing because somebody who does was studied closely. The method is his.
The demo linked below is the real front end, running on mock data. The backend is deliberately not part of it — the client's system is theirs, and a portfolio piece has no business carrying it. Everything on screen is the interface as built: the lead desk, the conversation threads, the call log, the qualification scoring. Sign in with any email. The prefix picks the role — admin@, sales@ or viewer@.
The promo film — how a lead becomes a booked site visit
Open the live demo — real front end, mock data
mockingbird.wishmaster.space
NEXT PROJECT ↘Eurobond BSE Listing Cinematic Coverage