tokens8%

01 / What BrainGenz is

An AI engineering studio in Lahore, working with teams in the US, Europe and the Gulf. Retrieval over private documents, speech recognition at real concurrency, and the internal tools that sit around them. Fifteen systems shipped so far - four of them our own, and labelled as ours. We design it, build it, then keep running it, and we publish the measurements, including the ones that show where it breaks.

02 / The part nobody quotes for

Anyone can launch it. Month seven is the hard part.

Building the thing is the easy half. What costs money is the system that quietly drifts three months after handover - slower every week, dearer every month, and nobody left who remembers why it was built that way. That second half is where we do our work, because it decides whether the first half was worth paying for.

03 / What that buys you

A number you can check, instead of an adjective.

These are real measurements from a system we run, not a projection. It keeps getting faster right up to 300 people using it at once. But somewhere between 200 and 250, the slowest replies jump from 32 milliseconds to 94 - nearly triple. That jump, not the average, is what sets your server bill for the rest of the year, so we publish it.

Citrinet-1024 on NVIDIA Riva, 4 model instances, gRPC over 10 channels, 60-second runs, p95 latency

AI engineering studio · Lahore

Still running in month seven

A small engineering team in Lahore. We design the system, build it, then run it, and we are still the ones answering for it long after the invoice cleared.

Six written rules, reviewed monthlyVercel, Railway, Hetzner, SupabaseA runbook with every handover

Four numbers, each with the conditions it was measured under. If you cannot go and check one of them, we should not have put it on the page.

15
Systems shipped, every one of them listed below
Including the four internal ones, which are labelled as internal
553/s
Speech chunks per second at 300 concurrent streams
Citrinet-1024 on NVIDIA Riva, 4 instances, gRPC over 10 channels, 60-second runs
<5¢
Per answer, retrieved and scored on eight dimensions
A retrieval assistant we run, measured on infrastructure we self-host
4
Clouds carrying production traffic today
Vercel, Railway, Hetzner and Supabase, per the stack on each case study

Named where the work is already public. The rest stay described by sector until the client says otherwise, which is a slower way to build a logo wall and the only honest one.

RetainQuranCV SummitClean Cooking AllianceAdvanced Agent MarketingCreatorLoopplus a healthcare client, a construction tendering platform, and a Shopify merchant

Work

Six carrying traffic today. Nine more shipped and handed over.

Open any of them for the problem, how it was built, what it costs to run, and the part we would still not let a model do unsupervised.

01Live

Retain Quran

An AI-powered mobile app helping Hafiz-e-Quran evaluate and improve their recitation accuracy through advanced speech recognition and ML models, built to serve millions during Ramadan.

General speech recognition transcribes what it hears into words. That is not the task here. Recitation evaluation has to judge whether what was recited was correct - pronunciation, articulation and the rules of tajweed - which is closer to assessment than to transcription.

Islamic Education / Mobile App
TensorFlowSpeech RecognitionCustom NLP ModelsGCPKubernetesCloud RunVertex AIPython
02Live

AIAutoestimate

A collision estimating platform that reads repair photos, decodes the VIN, and builds a priced line-item estimate with grounded OEM part pricing an estimator can defend.

The obvious build is a vision model over damage photos returning a parts list. That version works, and it is not an estimate. Four independent reviews of one real shop estimate against its own photo set agreed that roughly 76% of the estimate's value came from parts and operations nobody could see: reinforcements behind a bumper, brackets, harnesses, the alignment a shop performs after a suspension hit.

Automotive / Collision Repair
Gemini 3 Flash (vision)Gemini 3.5 Flash (grounded search)Claude (native web search)GPT (Responses API)LangChainReact 19ViteMUI v7
03Live

Tomogi

A neighbourhood platform where people post local tasks, neighbours take them on, and completed work builds a reputation that is earned rather than claimed. Web, plus Flutter applications for Android and iOS.

Tomogi describes itself as a gig layer for neighbours: post a task, have someone nearby take it on, and build a local reputation from work actually completed. That last clause is the one that constrains everything. A reputation that can be inflated is worth nothing, so the system has to be able to say who did what, and be right about it.

Social Networking / Local Services
Next.js 14React 18TypeScriptTailwind CSSTanStack QueryMotionFlutterDart
04Shipped

User Insight Hub

A research assistant for teams working from interview transcripts. It answers questions with inline citations and generates 59 kinds of structured research artifact across multiple datasets at once.

A research team finishes forty interviews and then loses weeks turning them into something a product decision can rest on. Any synthesis written by hand is also impossible to trace: six months later nobody can tell whether a claim came from a participant or from the person writing the slide.

AI & Research Tools
LangGraphIntent routingTemplate registryMilvus / Zilliztext-embedding-3-largeQuote extractionCitation validationGPT-4o
05Live

PSX Copilot

A chat assistant for Pakistan Stock Exchange investors. Ask in plain words and it answers with live market prices, analysis drawn from official filings with citations, paper trading and strategy backtesting.

In a finance product, one wrong number costs more trust than a hundred good answers earn back. So the model is never the source of a figure. Every reply about money is composed from a fact pack that code assembled first: execution results, cash balances, costs, timestamps.

FinTech / Capital Markets
LangGraphLLM routingEmbedding relevance gateTyped actionsOCR pipelineHybrid keyword + semantic searchPer-citation verificationLive PSX feed
06Shipped

n8n Workflow Systems

Self-hosted n8n running production operations across sales, marketing and back-office work. The flagship pipeline turns call recordings into graded summaries with objections and action points, delivered to Slack and the CRM; the same platform runs the scheduled, integration-heavy work that quietly eats a team's week.

Reviewing sales calls to coach from them requires someone to listen to them. At any real call volume that quietly stops, so coaching becomes sporadic and the calls that most needed review are the ones nobody got to.

Sales & Marketing Operations
n8n (self-hosted)JavaScript code nodesValidated webhooksWhisper-class transcriptionOpenAI modelsStructured output with retrySlackGoHighLevel
07Live

CVSummit

An AI-powered resume optimization platform that analyzes, enhances, and perfects CVs for top-tier job applications in seconds.

Resume optimisation sounds like a language-model problem and mostly is not. The model is perfectly capable of improving a bullet point once it can see one. Getting to that point is the work: a resume arrives as a PDF or a DOCX laid out in two columns, or a table, or a template that renders beautifully and stores its text in an order no human would read it in.

HR Tech / Career Services
OpenAI GPT-4LangChainNLP ModelsReactNext.jsTypeScriptTailwind CSSNode.js
08Live

CreatorLoop

An AI-powered platform that generates stunning, photorealistic car images from user prompts, enabling automotive enthusiasts and businesses to visualize custom vehicle designs.

The interesting part of this project was not prompting. Stable Diffusion produces good car imagery with reasonable inputs. The difficulty is that every image costs real GPU seconds, GPU capacity is expensive and slow to acquire, and users expect an interactive experience from a workload that is anything but.

AI / Automotive / Creative Tools
ComfyUIStable DiffusionCustom LoRA ModelsReactNext.jsTypeScriptTailwind CSSPython
09Shipped

Emerge Insights

A multi-tenant platform that classifies a builder's tender drawing set by trade and extracts quantities into a tender-ready Excel workbook, labelling every number by how it was obtained.

Taking off quantities from a tender drawing set is expensive specialist work, and automating it with a vision model alone would be worse than useless. The output goes into a tender someone is accountable for, so a value nobody can verify has no value.

Construction / Quantity Surveying
Next.js 16React 19TypeScriptFastAPIPython 3.13UvicornClaude visionPDF text-layer parsing
01 / 09

Also shipped

Nine more that shipped and were handed over, with runbooks. Point at one for the detail.

Stack

What we actually build with.

Every one of these is running on at least one system listed above, not on a list of things we would be willing to learn.

How we work

Six rules, not six values.

Anyone can promise discipline. Ours is checkable by ten in the morning. We review these monthly, on the principle that a rule everyone ignores is either wrong or needs tooling to enforce it - so you fix one or the other, rather than re-announcing it.

  • 01
    No card, no code.Everything you ask for is written down before anyone starts - including the thing mentioned in passing near the end of a call. Nothing you are paying for lives only in somebody's memory of a conversation.
  • 02
    The board is right by 9:55.Standup is ten minutes of reading the board together, not ten minutes of recalling things out loud. When you ask where something has got to, the answer is on a screen rather than in someone's head.
  • 03
    Blocked means flagged inside the hour.With one comment naming exactly who or what unblocks it. A silent blocker is the only kind we treat as a mistake, because that is how a week quietly disappears into something nobody mentioned.
  • 04
    One change, one subject.Four unrelated fixes shipped together cannot be reviewed properly and cannot be undone cleanly. Kept apart, the day something breaks we reverse the exact change that broke it, in minutes rather than in an afternoon of guessing.
  • 05
    Logins, payments and deletion get a named reviewer.Assigned the moment the work opens and answered within 24 hours. Those three are where a mistake costs you customers rather than an afternoon, so none of them ever ships on one person's judgement alone.
  • 06
    Nothing lives only on a laptop.Pushed before anyone logs off, every day. No part of the thing you paid for can go home in someone's bag - or leave with them.

Capabilities

Six things we have shipped more than once.

That is the whole list. If your problem is not on it we will say so, because the alternative is learning on your budget and calling it a discovery phase.

Retrieval over private documents

The answer and the source procedure side by side, so the agent can check it before repeating it to a patient. Every response scored on eight fixed dimensions, and the low scores come back as proposed edits to the knowledge base. Running under five cents a query in production.

Speech recognition at real concurrency

Fine-tuned Citrinet-1024 on Uthmani Arabic, served through NVIDIA Riva, multi-region on GKE. We benchmarked A100 against T4 and L4, one model instance against four, and measured from Dammam, Singapore and Toronto rather than from a load generator sitting next to the cluster.

Vision pipelines that produce numbers

VIN plates and odometers read from photographs, decoded against NHTSA vPIC with two commercial decoders behind it, because one decoder is a single point of failure for an entire estimate. Damage arrives as bounding boxes carrying component, operation, material and severity, which is enough structure to calculate labour from.

Agents wired into systems that bill people

Sales, migration and operations agents connected to CRMs, storefronts and dialers. Every write path gated and schema-checked, so an invalid generation is re-prompted instead of landing in a customer record for somebody to unpick three weeks later.

Workflow automation with a failure plan

n8n and Zapier systems built so the worst case is a queued item somebody can see, rather than a silent drop nobody notices for a fortnight. Error branches, retries, idempotency keys, and a runbook handed over with the credentials.

The unglamorous half

Numbered migrations that run identically in staging and production, nobody editing a live database by hand, billing reconciled from Stripe webhooks rather than from whatever the browser claimed, and cost per request tracked as a number somebody owns.

Method

Five habits that decide whether it still works in month seven.

Shipping quickly is easy to promise and impossible to check. What costs money is a system that drifts three months after handover. We review these monthly, on the principle that a rule everyone ignores is either wrong or needs tooling to enforce it, so you fix one or the other and never just re-announce it.

01

Scoping names what we will not automate

Before anything is built we write down which parts of your workflow should stay with a person, and where accuracy degrades enough to matter. That belongs in a scoping conversation, not in a discovery you make in month four with the invoice already paid.

02

One adapter per outside dependency

Every model vendor, VIN decoder and payment provider sits behind a single file. When a provider has a bad hour, adding a fallback is a one-file change. That is the difference between an outage costing an afternoon and an outage costing a rewrite.

03

Answer quality is a number, not an impression

Outputs are scored on fixed dimensions and the failures feed back into the corpus or the prompt set. A system with no score cannot tell you it is degrading, so it degrades quietly until a customer does the telling for it.

04

Numbered migrations, no hand edits

Schema changes move through the same numbered path in staging and in production, and nobody edits a production database by hand. Before we rewrote one repository's history we bundled all 46 refs to disk first, because "we can always recover it" is only true if somebody actually did.

05

Handover assumes we leave

Runbooks, exportable data, documented schema, credentials you hold. We would rather you could take the whole thing to another team and decide not to. A studio that makes itself hard to replace has optimised for the wrong thing.

Evidence

Ask us for these six things. Ask our competitors too.

This list comes from what enterprise technology buyers say builds their confidence, and from the documented red flags that mark a vendor who has never run anything in production. We would rather hand you the checklist than be graded on a pitch deck.

01

Named references, including the old ones

Not only this quarter's happy client. People still running systems we built a year or more ago, who have been through a model change and a vendor outage with us. References drawn only from first-year clients skip the part where you find out what maintenance actually feels like.

02

The first ninety days, incidents included

What broke after go-live, how long it took us to notice, and what we changed so it couldn't happen twice. Every production system has a bad first week. A vendor who can't describe theirs either hasn't shipped or isn't telling you.

03

Benchmarks with the method attached

Throughput and latency with the hardware, protocol, concurrency and duration they were measured under, so you can reproduce the number or argue with it. A latency figure without its method is marketing.

04

Production outcomes, not pilot results

A case study that ends at "the pilot worked" is hiding the hard part. Ours say what the system does in production, what it costs per request to run, and which parts still need a human.

05

Data portability in writing

Your data, your schema, your migration history, exportable on request, documented well enough that another team could take over. We'd rather you could leave and choose not to.

06

The limitations, before you ask

What the system won't do, where accuracy degrades, and which parts of your workflow we'd tell you not to automate at all. That conversation happens during scoping, not after the invoice.

Benchmark

A latency number without its method is marketing.

So here is one of ours in full, with the hardware, the protocol, the concurrency and the run length attached. Read across, and watch the p95 column.

RetainQuran AI · Citrinet-1024 on NVIDIA Riva · 4 model instances · gRPC, 10 channels · 60-second runs

Throughput, share of peak22%5059%15088%250100%300Concurrent streams
Peak throughput
553 chunks/s
at 300 streams
p95 breaks at
250 streams
31.8 ms to 94.4 ms
Comfortable ceiling
200 streams
p95 31.8 ms
Throughput and latency against concurrency
Concurrent streamsThroughputAvg latencyp95 latency
50119 chunks/s14.7 ms19.0 ms
100228 chunks/s17.2 ms21.9 ms
150328 chunks/s21.2 ms25.0 ms
200414 chunks/s27.1 ms31.8 ms
250486 chunks/s36.1 ms94.4 ms
300553 chunks/s44.1 ms158.7 ms

Read across, and note where p95 turns: throughput keeps climbing past 250 streams but tail latency stops being comfortable. That's the number that decides how many replicas you run, which is why we publish it rather than quoting the average alone.

Contact

Tell us what breaks, and who it breaks for.

Two paragraphs is plenty. If it is not work we should take on, the reply will say so and point you at someone better suited, which costs us a lead and saves you a quarter.

sales@braingenz.com

This form validates and confirms locally. The POST target is deliberately unwired until you choose the inbox, because a form that quietly discards enquiries is worse than no form at all. Email works today.

No tracking on this form.