ReliableVoice AGI.
Voice AI that demos well is easy. Voice AI you can trust with a million real conversations is a research problem. Bryn Labs exists to solve it, with evals that predict production, observability that misses nothing, and agents that learn from every failure.
The last mile of voice AI isn't a feature. It's a science.
Foundation models made voice agents easy to start and hard to trust. A stack of probabilistic systems fails in ways no unit test anticipates: fragmented turns, phantom interruptions, policies that bend under pressure, silence that reads as thought. The gap between a great demo and a dependable agent is where deployments die. That gap is our entire research agenda.
An agent that is 95% right sounds finished and is not. The last 5% is rare, strange, and expensive, and it only shows up at production scale.
STT, LLM, TTS, and telephony are four probabilistic systems in a trench coat. Every layer can be individually fine and jointly wrong.
Callers do not average their experiences. One bad call undoes a hundred good ones, so reliability has to be engineered call by call.
Six threads. One goal.
Reliability is not one problem. It is a stack of them, from the acoustic frame to the final verdict. We work every layer.
Measurement is the product. Scenario suites, adversarial personas, and verdict systems that predict production behavior, not leaderboard behavior.
Every call, fully measured: latency, interruptions, policy adherence, outcomes. Computed on all traffic, not a sample, so the tail has nowhere to hide.
Failure as fuel. Loops that turn a flagged call into a diagnosis, a fix, and a passing re-eval, shrinking the class of mistakes an agent can make twice.
Continuity across conversations. What an agent should carry between turns, calls, and callers, and what it must provably forget.
Reality, rehearsed. Synthetic callers with human impatience: accents, interruptions, dead air, network glitches. Production chaos on demand.
Hearing before reasoning. Turn-taking, diarization, voice activity, pronunciation: the acoustic substrate every downstream judgment depends on.
A decade in speech research.
The acoustic-modeling lineage behind the lab, from today's work back through noise-robust recognition at King's College London to subspace Gaussian mixtures at IIT Madras. Year by year, paper by paper.
On the bench right now.
Recovering who-said-what on stereo calls by anchoring diarization to per-channel voice activity instead of trusting STT speaker labels.
Rebuilding true conversational turns when transcription splits them mid-thought, so downstream metrics judge the conversation that actually happened.
Generating simulated callers that probe an agent where it is weakest: impatient, ambiguous, and off-script, at a scale no QA team can match.
Closing the loop from a failed call to a prompt fix to a passing re-eval, with humans approving direction instead of hand-writing every change.
Compressing an organization’s full policy surface into the small, mode-aware rulesets an agent can actually follow mid-conversation.
Catching mispronounced names, numbers, and domain terms in agent speech with phoneme-level acoustic models.
Naming the ways voice agents break: a shared vocabulary of failure modes, detection strategies, and fixes, grown from production evidence.
Superbryn is our lab bench. And our proof.
Most labs publish papers. We ship into Superbryn, the voice-AI reliability platform teams use to test agents before launch and watch them after. Every pillar above runs there against real traffic, which means our research is graded by production, not by benchmarks.
Scenario suites and personas run as real phone calls against your agent: before launch, on demand.
Every live call analyzed across dimensions the moment it ends. No sampling, no blind spots.
Latency, interruptions, sentiment, compliance. Decomposed per turn, comparable across agents.
Feedback on verdicts tunes the metrics themselves, so measurement improves with use.
Lab principles.
No claim ships without an eval behind it. If we cannot measure an improvement, we do not call it one.
Benchmarks are rehearsals. The score that counts comes from live traffic, on calls nobody scripted.
Every bad call is collected, clustered, and named. A failure mode we can name is a failure mode we can retire.
From VAD frames to verdicts, the same people own the whole pipeline. Nothing gets lost between teams.
A small lab with production stakes.
Bryn Labs is kept deliberately small: research owned end to end by the people who did the original work, graded by live traffic rather than citations.
“Voice agents fail in production because they weren’t built to learn from real-world conditions. We’re changing that.”
14+ years of research in speech recognition, noise-robust acoustic modeling, and assistive voice technology. PhD from IIT Madras (2011–2018), followed by postdoctoral research at King’s College London and enterprise NLP work at ZS Associates and Uniphore. Published extensively in IEEE and INTERSPEECH on speaker normalization, low-resource language modeling, and speech recognition for impaired speakers.
+ Behind the founders: the Superbryn engineering team. Every line of lab research is hardened, deployed, and scaled with the builders of the platform.
Pre-seed led by Kalaari through its CXXO initiative. Read “Why we invested in Superbryn”.
Selected for the Google for Startups Accelerator: India, 2026 class.
190+ citations across IEEE and INTERSPEECH in speech recognition and acoustic modeling.
What is Bryn Labs?
Bryn Labs is the research team behind Superbryn. We study why voice agents fail in production and build the science that makes them reliable enough to trust with real conversations: evals, observability, self-learning, and memory.
What does “Reliable Voice AGI” actually mean?
A voice system you can hand a real task without supervising every call. Not a model that occasionally dazzles, but a system that holds goals across a conversation, knows when it is failing, and gets measurably better every week it runs. We treat reliability as the capability, not a property you add later.
How does the lab relate to Superbryn, the product?
Superbryn is where our research ships. Every pillar, from simulation to dimension analysis to self-improving metrics, runs inside the platform against real customer traffic. The product is both our lab bench and our proof: research that does not survive production does not survive here.
Do you publish your research?
We publish working artifacts first: our open blueprint library, failure taxonomies, and technical notes as they mature. We favor things teams can use over papers that gather citations.
Can my team work with the lab?
We partner with a small number of teams running voice agents in production as design partners for new evals and observability research. If that is you, write to research@superbryn.com.
Are you hiring?
The lab is kept deliberately small. If you have done serious work in speech, evaluation, or agent reliability and want your research graded by production traffic, we want to hear from you.
Building a voice agent that can't afford bad calls?
We take on a handful of design partners for new evals and observability research. If your agent handles real customers, we want to study its failure modes and retire them.