How 11x ships trusted voice and chat AI agents with Bluejay
With Sachi Angle, Product Lead, and Muizz Matemilola, Member of Technical Staff
1,100+
hours of manual QA time saved through thousands of realistic simulations and accurate evals
Daily
simulation runs validating every agent change before it reaches a customer’s phone line
Yes, faster
customers adopt with confidence built through red teaming and automated reporting
About 11x
The AI growth company. Maker of Julian, the conversational agent that handles inbound demand for go-to-market teams. 11x supports customers like Xerox, Checkr, Canibuild and more. They are backed by a16z and Benchmark and have raised over $70M. 11x.ai
11x builds AI agents for go-to-market teams. Its flagship conversational agent, Julian, handles inbound demand for customers across voice, chat, phone, email, and the web, conducting entire conversations on a customer’s behalf to reach an end goal, whether that’s qualifying a lead, booking a meeting, or closing a sale.
That’s a high-trust job. When a customer hands Julian their inbound pipeline, every conversation happens in their name, with their prospects, in the wild.
“I came into 11x obsessing over how we maintain that trust for our customers when we’re deploying agents for them. AI agents are non-deterministic by design.”
Bluejay is how 11x turns that non-determinism into confidence: simulating thousands of realistic conversations, environments, and edge cases before an agent ever takes a live call.
The problem: you can’t manually test what you’ve never seen
Before Bluejay, 11x relied on a combination of internally built testing infrastructure and a lot of manual testing. That worked for known cases and the edge cases the team could think of. It broke down everywhere else.
“One thing that is very challenging is testing for situations that you’ve never encountered before,” said Muizz Matemilola, a member of technical staff on the Julian team.
Manual testing had a second cost: iteration speed. Validating a single change meant calling the agent over and over, one conversation at a time.
“Bluejay tightened our iteration loop on agent behavior. Instead of manually calling the agent over and over to test one change, we launch a batch of simulations and validate across all of them at once.”
Before Bluejay
×Internal testing stack plus heavy manual testing
×Edge-case coverage limited to scenarios the team could imagine
×One change validated one phone call at a time
×One-off production issues were hard to reproduce
×Reactive: fixing problems after they surfaced
After Bluejay
✓Simulations across environments, industries, and audio conditions
✓Stress tests run until agents meet the bar, before deployment
✓Batched simulations validate every change at once
✓Production feedback loops directly into new test scenarios
✓Proactive: problems pinpointed at scale before customers hit them
Simulating the real world, before it happens
Bluejay lets 11x recreate the messy conditions of real conversations quickly and at scale, across every surface Julian operates on. The same simulation infrastructure that stress-tests voice agents on the phone also tests 11x’s chat agents, so every channel a customer’s prospect might use gets the same rigor. “It’s a great hack in quickly recreating a lot of simulations really fast,” Angle said. “Bluejay’s been great in having us pinpoint the problems that are actually existing at scale, versus just reacting to problems that may come up one-off.”
One example: noisy environments. How does a voice agent perform when there’s conflicting audio coming in from the background? A busy dealership floor, a car radio, a second conversation nearby.
The team had fixes they believed in, but no good way to test them. Recreating a one-off production condition by hand doesn’t scale.
“We approached Bluejay saying, hey, we’ve got some things that we think work, we want to test them out,” Angle said. “Very instantly, they responded: guess what, we’ve got something for you. You can simulate not just the conversation, you can simulate all of these background sound effects.”
The team ran their improvements through simulated noisy calls and got fast, concrete confirmation that the fixes and enhancements were working.
“
We came to them with one problem; solved it. Then discovered more; solved those. And we’re confident it’ll continue to be that way.
Muizz Matemilola
Member of Technical Staff, 11x
Red teaming: proving agents are safe, not just capable
Performance is only half the trust equation. An agent that handles conversations brilliantly but can be manipulated into saying something harmful is still a liability. So alongside behavioral simulations, 11x uses Bluejay’s automated red teaming suite to verify that Julian agents are content safe and can stand up against adversarial pressure before they ever reach a customer’s phone line.
Bluejay’s red teaming probes agents against the OWASP LLM Top 10, the industry-standard taxonomy for LLM security risks. But the OWASP LLM Top 10 has no category for harmful content, so Bluejay grades content-harm findings on their own dedicated scorecard, keyed by harm category and aligned to the MLCommons AILuminate hazard taxonomy. On the secondary framework map, those findings land on MITRE ATLAS AML.T0048, External Harms.
OWASP LLM Top 10MLCommons AILuminateMITRE ATLAS
The result is coverage that spans both dimensions of agent safety: security vulnerabilities like prompt injection and data leakage on one side, and content harms like toxic or dangerous outputs on the other, each graded against the framework built to measure it.
For 11x, that rigor pays off directly in sales conversations. Enterprise customers evaluating Julian want evidence, not assurances, that an agent is ready to represent their brand. Bluejay’s reports give 11x exactly that: a documented record of the scenarios, attacks, and harm categories the agent has been tested against, mapped to frameworks their security and compliance teams already recognize.
Evals that evolve with the customer
Underneath all of this testing sits a philosophy about what an eval suite should be. At most companies, evals accumulate: every incident adds a test, nothing ever gets removed, and over time the suite measures what mattered six months ago rather than what matters now.
11x treats its eval suites as living systems instead. When a customer’s priorities shift, when Julian takes on a new workflow, or when a scenario stops reflecting how real callers behave, the corresponding simulations get rewritten, re-weighted, or retired. Because Bluejay can generate and run new scenario batches on demand, updating a suite costs minutes, not sprints, so keeping evals current is a habit rather than a project.
“
Most teams treat evals like a checklist that only grows. But customers evolve, priorities shift, and static evals quietly stop measuring what’s most important. Our agents must stay buoyant to that. Bluejay is what makes our living, dynamic eval suites possible.
Francisco Izaguirre
Engineering Manager, 11x
A self-healing loop: agents improving agents
Today, Bluejay is integrated directly into the Julian development cycle. When an agent surfaces feedback that needs addressing, an internal agent makes a change to the prompt. That change is immediately kicked off to Bluejay, tested across the full range of scenarios, and returned with results and suggestions. The agent picks it back up and runs the cycle again, over and over, until the team is at a good spot.
The self-healing loop
Feedback surfacesa call or customer conversation flags something
→
Agent updates the promptan internal 11x agent makes the change
→
Bluejay simulatesthe change runs across the full range of scenarios
→
Results & suggestionsthe agent picks them up and iterates
The cycle repeats until quality clears the bar
The team runs simulations at least daily. Often the trigger is a customer call. “Pretty frequently after we jump on a customer call, we get new ideas about something we want to test,” Angle said. “We’ll go run some simulations and present it to the customer.”
That rhythm answers the two questions that matter most in agent development: has my change worked, and what scenarios do I need to be building for next?
It also explains why the team treats Bluejay as durable infrastructure rather than a point solution.
“
We like using Bluejay as a headless eval toolkit. Models evolve, voices evolve, with Bluejay a good conversation becomes timeless.
Francisco Izaguirre
Engineering Manager, 11x
The impact: customers deploy with confidence
For 11x, the impact of Bluejay is twofold.
Customers say yes faster. “Customers allow us to deploy because they have a lot of confidence,” Angle said. “They appreciate they’re taking a risk by putting their agents out in the wild, and they see we’ve tested across all these different scenarios that they probably haven’t even thought of.” For enterprise buyers, Bluejay’s reports turn that confidence into documentation: proof of readiness that 11x can put in front of security and procurement teams before deployment.
Course correction is fast. When a genuinely new scenario does show up in production, one nobody anticipated, the team can reproduce it in simulation, fix it, and verify the fix quickly, rather than waiting for it to recur on live calls.
And the math is simple. Validating every change across a batch of simulations instead of one phone call at a time adds up: so far, it has saved the team over 1,100 hours of manual testing.
And the speed of shipping never had to slow down to get there. “It’s been great to have Bluejay as a partner, letting us go as fast and keeping that momentum on delivering agents, while still promising that the agents are going to be doing what we’ve designed them to do,” Angle said.
That word, partner, is deliberate. As Muizz put it:
“
Bluejay is not a tool we use. They are partners.
Muizz Matemilola
Member of Technical Staff, 11x
What’s next
The next chapter is expanding the self-healing loop, and going deeper into observability to power it. Angle is most excited about locking in the loop end to end: a call goes out, feedback comes back, an internal agent makes changes, Bluejay tests them, and the cycle continues without a human bottleneck.
Making that loop truly autonomous means richer visibility into what’s happening inside production conversations. The more precisely 11x can observe where an agent hesitated, misunderstood, or lost a caller, the better the signal that feeds back into simulation, and the faster each cycle converges on a fix. Deeper observability turns the loop from reactive to genuinely self-improving: production doesn’t just surface problems, it teaches the system what to test next.
The roadmap also follows trust. As voice and chat agents expand in capability and the tools they can use, 11x’s customers are trusting Julian with more of their workflows. Every new workflow brings new edge cases, and a new set of simulations to run.
Key takeaways
Simulate what you can’t anticipate. Manual testing covers the scenarios you can imagine. Simulation across environments, industries, and audio conditions covers the ones you can’t.
Batch your validation. Testing one change with one call at a time throttles iteration. Launching a batch of simulations validates a change across every scenario at once.
Close the loop with agents. Prompt changes flow automatically into Bluejay, results and suggestions flow back, and the cycle repeats until quality clears the bar.
Red team before your customers’ security teams do. Automated red teaming, graded against OWASP, MLCommons AILuminate, and MITRE ATLAS, produces the evidence enterprise buyers ask for anyway.
Test, monitor, and improve your way into customer trust. Customers deploy faster when they can see their agent has already survived the scenarios they hadn’t even thought of.
Thank you to Sachi, Muizz, Francisco, Aaron, and Santi for sharing 11x’s story.
Test, monitor, & improve your way into customer trust
Ship conversational AI that earns trust at every turn.