Cartesia needed a better first screen for contractor evaluator candidates. In four weeks, Jonathan Ramos reviewed 43 Ribbon interviews and hired 20 people without booking 43 live calls.

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.
Block quote
Ordered list
Unordered list
Bold text
Emphasis
Superscript
Subscript
"The biggest value for us is scale. I would not have been able to conduct 43 mini interviews in a live setting over the past month."
Jonathan Ramos, Cartesia
Cartesia was hiring contractors for evaluator work. The applicants represented about 20 language backgrounds, and English fluency mattered. Jonathan wanted to hear how people described their past work, not guess from a cover letter.
Scheduling every first screen himself was impossible. Skipping the conversation was not much better.
Ribbon moved that first conversation out of Jonathan's calendar. Candidates could complete it on their own time, then he could review the interview when he had room to pay attention.
"The conversational format gives candidates the opportunity to speak about their experiences in their own words, which gives a lot more signal than their initial cover letter," he said.
By August 21, 43 candidates had completed a Ribbon interview. Jonathan personally reviewed every one. He hired 20 people through the process.
He watched recordings at 1.5x speed and skipped to the sections where the candidate was speaking. Before pressing play, he read Ribbon's summary for a quick view of the person's background, relevant experience, and any questions they had asked. The transcript was there when he needed it, but the summary usually told him where to look.

Jonathan was still the decision-maker. That became especially clear when a score did not match what he saw in the interview. One candidate he considered an extremely strong fit received a default score of 3.8 out of 5. He checked the evidence and timestamps, then trusted his own review.
Jonathan did not settle for the first interview setup. He ran three tests back to back and changed the prompt after each one.
The first version used Ribbon's short default introduction. It asked the questions, but gave the candidate almost no context about Cartesia or the role.
The second explained how the screen would work. Better, but still one-way.
For the third, Jonathan added company and role context, mentioned that the interviewer was using one of Cartesia's own voices, and invited the candidate to ask questions before the screen began.
The conversation changed immediately. The candidate asked about Cartesia's products, and the interviewer accurately explained Sonic and Ink. It answered questions about working hours, meetings, and tooling. When the candidate asked about pay, it handed the question back to the hiring team instead of making something up.
The lesson was practical: give the interviewer the same context you would give a recruiter before a call.
Cartesia builds real-time speech and transcription models for voice agents. Jonathan listened as both the hiring manager and someone who builds the technology underneath the conversation.
He noticed background artifacts, delays between turns, pacing, and whether the voice reacted naturally after a detailed answer. He tested several voices and eventually chose Rafael.
A voice can be perfectly clear and still feel strange. A pause runs too long. A response to a thoughtful answer lands flat. Jonathan noticed those moments and sent detailed feedback. He also offered to recommend more Cartesia voices and help test them in several of the languages Ribbon supports.
Ribbon took 43 repetitive first screens off Jonathan's calendar. It did not make the hiring decisions. He reviewed each interview, checked the evidence, and chose who moved forward.
The process had drop-off too. About 23 invited candidates did not complete an interview, and one person declined because the screen used AI. Those numbers belong in the story alongside the 43 completions and 20 hires.
For Jonathan, the trade was still clear. He got a much better read on communication than he could get from a cover letter, without spending the month in back-to-back screening calls.
Jonathan's feedback was candid. Generated scores did not always match his assessment. Some integrity flags looked like false positives. Candidate questions needed better follow-up, and he wanted more control over the intent behind each interview question.
A real hiring process surfaced rough edges that a polished demo would miss. Jonathan's team made 20 hires. His feedback now feeds directly into scoring, follow-ups, integrity checks, and the question editor.
If 43 first-round calls would swallow your month, talk to us. We will help you set up the screen; your team keeps the final call.
Cartesia builds real-time voice and transcription models. Sonic handles text-to-speech, while Ink handles speech-to-text for responsive voice applications.