A two-hour call, 24,106 speaker turns

In short. FlowBridge comes from two separate needs: transcribing calls reliably and dictating text on iPhone. The two paths meet on the Mac. The question that holds them together is not how to get a text, but how far that text can be trusted.

The file was there, the conversation wasn't

A 130-minute recording came back from the transcriber split into 24,106 speaker turns: about three changes of voice per second. The system had produced a text file, but not a conversation anyone could read. The other person's voice, played through the Mac's speakers, had come back into the microphone. The program heard it on both channels and broke the sentences into fragments.

When an error reaches the decision log

The project had started much more simply. OBS saved the call audio; a script waited for the file to be complete, sent it off for transcription and archived the result. Then those transcripts started being used to reconstruct and cite decisions. “GreenWatt” misheard as “Greenback” made it into three entries of the decision log, even though nobody had ever proposed that name. From then on the problem was no longer getting text automatically. It was being able to tell how far to trust that text.

What the transcriber does today

The transcriber grew around that question. It uses a vocabulary for the names speech recognition often gets wrong; it checks the corrections the AI proposes and drops them if they change figures, speakers or too much content. It keeps the link to the recording and never rewrites an archived transcript: some decisions cite it by line number.

It also had to learn to survive ordinary failures. A 26-minute call, for example, had been transcribed twice and then set aside because of an error while generating the title. Today the job resumes from the stage it reached and recognises audio it has already processed, without paying for the same transcription again.

On iPhone: the app listens, the keyboard types

FlowBridge for iPhone started separately, for a different need: speaking and seeing the words appear where you are typing. At first it ran only on the device, with Whisper. The challenge was making that simplicity live with the limits of iOS: a regular app can't freely insert text into other apps, while a custom keyboard is not the right place to record audio and load a speech model.

So FlowBridge splits the work: the app listens and transcribes, the keyboard inserts the text. Then came launch shortcuts, the Dynamic Island, recovery of an interrupted dictation and an optional cloud mode, with consent and a personal key. Speed took real work too: on iPhone the first model load could take tens of seconds, so it had to be prepared before the user pressed Record.

On the Mac the two paths meet

Only later did the two paths start meeting on the Mac. There FlowBridge brings together local dictation, call recording and a journal. The new recorder captures microphone and system audio separately; the transcriber tries to remove the echo from a copy, leaving the original untouched. The engine wasn't chosen because it removed more decibels of echo, but because in testing it kept more of the speaker's words without attributing the other person's words to them.

Where it stands

That is the challenge holding the project together: making voice immediately useful without hiding the errors behind nicely formatted text. The call pipeline is in production; the apps and the journal across Mac and iPhone are still in full on-device testing. It is an important line, because FlowBridge was born precisely from the times a result looked ready and a real call proved otherwise.

CodeFlowBridge for iPhone on GitHub ↗

FAQ
Why doesn't the iPhone keyboard record the audio?
A custom keyboard is not the right place to record audio and load a speech model. That's why the app listens and transcribes, and the keyboard only inserts the text.
Does FlowBridge send your voice to the cloud?
Not necessarily. It was built to run only on the device, with Whisper. Cloud mode is optional and requires consent and a personal key.
How do you stop the AI from changing the meaning of a transcript?
Corrections proposed by the AI are checked and dropped if they change figures, speakers or too much content. An archived transcript is never rewritten.
Is it already in use?
The call pipeline is in production. The apps and the journal across Mac and iPhone are in full on-device testing.

Well-formatted text is not yet reliable text. The difference is knowing where it can be wrong.

Got a process that looks ready until a real case shows up?

Start with the pre-diagnosis