Prakhar Goel

@prakhar98

Streaming Through Failure: Designing Reliable Transcription Streams

Submitted Sep 27, 2026

Reconnecting a WebSocket restores the connection, but not the state of an interrupted stream. At Adalat AI, we faced this in live courtroom transcription; I’ll show how we retransmit audio the server missed and replay transcript updates the client missed after a network drop.

What problem are you addressing?

At Adalat AI, we work to reduce India’s judicial backlog by helping courtrooms record proceedings more efficiently. In lower courts, judges often write proceedings by hand; stenographers take shorthand and prepare a fair copy later.

Our live-transcription system is now used across 11 states. In the last quarter alone, judges recorded more than 1.5 million minutes (~25,000 hours) on our platform. The system captures audio in the browser, streams it to a transcription backend, and returns text while the hearing is underway.

That workflow depends on a two-way stream over internet connections that can be unstable. A WebSocket can fail silently, and reconnecting does not tell us which audio chunks the backend processed or which transcript updates reached the browser. Retransmitting from the wrong point can duplicate work; resuming too far ahead can leave gaps. The engineering challenge is to keep one logical transcription session intact across connections.

Intended audience

Developers, Platform Engineers

Level:

Intermediate

Practical takeaways

  1. Design recovery across both sides of the connection: the client preserves queued and recently sent audio, while the server preserves the transcription session and recent transcript output. After reconnecting, the server reports the last input sequence it received and the client reports the last output sequence it received, allowing each side to retransmit or replay what the other missed.
  2. Design the UX for slow networks: measure audio delivery latency to your servers and use that signal to communicate connection quality when live transcription falls behind.

What will you share?

The session follows the evolution of our reliability design as each fix exposed another failure. We began by trusting WebSocket close events and browser network signals, only to discover that a broken connection could remain apparently open. Application-level keepalives gave us a reliable way to detect these silent failures and reconnect.

Reconnecting restored the transport, but it did not restore the stream. We introduced sequence tracking in both directions, preserved session state across connections, and retained bounded buffers so the client could retransmit missing audio while the server replayed missed transcripts.

Finally, chunk acknowledgements gave us a measure of real delivery latency that the interface could use to reflect changes in connection speed.
A simulated demo will introduce network failures at each stage and show the system detecting the failure, reconnecting, and recovering the stream.

Experience with this problem

Hard-earned engineering lesson

What approaches failed or created unexpected problems?

We initially trusted WebSocket close events and browser network signals to tell us when connectivity was lost. In practice, switching networks or losing internet access could leave the socket apparently open, and calls to send() could continue without proving that the server received the data. Even calling close() did not guarantee that the browser would promptly invoke the close handler.

What will you do differently today?

  • Begin with the people using the dictation system and the network conditions available to them.
  • Build for real-world conditions such as low bandwidth, high latency, network switching, and extended outages.
  • Read the browser WebSocket specification carefully instead of assuming API behavior. For example, opening a WebSocket provides no configurable connection timeout, so the application must enforce one.
  • Test simulated network degradation early in development.

What trade-offs did you consider?

We needed to send a continuous audio stream from the browser while returning partial transcripts over the same session. We considered three transport approaches:

  • HTTP requests: simple to implement and operate, but sending audio as repeated requests adds overhead and makes ordering, retries, and continuous transcript delivery harder to coordinate.
  • gRPC streaming: provides efficient bidirectional streaming and strongly typed contracts, but browser support and the additional proxy layer made it less suitable for our client-facing connection.
  • WebSockets: provide a browser-native, full-duplex connection with low per-message overhead. We chose them for the live stream, while accepting that liveness detection, retransmission, replay, buffering, and recovery semantics would need to be implemented at the application layer.

Within the WebSocket design:

  • Bounding client queues and server replay buffers prevents unbounded memory growth, but limits how far back a disconnected session can recover.
  • Keepalives improve failure detection during idle periods but add traffic.

How can this help other practitioners?

  • A practical approach to building high-throughput bidirectional streams over unreliable networks using logical sessions, sequence positions, and bounded buffers on both sides.
  • A method for detecting silent failures, measuring connection quality, and testing recovery through simulated disconnections, network switches, latency, and bandwidth degradation.

Current state

Production experience

Tags

#reliability #websockets #realtimesystems #streaming #networking #platformengineering #casestudy #failurestory

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy