How we work
A controlled workflow from source intake to final delivery — documented, reviewable and designed for difficult archival audio.
Difficult archival audio cannot be resolved in a single processing step. R2 Mechanics combines source preservation, independent transcription and speaker-analysis paths, expert review and structured delivery in one traceable workflow.
This is how your material moves from original source to validated, structured output.
Step 1
Before processing begins, we define the source material, handling requirements, confidentiality conditions, deliverables and project scope.
Every project begins with a documented scope that defines the processing conditions and expected outputs.
Step 2
A controlled working copy is created from the source material. The original remains unchanged and is preserved as the reference source throughout the project.
Working copies and processing stages remain documented and traceable.
Step 3
The material is examined through multiple independent transcription, alignment and speaker-analysis paths. Each contributes evidence to the final result rather than being treated as automatically correct.
Agreement across independent paths strengthens the result. Divergences identify passages that require closer review and remain visible in the processing record.
Speaker attribution is analysed as a dedicated task, separate from speech recognition itself.
Diagnostic views make difficult regions of a recording visible early in the workflow. This allows additional audio preparation, re-analysis and human review to be concentrated on the passages that require it rather than applied uniformly across the entire recording.
Step 4
Results from the independent analysis paths are compared and reviewed. The strongest supported result is selected as the validated project baseline.
Selection is deliberate and traceable. The validated transcript remains linked to the evidence and processing history behind it.
Step 5
The validated transcript is converted into structured, navigable project outputs with timestamps, speaker attribution and the formats defined in the project scope. Where editorial refinement is required, it is performed as a separate layer. The source-faithful baseline remains preserved and unchanged.
Source-faithful and editorial outputs remain clearly separated and version-linked throughout delivery.
Architecture
R2 Mechanics treats audio and video as a navigable evidence timeline rather than reducing the source to a single transcript. The original stays the reference. Speaker, language and transcription evidence are mapped on one timeline first; anything optional is added on top of it, never in its place.
Speaker regions, language regions and the origin of the selected text are mapped along the original media timeline, together with passages that deserve review. Every passage links back to the exact point in the source.
For long, multilingual or hard-to-transcribe recordings, visual maps make this structure visible at a glance. Review indications are navigation cues for human reviewers, not a statement of correctness.
Several independent transcription perspectives can be applied to the same material. Differences are compared and handled in a documented way, and the provenance of the selected text can be preserved.
Difficult passages do not have to disappear invisibly into the final text.
Several languages in one recording, including rapid switches, are structured on a language timeline with language filters. Specialised workflows exist for multilingual material.
Translation is an optional derived access layer. The source-language transcript remains preserved as the reference, and the archive can switch between the source-language and translated view.
Existing transcripts, PDFs, scans and archival documentation can be text-extracted and, where required, OCR-processed in a supporting workflow before comparison with the media-derived transcript. This can support review, alignment and the identification of discrepancies.
Original reference documents remain unchanged; extracted or OCR-derived text is treated as a derived working representation. The media source stays the primary reference, and nothing in the transcript is overwritten automatically. This branch is optional.
Chapters, summaries, entities, context notes or a focused analysis can be added where a workflow asks for them.
Source evidence and interpretation stay separate: the LLM layer is not the source of the transcript and not a source authority.
For quick access to a recording: synchronized player, transcript and navigation, with an optional focused summary and without the full editorial cascade.
It skips the extended optional analysis, not the evidence timeline. The maps and navigation available for the material remain usable.
Our principle
The evidential record comes first. Editorial refinement, presentation and publication layers are built from that foundation without replacing it.
Mixed-language recordings require more than multilingual speech recognition. Speaker changes, language transitions, code-switching and recording quality must be assessed together to maintain continuity across the source.
For multilingual and mixed-language collections, the processing strategy is adapted to the actual language, speaker and recording conditions of the material. We adapt the analysis to your material — not the other way around.