How we work

What happens to your material — step by step.

A controlled workflow from source intake to final delivery — documented, reviewable and designed for difficult archival audio.

Difficult archival audio cannot be resolved in a single processing step. R2 Mechanics combines source preservation, independent transcription and speaker-analysis paths, expert review and structured delivery in one traceable workflow.

This is how your material moves from original source to validated, structured output.

Step 1

Scope & Protect

Before processing begins, we define the source material, handling requirements, confidentiality conditions, deliverables and project scope.

Every project begins with a documented scope that defines the processing conditions and expected outputs.

Step 2

Preserve the Source

A controlled working copy is created from the source material. The original remains unchanged and is preserved as the reference source throughout the project.

Working copies and processing stages remain documented and traceable.

Step 3

Analyse Through Independent Paths

The material is examined through multiple independent transcription, alignment and speaker-analysis paths. Each contributes evidence to the final result rather than being treated as automatically correct.

Agreement across independent paths strengthens the result. Divergences identify passages that require closer review and remain visible in the processing record.

Speaker attribution is analysed as a dedicated task, separate from speech recognition itself.

Diagnostic views make difficult regions of a recording visible early in the workflow. This allows additional audio preparation, re-analysis and human review to be concentrated on the passages that require it rather than applied uniformly across the entire recording.

Step 4

Review, Validate & Select

Results from the independent analysis paths are compared and reviewed. The strongest supported result is selected as the validated project baseline.

Selection is deliberate and traceable. The validated transcript remains linked to the evidence and processing history behind it.

Step 5

Structure and Deliver

The validated transcript is converted into structured, navigable project outputs with timestamps, speaker attribution and the formats defined in the project scope. Where editorial refinement is required, it is performed as a separate layer. The source-faithful baseline remains preserved and unchanged.

Source-faithful and editorial outputs remain clearly separated and version-linked throughout delivery.

Architecture

How R2 Mechanics works

R2 Mechanics treats audio and video as a navigable evidence timeline rather than reducing the source to a single transcript. The original stays the reference. Speaker, language and transcription evidence are mapped on one timeline first; anything optional is added on top of it, never in its place.

How R2 Mechanics works: conceptual architecture Conceptual diagram. Source media is analysed on a media and timeline layer. Speaker regions, language regions and transcription evidence are combined on one evidence and source timeline. From this timeline the workflow delivers an interactive archive and structured exports. Optional translation and optional LLM analysis are separate derived layers. A separate optional branch extracts text from reference material, uses OCR where required, aligns it with the media-derived transcript and supports review. SOURCE MEDIA audio · video original preserved MEDIA & TIMELINE ANALYSIS speech · timing · segmentation Speaker regions who speaks when Language regions incl. rapid switches Transcription evidence several independent perspectives EVIDENCE & SOURCE TIMELINE speaker · language · text source review-worthy regions navigation back to the source Interactive Archive player · transcript maps · filters Structured Export JSON · SRT · VTT Optional Translation derived view; source text stays reference Optional LLM Analysis chapters · summaries entities · notes OPTIONAL · REFERENCE-ASSISTED REVIEW Reference material transcripts · PDFs scans · documentation Text extraction / OCR OCR only where the material requires it Reference alignment compared with the media-derived transcript Review support discrepancy identification
Conceptual overview. Dashed elements are optional and depend on the project.

Technical architecture & public methodology on GitHub →

Evidence Mapping & Timeline Intelligence

Speaker regions, language regions and the origin of the selected text are mapped along the original media timeline, together with passages that deserve review. Every passage links back to the exact point in the source.

For long, multilingual or hard-to-transcribe recordings, visual maps make this structure visible at a glance. Review indications are navigation cues for human reviewers, not a statement of correctness.

Multi-engine processing

Several independent transcription perspectives can be applied to the same material. Differences are compared and handled in a documented way, and the provenance of the selected text can be preserved.

Difficult passages do not have to disappear invisibly into the final text.

Multilingual processing & translation

Several languages in one recording, including rapid switches, are structured on a language timeline with language filters. Specialised workflows exist for multilingual material.

Translation is an optional derived access layer. The source-language transcript remains preserved as the reference, and the archive can switch between the source-language and translated view.

Reference documents, OCR & alignment

Existing transcripts, PDFs, scans and archival documentation can be text-extracted and, where required, OCR-processed in a supporting workflow before comparison with the media-derived transcript. This can support review, alignment and the identification of discrepancies.

Original reference documents remain unchanged; extracted or OCR-derived text is treated as a derived working representation. The media source stays the primary reference, and nothing in the transcript is overwritten automatically. This branch is optional.

Optional LLM analysis

Chapters, summaries, entities, context notes or a focused analysis can be added where a workflow asks for them.

Source evidence and interpretation stay separate: the LLM layer is not the source of the transcript and not a source authority.

Fast media workflow

For quick access to a recording: synchronized player, transcript and navigation, with an optional focused summary and without the full editorial cascade.

It skips the extended optional analysis, not the evidence timeline. The maps and navigation available for the material remain usable.

Our principle

Evidence before polish.

The evidential record comes first. Editorial refinement, presentation and publication layers are built from that foundation without replacing it.

When languages shift.

Mixed-language recordings require more than multilingual speech recognition. Speaker changes, language transitions, code-switching and recording quality must be assessed together to maintain continuity across the source.

For multilingual and mixed-language collections, the processing strategy is adapted to the actual language, speaker and recording conditions of the material. We adapt the analysis to your material — not the other way around.

Questions about how your material would be handled?

Discuss your collection →