TranscriptLab Nova is a self-hosted transcription workspace for recorded classes and lectures. It is designed for homelab deployment and built as an open-source, AI-agent-driven project.
- Organizes recordings into folders and projects
- Uploads audio and video files in batches
- Queues transcription jobs for background processing
- Lets users review transcripts alongside media playback
- Exports transcripts in PDF, Markdown, TXT, and HTML
- Surfaces storage usage for folders and project workspaces
Implementation is underway. Phases 1-3 (scaffold, shared contracts, backend foundation) are complete. The backend and frontend projects build and the database schema is ready.
AGENTS.mdclass-transcriber-shared-api-contract.mdclass-transcriber-frontend-prd.mdclass-transcriber-backend-prd.mdclass-transcriber-frontend-tech-stack-requirements.mdclass-transcriber-backend-tech-stack-requirements.md
- Frontend: React, TypeScript, MUI, SWR, wouter
- Backend: ASP.NET Core Minimal APIs, EF Core, SQLite, BackgroundService
- Storage: local filesystem for media, prepared audio, and exports
- Processing: in-process queue with conservative concurrency for homelab hardware
- Intel i3-12100F
- 16 GB RAM
- Intel Arc A310 4 GB
Default processing assumptions:
- one transcription job at a time
- optional GPU acceleration
- conservative model selection
Before implementing anything:
- Read
AGENTS.md. - Read the relevant PRD.
- Read
class-transcriber-shared-api-contract.md. - Read the relevant tech stack requirements file.
Do not implement against assumptions that are not documented.
- Repository scaffolding
- Shared contracts and types
- Backend storage, folders, and settings
- Batch upload and project creation
- Queue worker and transcription pipeline
- Frontend pages and data flows
- Exports, playback polish, and Playwright validation
MIT
- .NET 10 SDK
- Node.js 24 LTS
- npm
- FFmpeg (for media processing)
cd src/ClassTranscriber.Api
dotnet runThe API runs at http://localhost:5000. Swagger UI is available at /swagger in development.
In local development, SQLite and all runtime artifacts are stored under the repo-root data/ directory rather than under src/ClassTranscriber.Api/.
When a WhisperNet model is selected for the first time, the backend can download the missing ggml-*.bin file into the configured models/ directory automatically.
The default upload request limit is 10 GiB. Override it with Uploads__MaxRequestBodySizeBytes if you need a different ceiling. Multipart upload buffering uses the configured storage temp directory by default.
TranscriptLab Nova can expose its four read-only transcript-source tools through
an opt-in, stateless MCP endpoint at /mcp. It is disabled by default. The
private MCP Compose overlay publishes this endpoint only on host loopback port
5001; the normal UI, REST API, and health endpoint remain on port 5000. CUDA
and OpenVINO images use the same MCP implementation and need no distinct setup.
For the private-server setup, trust boundary, verification, and teardown, see
docs/mcp-transcript-source-runbook.md.
An independently managed host client connects to http://127.0.0.1:5001/mcp.
Its installation, credentials, lifecycle, health, and entitlement are outside
this repository.
Runtime data is split by environment:
- Local development:
- SQLite database:
./data/transcriptlab.db - uploads, extracted audio, transcripts, exports, temp files, and downloaded models:
./data/
- SQLite database:
- Docker / container runtime:
- SQLite database:
/data/transcriptlab.db - uploads, extracted audio, transcripts, exports, temp files, and downloaded models:
/data/
- SQLite database:
The app intentionally keeps the database and filesystem artifacts under the same configurable base path so a single Docker volume captures the whole workspace.
The backend supports multiple transcription engines. Use GET /api/settings/options to query the currently available engines and their models at runtime.
Uses the official SherpaOnnx .NET runtime through an isolated helper worker process.
Model download:
# Download all registered models (small + medium)
./scripts/download-sherpa-models.sh all
# Or download a single model
./scripts/download-sherpa-models.sh smallModels are placed under the configured path (default /data/models/sherpa-onnx/<model>/). When auto-download is enabled, the backend can also fetch a missing registered model on first use.
Model directory layout:
Each model directory must contain a config.json describing the model backend and file names. Two layouts are supported:
Whisper backend (encoder/decoder pair):
/data/models/sherpa-onnx/small/
├── config.json
├── tiny-encoder.onnx
├── tiny-decoder.onnx
└── tiny-tokens.txt
SenseVoice backend (single model):
/data/models/sherpa-onnx/<model>/
├── config.json
├── model.onnx
└── tokens.txt
The config.json selects the backend and maps file names. Example for the whisper backend:
{
"backend": "whisper",
"encoder": "tiny-encoder.onnx",
"decoder": "tiny-decoder.onnx",
"tokens": "tiny-tokens.txt",
"task": "transcribe"
}Configuration in appsettings.json:
{
"Transcription": {
"SherpaOnnx": {
"ModelsPath": "/data/models/sherpa-onnx",
"Provider": "cpu",
"NumThreads": 4,
"AutoDownloadModels": true,
"LogSegments": false,
"ModelDownloadBaseUrl": "https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models"
}
}
}Uses the Whisper.net managed library with the Whisper.net.Runtime (CPU) native backend through an isolated helper worker process. Models use shared ggml-*.bin assets and are auto-downloaded on first use.
No extra setup needed. The NuGet packages are included in the project.
Uses the Whisper.net managed library with the stable Whisper.net.Runtime.Cuda native backend through the same isolated helper worker process.
Prerequisites:
- NVIDIA GPU with CUDA support
- CUDA runtime/driver libraries visible to the app
- For Docker: NVIDIA Container Toolkit and GPU device exposure to the container
Models use shared ggml-*.bin assets and are auto-downloaded on first use. The backend probes for CUDA runtime libraries before each job and returns a clear failure if the host/container cannot load them.
Uses the Whisper.net managed library with the Whisper.net.Runtime.CoreML backend through the same isolated helper worker process.
Prerequisites:
- Native macOS ARM64 / Apple Silicon process
- FFmpeg installed on the host, for example with
brew install ffmpeg - Shared
ggml-*.binWhisper model file - Matching CoreML encoder package next to the ggml file, for example:
models/
├── ggml-large-v3-turbo.bin
└── ggml-large-v3-turbo-encoder.mlmodelc/
Linux Docker containers on macOS do not receive direct Apple Metal/CoreML/ANE access, so Docker-only Whisper should be treated as CPU-only on Apple Silicon. For Mac mini M4 deployments, prefer a native macOS publish or run a native host transcription sidecar and call it from Docker through an OpenAI-compatible endpoint.
Uses the supported OpenVINO Whisper sidecar for Intel GPU acceleration. The backend starts a local FastAPI sidecar process that loads OpenVINO Whisper models through the openvino-genai Python package and exposes an OpenAI-compatible /v1/audio/transcriptions endpoint.
Prerequisites:
- Python 3 with the sidecar requirements from
src/ClassTranscriber.Api/Tools/requirements-openvino-sidecar.txt - Intel GPU runtime support visible to the host or container
OpenVINO Whisper models are stored under data/models/openvino-genai/ and can be downloaded on first use or through POST /api/settings/models/manage.
OpenRouter is available as a first-class remote engine when a server-side API key is configured. TranscriptLab discovers the current speech-to-text model catalog at runtime, stores the selected model with each project, and keeps remote models out of the local Model Manager.
Configure the key with an environment variable so it is never committed:
export Transcription__OpenRouter__ApiKey=YOUR_OPENROUTER_API_KEYOptional settings can override Transcription__OpenRouter__BaseUrl, FallbackModels, and TimeoutSeconds. BaseUrl must be an absolute HTTPS URL. All hosted transcription audio is prepared as lossless FLAC. OpenRouter's ordinary transcription catalog remains dynamic. Its verified long-form word-timestamp path is limited to openai/whisper-large-v3 and openai/whisper-large-v3-turbo: it sends sequential 600-second cores with up to two seconds of extraction overlap, recursively splits an encoded part at or above 24,000,000 bytes, and only sends FLAC parts strictly below that limit. It checkpoints successful chunks, sums actual reported usage cost, retries only 429/503 up to three total attempts honoring Retry-After, and treats timeouts as fatal.
Configure Transcription__Xai__ApiKey to enable the Xai engine and grok-stt-1.0. TranscriptLab sends one whole prepared FLAC request to xAI /v1/stt, preserving native speaker identities across long classes. The FLAC is checked against xAI's 500 MB limit; recordings over the limit fail validation and are never silently chunked or rerouted.
Direct xAI STT cost is displayed as an estimate using the configured hourly-rate snapshot (default $0.10/hour). OpenRouter-reported costs are displayed as actual. Optional speaker-role attribution sends timestamped speaker turns, never audio, to google/gemini-3.7-flash through OpenRouter and fails open to the original Speaker N labels.
When diarization is enabled, choose its source explicitly. Provider mode uses xAI's native speaker IDs and is offered only for direct Xai with grok-stt-1.0; Local mode runs TranscriptLab's local diarizer and exposes the Basic and Improved choices. xAI timing is available only for the two verified OpenRouter models when direct xAI is configured: OpenRouter wording stays unchanged, then a whole-FLAC xAI timing request assigns speakers by greatest positive overlap, then nearest interval within one second. An explicit xAI timing failure fails the job rather than falling back.
The Settings page keeps Settings, Local Model Manager, and System Capabilities as separate tabs. Model catalog and capability data load only when their tab is first selected. /api/diagnostics remains a separate route. GET /api/settings/capabilities returns sanitized provider, compute, and CPU-summary fields only; it never returns credentials, URLs, paths, raw provider responses, stack traces, or raw exceptions.
Optional paid smoke tests are intentionally skipped when credentials or approved disposable media are unavailable. Skipping that optional class is recorded as skipped, not as a passing full-provider test; deterministic mocked coverage remains required.
Configuration for all WhisperNet engines in appsettings.json:
{
"Transcription": {
"LogSegments": false,
"WhisperNet": {
"AutoDownloadModels": true,
"LogSegments": false
}
}
}When Transcription:LogSegments or an engine-specific ...:LogSegments option is set to true, the worker logs each decoded transcript segment to container stderr so it appears in docker compose logs -f. The default is false to avoid noisy logs on long recordings.
Publish the native app and self-contained WhisperNet worker for Apple Silicon:
./scripts/publish-macos-coreml.shSuggested host layout:
/Users/transcriptlab/transcriptlab/
├── app/
│ ├── ClassTranscriber.Api
│ └── ClassTranscriber.WhisperNet.Worker
└── data/
├── transcriptlab.db
├── uploads/
├── audio/
├── transcripts/
├── exports/
├── models/
└── temp/
Recommended runtime environment:
export ASPNETCORE_ENVIRONMENT=Production
export Storage__BasePath=/Users/transcriptlab/transcriptlab/data
export Transcription__FFmpegPath=/opt/homebrew/bin/ffmpeg
export Transcription__WhisperNet__WorkerPath=/Users/transcriptlab/transcriptlab/app/ClassTranscriber.WhisperNet.WorkerUse an absolute FFmpeg path because macOS services launched through launchd may not inherit the same shell PATH as an interactive terminal. On Apple Silicon Homebrew, FFmpeg is usually at /opt/homebrew/bin/ffmpeg.
cd src/frontend
npm install
npm run devThe dev server runs at http://localhost:5173 and proxies /api requests to the backend.
# Backend
cd src && dotnet build ClassTranscriber.slnx
# Frontend
cd src/frontend && npm run build
# Tests
cd src && dotnet test ClassTranscriber.slnx
cd src/frontend && npm testdocker compose up --buildThe default image uses Dockerfile and is intended for CPU-only runs. The application runs at http://localhost:5000 with data persisted in a Docker volume.
Large uploads use the same 10 GiB default ceiling in Docker through appsettings.json. Docker images set ASPNETCORE_TEMP=/data/temp so multipart upload buffering uses the mounted data volume instead of the container overlay. Override Uploads__MaxRequestBodySizeBytes in Compose if needed.
To try NVIDIA CUDA inside Docker, use the optional override:
docker compose -f docker-compose.yml -f docker-compose.cuda.yml up --buildThe override switches the build to Dockerfile.cuda, which uses an NVIDIA CUDA runtime base image and installs aspnetcore-runtime-10.0 inside it. This requires the NVIDIA Container Toolkit on the host so the container can access the GPU and driver libraries.
To try Intel OpenVINO inside Docker, use the optional override:
docker compose -f docker-compose.yml -f docker-compose.openvino.yml up --buildThe OpenVINO override switches the build to Dockerfile.openvino, exposes /dev/dri, and sets Transcription__OpenVinoWhisperSidecar__Device=GPU by default.
For Intel Arc hosts, the OpenVINO image also needs the system Intel OpenCL libraries to take precedence over older Intel compiler libraries inherited from the base image. Dockerfile.openvino now enforces that linker order, which was required to restore GPU compilation inside Docker on the Arc A310 used to develop and validate this project.
For Intel Arc hosts, /dev/dri is the device mapping you want. It exposes both the card* and renderD* nodes that OpenVINO uses. In the validated Docker setup for this repo, the Arc A310 is exposed as GPU inside the container. If a host enumerates multiple Intel GPUs differently, inspect the sidecar /devices output before overriding the device name.
OPENVINO_DEVICE=GPU docker compose -f docker-compose.yml -f docker-compose.openvino.yml up --buildTo inspect the host DRM nodes:
ls -l /dev/driFor CasaOS custom installs, use docker-compose.casaos.yml. It is image-based rather than build-based and defaults to the published CPU package:
ghcr.io/snavatta/transcriptlab-nova-cpu:latestThe CasaOS file stores all runtime data under:
/DATA/AppData/$AppID/dataIf you want a GPU-backed CasaOS install, change the image in that file to one of:
ghcr.io/snavatta/transcriptlab-nova-cuda:latestghcr.io/snavatta/transcriptlab-nova-openvino:latest
For OpenVINO on CasaOS, also add:
devices:
- /dev/dri:/dev/dri
environment:
Transcription__OpenVinoWhisperSidecar__Device: GPUIf a host reports multiple Intel GPUs in the sidecar /devices output, override the device value accordingly.
For CUDA on CasaOS, the host still needs NVIDIA Container Toolkit and GPU runtime support.
The repository is set up to publish three public GHCR images from GitHub Actions:
ghcr.io/<owner>/transcriptlab-nova-cpughcr.io/<owner>/transcriptlab-nova-cudaghcr.io/<owner>/transcriptlab-nova-openvino
Each package is published for linux/amd64 and receives:
latest- the pushed semver tag such as
v1.2.3 - an immutable
sha-...tag
Example pulls:
docker pull ghcr.io/<owner>/transcriptlab-nova-cpu:latest
docker pull ghcr.io/<owner>/transcriptlab-nova-cuda:latest
docker pull ghcr.io/<owner>/transcriptlab-nova-openvino:latestExample runs:
# CPU
docker run --rm -p 5000:5000 -v transcriptlab-data:/data \
ghcr.io/<owner>/transcriptlab-nova-cpu:latest
# CUDA
docker run --rm -p 5000:5000 -v transcriptlab-data:/data \
--gpus all \
-e NVIDIA_VISIBLE_DEVICES=all \
-e NVIDIA_DRIVER_CAPABILITIES=compute,utility \
ghcr.io/<owner>/transcriptlab-nova-cuda:latest
# OpenVINO
docker run --rm -p 5000:5000 -v transcriptlab-data:/data \
--device /dev/dri:/dev/dri \
-e Transcription__OpenVinoWhisperSidecar__Device=GPU \
ghcr.io/<owner>/transcriptlab-nova-openvino:latestIf your GitHub Packages defaults keep newly published container packages private, set the package visibility to Public after the first publish.
Host prerequisites:
- CPU image:
- no GPU runtime required
- CUDA image:
- NVIDIA GPU
- NVIDIA Container Toolkit
- driver/runtime libraries available to the container
- OpenVINO image:
- Intel GPU or supported Intel accelerator
/dev/dridevice exposure into the container- host graphics stack/driver support compatible with OpenVINO GPU execution
- Intel GPU or supported Intel accelerator
/dev/dridevice exposure into the container- host graphics stack/driver support compatible with OpenVINO GPU execution
- Python OpenVINO GenAI runtime supplied by the dedicated image