Use and evaluate offline dictation¶
Desktop includes local Whisper dictation. Open Settings → Global → Dictation, enable it, and choose the microphone, speech model, and spoken punctuation preference. The composer microphone starts recording; Pause permits edits; Stop releases capture. The Settings microphone test discards its text when you leave. Recording requires an explicit action and OS microphone permission. Audio stays on the machine and is discarded after transcription. Recognized text remains a draft until Send.
Tiny English (~78 MB) ships with Desktop for offline first use. Base English (~148 MB) is optional: download it from Settings or choose a verified local model file. Transfers use fixed URLs, sizes, and SHA-256 checks; partial files are never published as models. An unplugged selected microphone produces an error instead of switching inputs.
Preferences and models live in COLOSSUS_HOME/dictation (normally ~/.colossus/dictation).
Desktop and the local TUI share these device-local choices. They apply to the next
recording, independently of workspaces and external agent targets. Model paths and
audio never enter the renderer or agent protocol. The shared native engine lives in
crates/colossus-native-dictation.
Dictate in the local TUI¶
The Desktop-bundled CLI includes dictation. Enable it explicitly for source builds:
GGML_NATIVE=OFF cargo build --locked -p colossus-cli --features dictation
./target/debug/colossus tui
Use /dictate on, then F4 to start, pause, or resume. Shift+F4 stops. Enter
finalizes pending speech before sending; multiline mode uses Ctrl+D. Active speech
updates cannot overwrite manual edits. Pause before editing. Escape or Ctrl+C cancels
recording; exit, session navigation, and operator decisions release the microphone.
Use /dictate settings for current choices and command help:
/dictate microphoneslists inputs;/dictate microphone NUMBERselects one./dictate microphone defaultrestores the OS default./dictate install tinyor/dictate install baseexplicitly downloads a pinned model./dictate model tinyor/dictate model baseselects an installed model./dictate punctuation onor/dictate punctuation offcontrols spoken commands./dictate offdisables future recording.
Headless CLI builds omit native audio dependencies. A CLI running over SSH captures on the machine running the CLI. Microphone access is not forwarded through the worker or agent connection. Native permissions, input devices, and latency still need live acceptance on each supported platform.
Package the included model¶
Desktop dev/build and macOS/Windows packaging stage Tiny English before native
assembly. release/dictation/models.json pins the immutable revision, size, and digest;
release/dictation/LICENSE-MIT retains the upstream license. Weights are ignored build
assets, never Git source or renderer assets. Stage a reviewed file without networking:
Downloaded and cached bytes are verified. Native resources include the model, license,
and provenance. Cargo defaults to GGML_NATIVE=OFF for CPU portability; an explicit
environment value can override it for local profiling. macOS signing grants audio
input to the app and bundled CLI. Renderer
chunk and fixture checks remain active, with a 4,050,000-byte total budget including the
regular dictation controls and lazy Settings page.
Build and install a candidate¶
Use the source toolchain, a C++ compiler, CMake, and libclang. Windows needs the MSVC C++ build tools; macOS needs Xcode command-line tools. Linux also needs ALSA development headers for the capture adapter. Dependencies come from the locked root Cargo graph. Build a portable CPU baseline from the repository root:
GGML_NATIVE=OFF cargo build --locked --release \
--package colossus-native-dictation --features probe --bin dictation-probe
On PowerShell, set $env:GGML_NATIVE = 'OFF' before the same Cargo command.
On macOS, evaluate --features metal separately after the CPU baseline. The
metal_requested output identifies the requested backend; confirm actual GPU use
with platform profiling. It does not establish accelerator availability.
The candidate is whisper-rs 0.16.0 / whisper-rs-sys 0.15.0, containing
whisper.cpp 1.8.3. The bindings are Unlicense, whisper.cpp is MIT, CPAL 0.18.2 is
Apache-2.0, and Rubato 0.16.2 is MIT. The
converted OpenAI Whisper weights
declare MIT. These GGML conversions are third-party assets; retain both their
conversion provenance and the original Whisper license when distributing them.
Download a model explicitly from the pinned conversion revision
5359861c739e955e79d9a303bcbc70fb988958b1. The probe contains no downloader, network
client, API credential handling, automatic updates, or cloud fallback.
| Candidate | File | Bytes | SHA-256 from the pinned LFS metadata |
|---|---|---|---|
| English tiny | ggml-tiny.en.bin |
77,704,715 | 921e4cf8686fdd993dcd081a5da5b6c365bfde1162e72b08d75ac75289920b1f |
| English base | ggml-base.en.bin |
147,964,211 | a03779c86df3323075f5e796cb2ce5029f00ec8869eee3fdfb897afe36c6d002 |
Choose an ordinary local file outside the checkout. The probe rejects missing,
empty, symlinked, non-regular, oversized, and hash-mismatched models. It loads the
exact verified bytes, without reopening the pathname. An operator-supplied hash
checks integrity; authenticity depends on reviewing the pinned source above. Hashing
an arbitrary downloaded file and supplying that result does not establish provenance.
The probe creates no model store. Remove its model by deleting that exact installed
file. Desktop and TUI use the reviewed catalog in release/dictation/models.json:
Tiny English is bundled with Desktop, and Base English is an explicit download.
Open Desktop with a microphone control¶
From the repository root on macOS or Linux:
On Windows, set $env:GGML_NATIVE = 'OFF', run npm ci --ignore-scripts in
apps/desktop, then npm run tauri:dev. The historical --dictation and
tauri:dev:dictation commands remain aliases. Building the probe alone does not launch
Desktop. Close another Desktop instance before launching the development app.
Enable dictation in Settings → Global → Dictation, use Test dictation to check your input, then return to the thread and click the composer microphone.
Spoken punctuation is enabled by default in Dictation Settings. Say
“period” or “full stop” for ., “comma” for ,, “question mark” for ?,
“exclamation mark” or “exclamation point” for !, “colon” for :, “semicolon”
for ;, and “new line” or “new paragraph” for line breaks. For example,
“Check the build period Is it ready question mark” becomes
“Check the build. Is it ready?”. The formatter works on recognized speech only;
typed text and pasted content keep their wording.
This mode treats punctuation names as commands, including in ordinary phrases. Say “a literal period of time” to keep “a period of time”, or turn Spoken punctuation off before starting recording to retain all punctuation names as words. The setting stays fixed during a recording session. Multiword commands can span transcript segments; editing a settled draft discards that suffix when its text or position changes. Recognition mistakes can still require Pause and an edit. This is formatting after local transcription, so it cannot recover a spoken command that the model did not recognize.
The microphone control pauses and resumes recording. Pause finalizes captured speech and releases the input device before the draft becomes editable; the composer protects an active partial from competing edits. Resume continues with the edited draft. The adjacent stop control finalizes speech and closes the recording session. A clear recording strip remains visible while capture is active. Its cyan waveform shows recent microphone loudness, with a glow that follows the input level and a clock counting active recording time. Silence settles into dots; Pause freezes and dims the waveform and stops the clock. Starting shows a loading indicator. Reduced motion keeps a stationary input meter and disables decorative animation. The display estimates a quiet background from recent input levels and suppresses activity near that level, so steady room noise usually settles into dots after a short warmup. This affects only the visualization; it does not gate recording or recognition. Loud or changing background noise can still move the meter.
The native callback coalesces RMS loudness into one byte from 0 to 255. The existing
session poll delivers at most one level update alongside its bounded transcript
events, so metering stays responsive while inference is running without queuing
telemetry. The renderer retains 96 loudness values for the display, independently
of composer updates. Samples and model data stay in native code. The level history
is not a recording or a frequency spectrum; it measures input strength, including
background noise. Stale levels decay visually, and Pause, Stop, navigation, or a
failure clears the input meter. Stopping during model initialization cancels it.
Native transcript updates remove Whisper's exact [BLANK_AUDIO] silence marker,
including empty final revisions that replace an earlier partial.
Send settles a FIFO audio boundary before using the existing prompt/run or Next up path. The same microphone stream stays open. Audio queued after the boundary goes into the next draft, including speech recognized while the previous message is submitting. A failed submission keeps the previous draft and appends subsequent speech for review. Native replies are serialized, and thread/workspace navigation cancels recording and ignores late updates from the old session. Opening another Desktop surface also stops capture, as does switching to Topology or Activity where the composer is hidden. Speech never submits a turn without an explicit Send or Redirect action.
Check pause/edit/resume, successive sends, failed submission, permission denial and recovery, device removal, model failure, renderer reload, and application close on each native platform. If capture is unavailable, check the system's default input device and microphone privacy settings for Colossus or the launching terminal. The app includes a macOS microphone usage description and an audio-input entitlement configuration; actual platform permission prompts and hardened packaging still need acceptance.
Exercise offline inference and capture¶
The release executable is under target/release/ unless
CARGO_TARGET_DIR overrides it. Windows adds .exe. With the model installed,
disconnect networking and run:
target/release/dictation-probe \
/path/to/ggml-tiny.en.bin \
921e4cf8686fdd993dcd081a5da5b6c365bfde1162e72b08d75ac75289920b1f
Model verification and initialization happen before microphone access. Capture starts only for this explicit live invocation. Permission/device failure produces a categorical error with OS-settings recovery guidance. This console probe does not yet distinguish every native permission state or provide the production consent UI.
Speak continuously, then enter these commands on stdin:
pausereleases the microphone, drains captured audio, and finalizes the tail.resumeopens a new stream while preserving monotonic segment IDs.flushsettles audio already consumed by the pipeline while capture continues.stopreleases capture, finalizes the tail, and terminates the local helper. EOF also stops; Ctrl+C terminates the console process.
flush is a probe segmentation control, not an agent submission. Audio still queued
at that instant belongs to subsequent updates. Desktop instead inserts a
FIFO capture marker and settles the current draft before using the existing Send path.
Stdout contains JSON cold-start and per-revision metrics. It omits transcript text by
default; add --show-text only when deliberately inspecting recognition. Do not
redirect private transcripts into durable logs. Inference time excludes capture,
queue wait, and endpoint delay; measure total visible latency separately.
For reproducible inference without a microphone, add --wav /path/to/fixture.wav.
The fixture must already be a mono, 16 kHz, signed 16-bit WAV of at most 60 seconds.
The probe reads it in bounded chunks and creates no recording or temporary audio file.
Use the public
JFK sample at whisper.cpp v1.8.3
for a credential-free smoke check; its SHA-256 is
59dfb9a4acb36fe2a2affc14bacbee2920ff435cb13cc314a08c13f66ba7860e.
Interpret the native boundary¶
Capture downmixes supported input formats and uses an FFT resampler to produce 16 kHz mono audio. Common rates from 8 to 96 kHz and at most eight channels are accepted. Each callback is limited to 8,192 mono frames, the capture queue to 256 blocks (at most 8 MiB), and the active segment to six seconds. Partials are recomputed at two and four seconds; the six-second update finalizes that segment. Pause/stop also settle short tails, with padding excluded from the reported audio duration.
Each update replaces the complete text for one segment_id and increasing
revision. Append a segment to a settled draft exactly once when is_final is true;
an empty final replaces an earlier partial. Text is bounded to 8 KiB per segment.
The probe retains no cumulative transcript. Its owned audio buffers are zeroized
on release, with no durable recording or transcript logging by default. The native
inference library may retain internal process memory until its helper exits.
The helper is another invocation of the same probe executable, with bounded private stdin/stdout framing. Audio and model bytes stay within native code. The renderer is not involved. The parent kills and reaps the helper on normal shutdown or failure, with a 30-second initialization timeout and a ten-second response timeout per decode. This avoids the unsafe closure ownership in whisper-rs 0.16.0's cancellation adapter. Capture overload, malformed input, oversized output, and inference failures stop the session rather than silently discarding speech or changing to a hosted service. Desktop window close, renderer reload, and normal application exit independently cancel a decode; the supervisor checks cancellation every 20 ms. Forced parent termination still requires platform process-lifetime containment before production enablement.
Recorded Linux smoke evidence¶
On 2026-10-03, the portable CPU release probe transcribed the public 11-second JFK
fixture in a container with --network none, using the pinned tiny English model
and no API key. It emitted six revisions and two settled segments, with transcript
text absent from default output. Initialization took 503 ms; individual decodes
took 1,404–2,149 ms on a five-vCPU Intel Xeon Platinum 8573C cloud host while other
repository validation was running. A separate less-contended replay took
1,020–1,469 ms per decode. Process RSS was unavailable in this sandbox, so no memory
claim is made.
This establishes offline execution and the transcript protocol, not microphone acceptance or supported-device performance. Both runs exceed the provisional one-second decode budget below. macOS and Windows capture, accuracy, memory, accelerator behavior, packaging, and long-session stability remain unmeasured.
Collect supported-device evidence¶
Record OS/version, CPU, RAM, microphone and sample rate, requested/observed backend, model digest, executable size, installed model size, cold-start time, p50/p95 decode time, visible partial/final latency, parent-plus-helper peak RSS, CPU/GPU use, and technical-term error rate. Include quiet speech, accents, silence/background noise, words spanning segment boundaries, a 30-minute session, and at least ten successive draft boundaries. Fixed non-overlapping windows can cut words; byte continuity tests do not prove recognition accuracy or word-level deduplication.
Use these provisional review thresholds, then explicitly agree the supported hardware floor: cold start at most five seconds, p95 decode at most one second for six seconds of audio, first visible partial at most three seconds, combined peak RSS at most 512 MiB, no capture overruns or growing memory during 30 minutes, and at most 10% word error on a reviewed technical dictation corpus. The fixed-window probe does not establish an utterance-end finalization target; production endpointing needs its own latency and accuracy evidence.
On both platforms, verify first-use consent, denial and recovery, absent/unplugged devices, invalid models, inference timeout/crash, pause/resume, stopping during speech, normal exit, and forced exit. Confirm the OS microphone indicator clears and repeat with networking disconnected and no OpenAI API key. Synthetic queue tests cannot substitute for native permission and shutdown acceptance.
Run deterministic checks with:
cargo test --locked --package colossus-native-dictation --lib
cargo test --locked --package colossus-native-dictation --features probe --lib
cargo clippy --locked --package colossus-native-dictation --features probe --all-targets -- -D warnings
The root Rust gate runs the first suite without audio/model build dependencies. The feature-enabled suite additionally tests capture bounds, resampler tails, continuous send boundaries, and cancellation. Renderer tests cover ordered revisions, settled drafts, failed sends, bounded output, and stale-session replies. Production enablement remains blocked on macOS/Windows measurements, native permission acceptance, forced-exit containment, model delivery, silence/endpoint behavior, and release packaging. No supported-device or packaging decision has been established by a Linux replay alone.