Turn audio into a searchable library with transcription, metadata enrichment, tagging, speaker identification, and artwork generation.
Find a file
Naomi Persephone Amethyst fce7f78531
Some checks failed
Test and publish / test (push) Has been cancelled
Test and publish / portability (macos-latest) (push) Has been cancelled
Test and publish / portability (windows-latest) (push) Has been cancelled
Test and publish / binaries (amd64, darwin) (push) Has been cancelled
Test and publish / binaries (amd64, linux) (push) Has been cancelled
Test and publish / binaries (amd64, windows) (push) Has been cancelled
Test and publish / binaries (arm64, darwin) (push) Has been cancelled
Test and publish / binaries (arm64, linux) (push) Has been cancelled
Test and publish / binaries (arm64, windows) (push) Has been cancelled
Test and publish / containers (ffmpeg) (push) Has been cancelled
Test and publish / containers (scratch) (push) Has been cancelled
Test and publish / release (push) Has been cancelled
Rehearse a written script through the passes a recording would meet
A script is the thing before the recording, and the questions worth asking of it
are the ones the pipeline asks afterwards: what is in this, what would it be
tagged, what would a listener be warned about. Until now the only way to find
out was to record it, import it, and read the entry -- and then live with a
transcript, an analysis and a review in the store for something that was never
in the library.

`inductor script draft.txt` runs the first pass and the review over written text
and reports what came back. It writes nothing: not the analysis store, not an
entry, not the registry, not a proposed tag. A test walks the whole tree before
and after and compares every path and size, because a rehearsal that left
anything behind would be an import.

The prompts are the library's own, unchanged, and so are the registry, the
models, the temperatures and the retry behaviour -- `.15` through `ChatJSON` for
the first pass, `.2` through `Chat` and `ExtractJSON` for the review, exactly as
the recorded path calls them. A rehearsal whose prompts had been tidied for the
occasion would answer a question nobody asked. It takes the same fork, too: under
two hundred characters there is not enough to be worth a first pass, so the
script is reviewed from itself, which is where `Analyze` already draws the line.

The text is turned into the shape a transcript arrives in and handed to
`Sentences`, so the splitting, the numbering and the `s001` ids a first pass is
told to cite all come from the recorded path's own machinery rather than from a
second implementation of it that would drift from the first.

What is deliberately withheld is the audio evidence. A script has no recording,
so `measured` and `heard` stay empty. The sentence timings exist to give
citations somewhere to point and to estimate a running time -- which the review
is told about, and which changes what it makes of repetition over forty minutes
against twelve -- but they are an assumption about pace, and `MeasuredBlock`
offers its figures to a model as facts about a recording. Filling one from the
other would be inventing evidence, which is the failure most of this pipeline's
guards exist to prevent. `--wpm` moves the assumption; nothing promotes it.

What is reported is each call with its model, its duration and its parsed
result, the usage and cost behind it, and `as_imported`: what `emit` would
actually put on the entry. That last part is the reason to report more than the
raw JSON. A review returns tag *names*; what reaches an entry is whatever the
registry and the creator's own map say those names mean, and a name that resolves
to nothing does not become a tag, it becomes a proposal waiting on an
adjudicator. So the tags are resolved through both, and the ones that would not
land are reported separately as `would_propose` -- the answer to "would this
script widen the vocabulary, and by what". The spoilers go through
`SpoilersFrom`, the synopsis through `Block`, and a `replace` verdict through
`ReplacementTitle`, so a retitle shows as a retitle rather than as a verdict to
be interpreted.

----

The first live run looked like it had hung, and the reason is worth recording
because it is not specific to this command.

`startRunWork` returns a no-op when no progress reporter is in the context, and
only `run` installs one. So a one-shot command prints the line that says what it
is about to do and then nothing at all -- while `Chat` retries five times behind
a fifteen-minute HTTP timeout, inside a `ChatJSON` that retries three times more.
A model thinking for four minutes and a model that will never answer look exactly
alike from outside, and the first is ordinary. Unbounded, one script could sit
for over an hour in silence.

So each stage now prints the prompt size and the token ceiling before it
dispatches -- the one property of a request that explains a slow answer and is
knowable before it arrives, and on a large library it is mostly the tag registry
rather than the script -- and says it is still waiting every fifteen seconds,
the interval a run's own heartbeat uses. A stage gives up after ten minutes.
That is far less patience than the pipeline has, and deliberately: a run is
thousands of recordings left going overnight, where a review worth ten minutes is
worth waiting for, and this is one file with somebody watching it. `--timeout`
moves the bound and `--timeout 0` restores the pipeline's own, which the options
struct spells as a negative duration because its zero already means "nobody
said". A deadline that fires says how to wait longer rather than reporting
`context deadline exceeded`, which names the mechanism and not the thing that
happened.

Tests cover the transcript shaping and its timings, the absent audio blocks in
the review prompt, both calls reporting with their usage, tag resolution
splitting into what would land and what would only be proposed, the solo fork
firing at the threshold, an unparseable review keeping its text alongside the
first pass that did succeed, the parsing seam, and a server that never answers
returning within its deadline.
2026-09-23 02:03:28 +00:00
.github/workflows Add cross-platform binaries and multiarch container publishing 2026-09-16 01:51:42 +00:00
cmd/inductor Initial release of Inductor 2026-09-15 06:19:35 +00:00
docs Rehearse a written script through the passes a recording would meet 2026-09-23 02:03:28 +00:00
examples/library Initial release of Inductor 2026-09-15 06:19:35 +00:00
internal/inductor Rehearse a written script through the passes a recording would meet 2026-09-23 02:03:28 +00:00
scripts Add cross-platform binaries and multiarch container publishing 2026-09-16 01:51:42 +00:00
.dockerignore Add cross-platform binaries and multiarch container publishing 2026-09-16 01:51:42 +00:00
.gitignore Add cross-platform binaries and multiarch container publishing 2026-09-16 01:51:42 +00:00
CONTRIBUTING.md Add cross-platform binaries and multiarch container publishing 2026-09-16 01:51:42 +00:00
Dockerfile Add cross-platform binaries and multiarch container publishing 2026-09-16 01:51:42 +00:00
go.mod Add cross-platform binaries and multiarch container publishing 2026-09-16 01:51:42 +00:00
go.sum Initial release of Inductor 2026-09-15 06:19:35 +00:00
LICENSE Initial release of Inductor 2026-09-15 06:19:35 +00:00
Makefile Initial release of Inductor 2026-09-15 06:19:35 +00:00
NOTICE Initial release of Inductor 2026-09-15 06:19:35 +00:00
README.md Rehearse a written script through the passes a recording would meet 2026-09-23 02:03:28 +00:00

Inductor

Inductor turns audio and source metadata into a Hypnotica content library. It transcribes recordings, measures audio, reviews transcript evidence, manages a controlled tag vocabulary, and writes portable YAML entries and artwork.

The application is written in Go. A Python worker runs transcription and speaker embeddings locally or over SSH. Existing inductor.yaml configuration, command names and flags, hypnotica/v1 documents, fingerprints, and stored enrichment results remain usable.

Where this sits

Hypnotica is the other half: it turns the same content tree into a static website with feeds, an offline player and a searchable catalogue. The two never meet — Inductor writes hypnotica/v1 YAML and knows nothing about websites, Hypnotica reads it and never transcribes anything — so either one can be replaced without touching the library between them.

Driftspace is the quickest way to start: a ready-made library directory with the layout, a tag vocabulary, a worked example and the working notes an agent needs to fill it. Clone it, point an agent at your audio, and you are past configuration.

git clone https://github.com/NaomiAmethyst/driftspace-template my-library

Build

Requires Go 1.24 or later. CI builds Linux, Windows, and macOS binaries for AMD64 and ARM64. FFmpeg and ffprobe are needed for audio inspection, conversion, and measurements. SSH and SCP are needed for a remote worker. Python 3.10 or later is needed on the worker machine.

go build -o bin/inductor ./cmd/inductor
bin/inductor --help

No Python installation is needed to run metadata, tagging, cache, or document maintenance commands. Worker scripts and model prompts are embedded in the Go binary.

Downloads and containers

Every successful push or pull request build provides six binary archives under GitHub Actions, retained for 30 days. Each archive includes license notices and comes with a SHA-256 checksum file. Windows downloads are ZIP files; Linux and macOS downloads are .tar.gz files. Pushing a v* tag also publishes the archives and checksums.txt on GitHub Releases.

Images are published to ghcr.io/naomiamethyst/inductor on successful pushes. Both variants support Linux AMD64 and ARM64:

Tag on the default branch Contents
latest, ffmpeg Inductor, FFmpeg/ffprobe, CA certificates, SSH/SCP
scratch Inductor and CA certificates, built from scratch

Each variant also gets <branch>-ffmpeg / <branch>-scratch, sha-<full-commit>-ffmpeg / sha-<full-commit>-scratch, and, for tag pushes, <tag>-ffmpeg / <tag>-scratch (for example, v1.2.3-scratch). Pull requests and manual workflow runs build images without publishing them.

docker run --rm -v "$PWD:/library" ghcr.io/naomiamethyst/inductor:ffmpeg check
docker run --rm -v "$PWD:/library" ghcr.io/naomiamethyst/inductor:scratch --help

The scratch image supports commands that need only the Go binary, including metadata maintenance and HTTPS API calls. Media processing needs the FFmpeg variant. Neither image includes Python or inference models: use a remote worker with the FFmpeg variant for transcription and voice embeddings. Native Windows builds also need a remote Unix worker; the local worker uses a POSIX shell and Unix file locking. Standalone binaries need FFmpeg/ffprobe installed separately for media operations.

To build containers locally, the default target includes FFmpeg:

docker build --target runtime-ffmpeg -t inductor:ffmpeg .
docker build --target runtime-scratch -t inductor:scratch .

Start a library

Copy examples/library into your library directory and edit inductor.yaml and sources/example.yaml. Each source needs audio, title, and author. Supply local audio paths; URL-only source records are reported but do not download audio. The example tag registry is deliberately small: extend it to fit your library.

bin/inductor -r /path/to/library check
bin/inductor -r /path/to/library run --dry-run
bin/inductor -r /path/to/library ingest --stage media --stage emit --no-covers

For transcription, configure transcribe.remote with an SSH host, or leave it empty to use a local CPU worker. The first worker start creates its virtualenv, installs dependencies, and copies the embedded scripts. Models download on first use. Set OPENROUTER_API_KEY for analysis and review. Configure enrich.comfy_url and the appropriate ComfyUI models for artwork.

bin/inductor -r /path/to/library run --no-covers --no-pages

This processes missing artifacts and writes entries. ingest also supports individual stages: media, transcribe, analyse, review, and emit. Use ingest --redo to include finished items; run checks missing artifacts even on finished entries, and run --redo ARTIFACT forces the named artifact to be rebuilt. --overwrite permits replacing existing summaries and spoilers. Creator descriptions, existing IDs, and extension fields are preserved.

run also checks sources and mapping targets, fills missing author tag mappings before processing, then generates and applies adjudication rulings with writes enabled before generating author pages. Adjudication includes stored reviewer-approved proposals, and backfill applies approved, renamed, or merged tags to recordings whose reviews requested them. Rejected and unreviewed proposals are not backfilled by run.

After processing, run audits transcript coverage, writes acoustic measurements and missing durations onto entries, repairs cover prompts and missing, invalid, or stale generated covers, and updates similar-voice relationships on author pages. It reports voiceprint verification, guest-speaker appearances (cameos), duplicates, and orphans. Transcript auditing reports recordings longer than two minutes with less than 90% coverage; it does not automatically retranscribe them. Manual covers are preserved. Duplicates and orphans are reported without merging or deleting anything. These steps still run when no recordings need processing.

Each additional step can be omitted:

Flag Effect
--no-check Skip the preflight report; bypassing source validation still requires --force.
--no-tagmaps Skip filling missing mappings and reconciling their item tags.
--no-adjudicate Skip generating and applying rulings.
--no-adjudicate-apply Generate and save rulings, but do not apply them.
--no-adjudicate-write Generate and save rulings, and preview application without writing registry, mapping, or item changes.
--no-review-proposals Omit stored review proposals from adjudication; item and tagmap proposals remain included.
--no-backfill Skip applying registered/adjudicated tags requested by reviews.
--no-acoustic-apply Skip writing acoustic metadata and missing durations.
--no-measured-tags Skip tagging recordings from their own measurements.
--no-transcribe-audit Skip the transcript coverage report.
--no-artwork-repair Skip library-wide cover prompt translation and cover repairs.
--no-cover-prompts Skip prompt translation while retaining cover repair from existing prompts.
--no-similar Skip writing similar-voice relationships on author pages.
--no-cameos Skip guest-speaker detection independently of voiceprint verification.
--no-voiceprint-verify Skip the final attribution report; voiceprint generation remains part of the recording graph.
--no-duplicates Skip the duplicate report.
--no-orphans Skip the orphan report.

Unapplied rulings are retained in the cache's rulings.yaml for a later run. --dry-run reports existing checks, missing mappings, artifact work, and pending adjudication without API calls or applying changes. It previews backfill, acoustic application, similar voices, and artwork repairs too. The adjudication apply/write opt-outs also prevent backfill writes. Existing --no-pages controls author page generation; --no-covers also skips cover prompt translation and artwork repairs. The configured enrich.covers: false disables these artwork repairs as well.

Tagmaps honor --author; with --only or --limit, they cover the selected recordings' authors. Unrestricted runs also fill mappings for finished authors. Adjudication, backfill, acoustic application, transcript auditing, artwork repair, similar voices, and the final reports cover the whole library. The recording graph honors selection flags, and --limit applies after checking for outstanding artifacts so complete entries do not prevent later missing work from being found.

During a run, a progress line prints every minute with elapsed time, active work, and artifact counts (cached, completed, running, pending, failed, blocked, and skipped). This continues through checks, tagmaps, adjudication, and author pages. Use --no-progress to disable these periodic lines; ordinary status messages, errors, and the final report remain visible. Dry runs omit the periodic lines. Use --verbose to log each dispatch and completion, including recording artifacts, review batches, tagmap requests, and author work. Failures and skipped results are identified explicitly. Both flags can be combined:

bin/inductor -r /path/to/library run --verbose --no-progress

Commands

Area Commands
Import and processing check, add, ingest, run, transcribe
Tags and decisions retag, fold, tagmap, adjudicate, registry, backfill, reconsider
Metadata and artwork retitle, authors, cover-prompts, artwork, attribute
Audio and speakers acoustic, acoustic-apply, measured, voiceprint, similar
Library maintenance paths, orphans, migrate, export, duplicates
Model evaluation compare, script

paths is what makes a library portable: it spells every asset reference and every placed symlink relative wherever it points inside the tree, and absolute wherever it points out of it. Run it after moving or cloning a library — a reference that has gone stale is reported by check, but a symlink that has gone stale is a cover that silently stops existing, and the only sign of it is a site build reporting artwork it cannot find.

registry is the one that edits the vocabulary directly, for the decisions a model should not be making: add, describe, remove, rename, merge, and bulk for a file of those applied in order. Each keeps the rest of the library in step — the items carrying the tag, the tags_added a run recorded, the proposals still waiting on it, and every creator map that points at it — and records a ruling saying what was decided, so a later adjudicate does not re-open it.

inductor registry add "Humour" --description "Comedy is part of the intent." --write
inductor registry merge "Toy" "Toys" --write
inductor registry bulk changes.yaml --write

measured is the exception to all of that. Every other tag in a library is somebody's claim — the creator's, in the source record, or a review model's, from reading a transcript — and both go through the registry and the adjudicator. These come from the instruments instead: whether there is a binaural beat in the file, at what rate, whether it holds still, whether anybody speaks, and how. There is no opinion to rule on, so they skip the model and the adjudicator, and measured --registry provisions the definitions they need.

Each rule declines rather than guesses when its guard fails: a recording whose dominant partial wanders has no carrier to report, and one the tagger never ran over cannot be called wordless. They are stamped with the ruleset that produced them, so a threshold that moves can find the tags it invalidated. And where a measurement contradicts a tag already on the entry, nothing is overwritten — the disagreement is reported, counted per creator, because a creator who labels tone tracks "binaural" when no beat is measurable is telling you something about the rest of their metadata.

inductor measured --dry-run     # what it would tag, and which claims it disputes
inductor measured --registry    # provision the vocabulary the rules need

script runs a written script through the passes a recording of it would meet and reports what came back, writing nothing. The prompts, the registry, the models and the temperatures are the library's own, so what it shows is what an import would produce -- including which of the review's tags resolve against the vocabulary and which would only become proposals. What it does not supply is the audio evidence: a script has no recording, so the measured and heard blocks stay empty rather than being filled from an assumed pace.

inductor script draft.txt --author somecreator --title "The Anchor Word"
inductor script draft.txt --no-review        # stop after the first pass
inductor script draft.txt --wpm 70           # a slower delivery, and a longer running time
inductor script draft.txt --timeout 0        # wait as long as a run would

Each stage prints the prompt size before it dispatches and says it is still waiting every fifteen seconds, because run is the only thing that installs a progress reporter and a model that thinks for four minutes otherwise looks exactly like one that has hung. A stage gives up after ten minutes by default; --timeout changes that, and --timeout 0 restores the pipeline's own patience, which is five retries behind a fifteen-minute HTTP timeout.

Every command accepts --help. Maintenance commands retain their original --write or --dry-run behavior; read their help before use. Output includes progress messages and a JSON summary. Exit status is 0 on success, 1 on an operation failure, and 2 for invalid arguments.

Library layout

inductor.yaml
sources/                         original metadata
content/tags.yaml                vocabulary and definitions
content/<creator>/_author.yaml
content/<creator>/<title>.yaml
content/<creator>/<title>.transcript.yaml
media/                           linked, copied, or converted audio
media/cover/                     artwork
state/transcripts/               reusable audio-keyed transcripts
state/enrichment/                reusable transcript-keyed analysis and review
state/decisions/                 durable tag decisions
.inductor/                       disposable indexes, measurements, queues, journals

Legacy caches in content/.hypnotica remain readable. Keep state/ with the library: it contains expensive results and editorial decisions.

See configuration, contracts and architecture, and development for details.

Testing

make test
make vet
make test-worker

Tests include thousands of comparisons against the frozen Python implementation, numerical audio fixtures, filesystem workflows, HTTP protocol tests, and race checks. Tests make no external API requests and do not download inference models. GPU inference and live provider integrations require a configured environment; the offline suite does not establish model quality or provider availability.

License

Copyright © 2026 Naomi Persephone Amethyst naomi@amethyst.name.

Inductor is licensed under GNU GPL version 3 only. See LICENSE, NOTICE, and third-party notices.