AKOÚŌ is a portable listening system for AI agents. It does not only ask what is inside a sound. It asks how an agent should listen, what kind of evidence is available, which claims are allowed, and which claims must remain unknown.
The public system contains 17 portable skills:
akouo-router: the meta-router that chooses listening modessignal-inspection-listening: technical signal and metadata earacoulogical-object-listening: perceptual object/auditum earmusical-aesthetic-listening: musical, aesthetic, and sound-design earembodied-affective-listening: body, pressure, affect, and attention eartransductive-media-listening: mediation, sensors, codecs, platforms, and models earforensic-archival-listening: evidence, archive, testimony, damage, and custody earecological-posthuman-listening: habitat, nonhuman, weather, field, and more-than-human earcritical-political-listening: power, platform, labor, race, class, gender, coloniality, access, and infrastructure earsymbolic-fictional-listening: declared fiction, myth, ritual, worldbuilding, and speculative earaudiovisual-scenic-listening: sound-image-text-scene, synchronization, captions, UI sound, and audiovisual phrasing earvoice-speech-listening: voice, speech, transcript, ASR, TTS, voice-agent, identity caution, and consent earaccessibility-normative-listening: hearing norms, captions, transcripts, haptics, sensory variation, device, fatigue, and access earmaterial-event-listening: vibration, resonance, duration, flux, material support, propagation, and event earmemory-lineage-listening: sound-memory ear for stored records, recurrence, kinship, lineage, and change over timesovereign-listening: covenant-aware ear for consent, withholding, opacity, retention, precision, and refusalreference-layer: the conceptual mapping skill that turns listening into concepts, methods, traditions, research routes, cautions, and adjacent modes
The private benchmark extension can add benchmark-listening for external-agent evaluation and database ingestion. That benchmark skill is not part of the public portable release.
Every listening output must separate claims into six categories:
heard: an embodied listener's attributable report of what was directly present to a declared perceptual aperture; machine outputs and supplied text are not hearings of the represented soundmeasured: produced by technical inspection, metadata, waveform, spectrogram, or file analysisinferred: plausible logical or technical deductions onlyinterpreted: cultural, theoretical, affective, aesthetic, ecological, archival, political, or contextual readingsspeculative: declared fictional, symbolic, mythic, imaginative, or possible-world readingsundetermined: what cannot be responsibly claimed from the available evidence
This taxonomy is the main safety system. It prevents agents from turning a prompt into audio evidence, a metaphor into a fact, or a cultural reading into a measurement.
Current producers also emit listening_context. It keeps four questions
separate: what the covenant permits, where the listener is positioned, what
the apparatus can sense, and what the evidence supports. Its apertures,
scales, sources, participants, authority, revision, and honest-absence fields
make those conditions machine-checkable. See ACCOUNTABLE_LISTENING.md.
Use this process for most sound tasks:
- Identify the sonic object.
- Identify the input type: audio file, prompt, transcript, field note, archive note, dataset description, spectrogram, waveform, video, metadata, model output, mixed, unknown, or other.
- Run
akouo-routerunless the user explicitly requests one mode. - Choose primary, secondary, and corrective modes.
- Run each selected listening mode with the shared JSON schema.
- Merge claims into the shared taxonomy.
- Write a synthesis that preserves differences between modes.
- Recommend the next mode or command.
Every command begins with a router planning pass, even when its mode chain is fixed. The planning pass supplies the evidence inventory, risks, and forbidden assumptions that the command's synthesis must respect, so akouo-router appears in skills_called for every command output except /one-sound-many-ears, whose comparative contract runs all modes unconditionally.
Host apps should consume AKOÚŌ as data:
akouo.manifest.jsoncarries the skill list with structured metadata (facets, cost tier, memory policy, corrective eligibility), the command chains, the Evidence Ladder, and command permission overrides. Validate withschemas/manifest.schema.json.presets/presets.jsoncarries named listening configurations for recurring use-cases; validate each entry withschemas/preset.schema.json. A preset names its command, mode chain, cost tier (light/standard/deep), memory policy, and perception passes; hosts map passes to their own backends.- Outputs pin their contract with
akouo_version, declare theirapparatus, declare the listener, link stored records throughmemory, and attribute claims to evidence, apertures, listening passes, temporal scales, alternatives, and actionability. - Current outputs add context v2 plus
listening_provenance,listening_passes, androute_decisions.ensembleis emitted only for an explicit plural-listening or ear-swarm declaration. Readers retain an explicit compatibility path for context v1. - Machine outputs, textual prompts, transcripts, field notes, and descriptions are attributable evidence, but are never placed in
heard; use measured, inferred, interpreted, or undetermined according to their actual basis. - Release validation resolves participant, listening-pass, route-decision, and influence references. Canonical
heardclaims must resolve to a human pass and human-report source; ordinary record links do not create listening influence or an ensemble.
Loading these files replaces hand-copied route tables, which drift. Prose in this guide explains the contract; the manifest is the source of truth.
AKOÚŌ is designed to be driven by other agents, apps, and frameworks. The consumption loop is:
- Route. Inject
skills/akouo-router/SKILL.mdand request a router output (schemas/router-output.schema.json) or, for autonomous handoff, an expanded routing plan (schemas/routing-plan.schema.json). The plan carriesevidence_level,claim_permissions,mode_chain,forbidden_assumptions, andstop_conditions. - Check stop conditions. If the plan says the needed evidence is unavailable, stop or gather evidence; do not run listening modes on imagined input.
- Listen. Inject only the
SKILL.mdfiles named in the mode chain, in role order (primary, secondary, corrective), each emittingschemas/listening-output.schema.json. Enforce the plan'sclaim_permissionson every output. - Map (optional). Inject
skills/reference-layer/SKILL.mdwhen the workflow needs concepts, methods, traditions, and research routes (schemas/reference-map.schema.json). - Merge. Wrap the run in
schemas/command-output.schema.json(orschemas/comparative-listening-output.schema.jsonfor/one-sound-many-ears), preserving each mode's claims and disagreements. Command outputs may carry the expanded plan in the optionalrouting_planfield; the reference app does this for/routeand/method. - Hand off. Pass
recommended_next_mode,recommended_command, remainingundeterminedclaims, and unmet stop conditions to the next agent or turn.
Each skill folder is self-contained: SKILL.md plus the references/ schemas are everything an external agent needs for that step. No step requires a specific model provider, and every step can be validated against the canonical schemas in schemas/.
Default routed pass. Use it when the user wants a responsible first analysis and has not specified a method.
Typical chain: router, primary mode, secondary mode, corrective mode.
Broad multimodal scan. Use it for deep analysis, early research, artwork review, sound design, or unknown sounds.
Typical chain: router, signal, acoulogical, musical/aesthetic when relevant, embodied, transductive, contextual mode, optional critical-political corrective.
Research-oriented command. Use it for essays, field notes, sonic methodology, artistic research, theory, or study planning.
Typical chain: acoulogical grounding, musical/aesthetic route when relevant, critical-political caution, ecological or symbolic route, reference mapping.
Technical inspection. Use it for metadata, waveform, spectrogram, clipping, loudness, noise, compression, repair planning, AI artifacts, or media-chain questions.
Typical chain: signal inspection, transductive media.
Conceptual mapping. Use it to identify methods, concepts, traditions, research questions, cautions, and adjacent modes. It should not become a bibliography dump.
Critique of simplistic sound-versus-vision claims. Use it when a text or project says sound is more embodied, authentic, immersive, or present than images.
Typical chain: critical-political, transductive-media, acoulogical-object, with musical/aesthetic grounding when relevant.
Declared speculative sonic worldbuilding. Use it for ritual, myth, games, film worlds, dreams, alien voices, hauntology, and sonic fiction.
Typical chain: symbolic-fictional, embodied-affective, ecological or critical corrective when needed.
Strict evidentiary pass. Use it for testimony, archives, damaged recordings, protest recordings, surveillance, legal stakes, or recordings of harm.
Typical chain: signal inspection, forensic-archival, critical-political.
Mediation-chain mapping. Use it for sensors, sonification, microphones, codecs, platforms, ASR, neural codecs, AI audio, and model outputs.
Typical chain: transductive-media, signal inspection, critical-political when stakes require it.
Voice and speech pass. Use it for spoken audio, transcripts, captions, podcasts, ASR, TTS, voice agents, cloning, identity caution, consent, and intelligibility.
Typical chain: voice-speech, transductive-media, accessibility-normative, critical-political when stakes require it.
Sound-image-scene pass. Use it for video, film, games, installation, UI sound, subtitles, captions, synchronization, offscreen sound, diegesis, and audiovisual phrasing.
Typical chain: audiovisual-scenic, acoulogical-object, voice-speech when speech or captions matter, critical-political when audiovisual assumptions need critique.
Accessibility and hearing-norm audit. Use it for captions, transcripts, haptics, sonic alerts, voice interfaces, assistive paths, sensory variation, fatigue, masking, and implied listener assumptions.
Typical chain: accessibility-normative, voice-speech when speech matters, embodied-affective when loudness or fatigue matters, critical-political.
Field recording and situated listening route. Use it for soundscape, acoustemology, aurality, listening-with, field notes, habitat, infrastructure, sensors, and fieldwork ethics.
Typical chain: ecological-posthuman, transductive-media, critical-political, material-event when resonance, vibration, or duration matter.
Sonic methodology and agent-handoff command. Use it for research design, artistic research, listening practice, workflow planning, and routing AKOÚŌ into other apps or agents.
Typical chain: router, acoulogical-object, critical-political, accessibility-normative, reference-layer.
Router-only handoff plan. Use it when another app, agent, benchmark runner, or framework needs a compact mode chain, evidence inventory, risks, and forbidden assumptions before doing the work.
Typical chain: akouo-router only.
Memory route. Use it to situate a sound in its lineage against a sound-memory store (akousma/akousmata-style records) and register the listening into the store.
Typical chain: router, memory-lineage, acoulogical grounding, signal-inspection corrective. Stop rather than write when no store is available; use the read-only recall preset for comparison without registration.
Sovereignty route (introduced in v0.7). Use it to listen under an explicit listening covenant: verify the covenant's identity and lineage, apply what the host's gates can enforce (refuse sources, ignore classes, withhold aspects, coarsen precision, retain nothing, honor quiet hours), carry every non-executable line as a commitment, and report withholding as honest, attributed absence — counted and named by rule, never described, and never confused with undetermined.
Typical chain: router, sovereign-listening, acoulogical grounding on what the covenant admits, signal-inspection corrective. With no covenant available, report exactly that and stop: sovereignty is opted into, never imposed.
Corpus-lineage route. Use it for training, fine-tuning, retrieval, annotation, filtering, provider disclosure, licensing, labor, consent, jurisdiction, and opt-out questions. Provider identity and model behavior never substitute for a verified corpus ledger; unknown inheritance stays unknown.
Typical chain: router, corpus-listening, critical-political corrective, and transductive-media mapping.
Comparative flagship command. Runs one sonic object through all fifteen public non-sovereign listening modes and compares contradictions, productive tensions, limits, and next steps. Parallel outputs remain plural listening; they are not an ear swarm unless later passes declare attributable influence.
For an unknown audio file, start with /listen or /full-ear.
For music, rhythm, harmony, pitch, timbre, production aesthetics, sound design, or creative usefulness, include musical-aesthetic-listening.
For waveform, spectrogram, loudness, clipping, codec, frequency, or repair questions, include signal-inspection-listening and /tech.
For texture, source ambiguity, sound objects, Foley, acousmatic sound, or morphology, include acoulogical-object-listening.
For bass, pressure, fatigue, pleasure, dread, alarm, dance, immersion, ASMR, or bodily force, include embodied-affective-listening.
For microphones, datasets, ASR, voice cloning, neural codecs, sonification, compression, and platforms, include transductive-media-listening.
For training data, fine-tuning, retrieval corpora, annotation, provider
disclosure, licensing, or unknown model inheritance, include
corpus-listening and /corpus.
For evidence, archives, testimony, damage, event reconstruction, or harm contexts, include forensic-archival-listening.
For field recordings, animals, weather, hydrophones, habitats, soundwalks, and more-than-human relations, include ecological-posthuman-listening.
For platform power, labor, race, gender, class, coloniality, surveillance, policing, accessibility, or extraction, include critical-political-listening.
For dreams, ritual, myths, fictional worlds, game audio, alien voices, and speculative scenes, include symbolic-fictional-listening.
For video, film, games, captions, subtitles, screen interfaces, sound-image synchronization, offscreen sound, or audiovisual phrasing, include audiovisual-scenic-listening.
For voice, speech, transcript, ASR, TTS, voice cloning, voice agents, prosody, intelligibility, or consent questions, include voice-speech-listening.
For captions, transcripts, haptics, deaf or hard-of-hearing access, sensory variation, masking, fatigue, alerts, or implied listener assumptions, include accessibility-normative-listening.
For vibration, resonance, propagation, duration, feedback, rumble, low frequencies, installation sound, materials, or processual sonic events, include material-event-listening.
For stored sound-memories, recurrence, lineage, series over time, archive comparison, or registering a listening into a store, include memory-lineage-listening and /remember.
- Run
/listen. - Read router output first.
- Check whether the selected corrective mode is strong enough.
- Inspect
undeterminedclaims before trusting the synthesis. - Continue with
/full-ear,/tech,/study, or a single mode if needed.
- Run
/listenor directly usemusical-aesthetic-listening. - Separate rhythm, pitch, timbre, texture, form, and production observations.
- Mark genre, tradition, instrument, culture, and scene as undetermined unless evidence exists.
- Add
signal-inspection-listeningfor measured tempo, tuning, dynamics, or spectral claims. - Add
critical-political-listeningif genre, cultural framing, platform, labor, or extraction matters. - End with sound-design utility: what the sound can do creatively without pretending to know its origin.
- Run
/study. - Use the acoulogical output as perceptual grounding.
- Use musical/aesthetic output when musical organization matters.
- Use ecological, political, archival, or fictional modes depending on the research question.
- Use
/referenceto turn the listening result into methods, concepts, questions, and cautions. - Keep citations and theory subordinate to the claim taxonomy.
- Run
/forensic. - Treat the sound as trace, not as proof.
- Separate measured file properties from event claims.
- Identify missing provenance, custody, editing chain, and corroboration.
- Use
critical-political-listeningto name institutional, surveillance, archive, or harm risks. - Keep speculative claims empty unless the user explicitly asks for declared speculation.
- Run
/one-sound-many-ears. - Compare what each mode reveals and hides.
- Look for contradictions between evidence, perception, music, affect, mediation, archive, ecology, politics, audiovisual scene, voice, access, material event, and fiction.
- Use the most responsible reading as a map, not a final truth.
- Choose the next mode based on the user's actual goal.
- Run
/routewhen another app or agent only needs a plan. - Preserve
available_evidence,unavailable_evidence,risks, andmust_not_assume. - Pass the primary, secondary, and corrective modes to the receiving framework.
- Use
/methodwhen the receiving system also needs research questions, access requirements, reference routes, and stop conditions. - Stop rather than analyze when the needed evidence is unavailable.
The optional private benchmark extension lets different agents run the same listening tasks and ingest their reports into a local benchmark database.
Typical benchmark process:
- Select a benchmark suite and case.
- Provide blind object names and blind prompts when testing audio understanding.
- Run
benchmark-listeningas an orchestration skill. - Run
akouo-routerand the selected listening modes, includingmusical-aesthetic-listeningfor music/aesthetic cases. - Produce a Markdown report with an embedded canonical JSON block.
- Ingest the report into the benchmark API.
- Compare models by claims, scores, flags, latency, suite, case, provider, and agent ID.
Benchmark data usually separates:
- model metadata: provider, model ID, modality, temperature, token limit, audio mode
- agent metadata: agent ID, agent type, notes
- sound object metadata: object label, input type, prompt, file metadata, tags
- run result JSON: router output, mode outputs, synthesis, benchmark metadata
- normalized claims: category, statement, confidence, basis, listening mode, source section
- scores: schema validity, claim discipline, evidence grounding, mode fidelity, audio specificity, musical specificity, aesthetic discipline, cultural caution, sound-design utility, poetic usefulness, and suite-specific axes
- flags: hallucination risk, source overreach, weak uncertainty, schema drift, and claim inflation
The benchmark should preserve the canonical JSON. Markdown is the human-readable wrapper; the JSON block is the ingestion source of truth.
The reference store implementation is the akousmata listening navigator
(github.com/sonicfieldlabs/akousmata): a local-first library app over the
shared store (earworm protocol) with filtering, tagging, manual human
memories, graph navigation, a wiki layer, and research sessions. Hosts that
implement /remember against an akousmata store should keep their record
cards compatible with it. Human listening events written there carry
listener.type: "human" — the same output contract this guide defines, with
a person at the apparatus.
- Run
/remember(or therecallpreset for read-only comparison). - Confirm store availability in the routing plan's evidence inventory; stop rather than invent records.
- Keep the fresh perceptual pass before memory comparison.
- Mark memory-derived claims with
source: "memory"; they never enterheardormeasuredfor the present sound. - Populate the
memoryblock (akousma_id,akousmata_refs,lineage_note) on outputs the host will store. - Treat disagreement between memory and present listening as a finding.
- Never claim that an agent heard binary audio if only text was supplied.
- Never treat a file name as ground truth in blind tests.
- Never convert genre, culture, tradition, ethnicity, or geography into a confident claim without evidence.
- Never let poetic or symbolic interpretation escape
interpretedorspeculative. - Always include meaningful
undeterminedclaims when evidence is missing. - Always keep mode outputs distinct before synthesis.
- Declare the apparatus when known; derive forbidden claims from the declared substrate (mono input forbids stereo claims, model-only perception forbids
measured). - Never invent stored records, identifiers, or lineage; absence from a store is not novelty in the world.