Skip to content

Repository files navigation

akouo

AKOÚŌ

Operational ears for the agentic era

AKOÚŌ is a multimodal listening system and prompt library for agentic workflows. It serves as a portable skill layer for AI agents (like OpenClaw, Hermes, Claude, Gemini, or OpenCode), allowing them to perform accountable, structured sonic inquiry.

Most audio AI systems ask what is inside an audio file. AKOÚŌ asks how an agent should listen, what each listening mode reveals, what it hides, and what must remain unknown.

AKOÚŌ does not pretend that agents hear like humans. It gives them accountable, switchable, operational ears.

This public release contains the portable AKOÚŌ skills, router, command definitions, schemas, and a local-first reference app for running the listening workflows.

Official public repository: https://github.com/sonicfieldlabs/akouo. Current release contract: v0.9.

The akouo-contract Python distribution packages this repository's canonical skills, commands, presets, schemas, and manifest for Oída and other local agent hosts. Package release 0.9.2 implements the akouo/v0.9 data contract; no skill fork is created inside a host application.

Version Status

  • Package 0.9.2 adds executable cross-reference checks and a canonical linked human/agent example without changing the v0.9 contract.
  • v0.9 distinguishes listening provenance from source provenance, represents several attributable listenings across time, makes route decisions and coded silence addressable, defines plural listening versus an ear swarm, and adds corpus-listening with /corpus.
  • v0.8 adds an accountable listening context: position, apertures, auditory scale, sources of listening, participants, action authority, revision, and honest absence are first-class data rather than prose conventions.
  • v0.7 adds sovereign-listening, the /covenant command, and an enforceable listening-covenant schema for consent, withholding, retention, precision, and quiet-hour rules.
  • v0.1 marks the first public portable-skills release.
  • v0.2 refines the listening modes with stronger theoretical distinctions, clearer evidence boundaries, and more explicit public-repo hygiene.
  • v0.3 adds musical-aesthetic-listening as a public mode for music, rhythm, pitch, harmony, texture, sound-design utility, poetic usefulness, and genre/cultural caution.
  • v0.4 expands AKOÚŌ into a more portable listening router for agentic workflows with voice/speech, audiovisual/scene, accessibility/normativity, and material/event listening.
  • v0.5 consolidates agentic routing: reference-layer becomes a portable meta-skill, routing plans carry evidence levels and claim permissions, the app gains deeper browser-side signal estimates, and SYSTEM_GUIDE.md documents the integration contract.
  • v0.6 instruments the system for host apps: a machine-readable contract (akouo.manifest.json), portable listening presets (presets/), the new memory-lineage-listening mode and /remember command for sound-memory stores, apparatus/listener/memory declarations on outputs, per-claim source and time_range, budget-aware routing plans, and command-level claim-permission overrides.

Sonic Field Labs Stack

AKOÚŌ is the shared listening contract across the stack. The current release contains 16 listening modes, 18 skills, and 19 commands.

Project Current integration
OÍDA 0.9.2 Loads the akouo/v0.9 contract and returns addressable route decisions, passes, provenance, and optional local capabilities.
Earworm 0.6.1 Persists addressable auditums and pre-capture decisions as akousma/v1.5 records, including revision and forgetting receipts.
Akousmata 0.6.1 Validates and navigates those records while keeping plural listeners distinct and requiring influence before declaring an ear swarm.
Algophony 0.5.2 Uses the listening contract for generation supervision, evaluation, and research workflows.
GERM 0.3.3 Connects cultivation sessions to OÍDA and Akousmata for listening and lineage-aware workflows.
ORAM 0.4.1 ORAM audio can be passed to any AKOÚŌ host; ORAM does not embed the contract directly.

Core Idea

A sound is never only a source. It can be approached as:

  • a signal measured in time and frequency
  • a perceptual object or auditum
  • a bodily and affective event
  • a technical or computational trace
  • an archive, testimony, or damaged record
  • an ecological relation
  • a cultural and political mediation
  • a musical and aesthetic organization
  • a voice, speech, transcript, or voice-agent relation
  • an audiovisual scene across sound, image, text, interface, and timing
  • an accessibility problem shaped by hearing norms, captions, haptics, devices, and sensory variation
  • a material event of vibration, resonance, propagation, duration, and process
  • a computational inheritance shaped by corpora, selection, annotation, labor, permission, and undisclosed lineage
  • a symbolic, fictional, or speculative world

akoúō keeps these dimensions distinct through explicit listening modes and a strictly enforced JSON claim taxonomy.

Claim Taxonomy

Every output must distinguish its findings into the following epistemic categories:

  • heard: an embodied listener's attributable report of what was directly present to a declared perceptual aperture; machine output, prompt, transcript, field-note, and description text belong to other categories according to their actual basis
  • measured: produced by file, signal, waveform, spectrogram, or metadata inspection
  • inferred: plausible logical deductions (not theory or culture)
  • interpreted: cultural, theoretical, affective, or contextual reading
  • speculative: fictional, symbolic, imaginative, or possible-world reading
  • undetermined: what cannot be responsibly claimed

This taxonomy is the main public contract of the system, preventing LLMs from hallucinating certainty or confusing a theoretical reading for a forensic measurement.

v0.9 Listening Provenance, Several Times, and Coded Silence

The v0.9 contract lets a report answer not only what an ear concluded, but which listening happened, when, through which sources and cuts, under which route decision, and whether another listening actually changed it.

  • listening-pass.schema.json keeps each listening attributable across live, captured, archival, seasonal, prospective, and re-listening moments.
  • listening-provenance.schema.json records sources of listening, technical and classificatory cuts, corpus lineage, and optional source-of-voice disclosure. Unknown lineage stays unknown.
  • route-decision.schema.json makes proceed, pause, defer, abstain, refusal, withholding, forgetting, and non-action addressable outcomes. Coded silence is therefore a decision, not a missing response.
  • ensemble.schema.json distinguishes plural listening from an ear swarm. Listener count or parallel execution is insufficient: a swarm requires attributable influence, preserved permission, preserved disagreement, and a dissolution rule.
  • corpus-listening and /corpus audit training, fine-tuning, retrieval, annotation, filtering, licensing, labor, provider disclosure, and jurisdiction without guessing protected corpus membership.

The claim model also carries stable claim ids, evidence and aperture references, pass attribution, auditory scale, alternatives, and actionability. A claim still never grants its own authority. The canonical human/agent paired example shows a human heard claim resolving to its participant, pass, aperture, and human-report source while an ordinary response_to record link remains neither pass influence nor an ensemble declaration.

v0.8 Accountable Listening

The v0.8 contract makes the conditions of a hearing inspectable. A current producer emits listening_context beside its six claim categories:

  • position states the listener's relation, situation, identity reference, and limitations without pretending to be a view from nowhere.
  • apertures inventory the channels that were available, degraded, unavailable, or withheld. A model observation is never silently promoted to a signal measurement.
  • auditory_scales and sources_of_listening state the temporal or infrastructural scale and evidence streams actually used.
  • participants keeps human, agent, hybrid, community, institutional, sensor, habitat, other-animal, ensemble, and other listenings attributable.
  • action_authority separates what a system can technically do from what it is authorized to do. Reference listeners default to observe_only.
  • honest_absences name unavailable, withheld, refused, unretained, and forgotten material without filling the gap. Epistemic uncertainty remains an undetermined claim.
  • revision makes re-listening additive and traceable instead of overwriting a prior account.

The four objects remain distinct: a covenant says what the listener may do; position says where and in what relation it listens; apparatus says what it can sense; and claims say what the available evidence supports. See ACCOUNTABLE_LISTENING.md for the producer and integration rules.

v0.7 Sovereign Listening

The v0.7 release adds the sovereignty layer — the Rights of the Audible made operational. It adds:

  • sovereign-listening: the fifteenth mode. It listens under an explicit listening covenant — a small, human-written declaration of what this ear will not listen to, will release after hearing, will not reveal, will not retain, will blur, or will refuse at certain hours, and why. The covenant does not guarantee obedience: rules a host can execute are enforced at its gates; every line it cannot execute is carried verbatim as a commitment and reported with the hearing. Withholding is honest, attributed absence — counted and named by rule and category, never described — and it is never confused with undetermined.
  • /covenant: the sovereignty command — verify and apply a covenant, listen to what it admits, report enforcement, withholding, and commitments.
  • schemas/covenant.schema.json: the parsed covenant (id, lineage via extends — e.g. the Algophonya manifesto —, executable rules with a small verb vocabulary: do_not_listen, ignore, do_not_reveal, do_not_retain, coarsen, quiet_hours, max_window, require_consent — and carried commitments).
  • An optional covenant block on listening outputs: the covenant's identity plus what was withheld under its rules, so every hearing can answer under which ethics was this listened? Sound-memory stores speaking akousma spec v1.3 keep the same block on records.
  • The default everywhere is no covenant: sovereignty is opted into by the operator, never imposed by the tool — and a covenant governs the listener that adopted it, protecting the listened-to; it is not an instrument for silencing others.

v0.6 Instrumented Listening

The v0.6 release makes AKOÚŌ consumable as data, not just prose, and connects it to sound-memory stores. It adds:

  • akouo.manifest.json: the machine-readable contract — skills with structured metadata (facets, cost tier, memory policy, corrective eligibility), command chains as data, the Evidence Ladder as data, and command permission overrides. Host apps load this instead of hand-copying tables that drift.
  • presets/presets.json: portable listening presets for recurring use-cases (basic, signal, field, music, voice, recall, remember, deep, forensic, access, fiction, generative, extended-spectrum), each with a mode chain, cost tier, memory policy, and perception passes.
  • memory-lineage-listening: the fourteenth mode. It listens WITH stored sound-memories (earworm-style akousma records): recurrence, kinship, lineage, and change over time — while keeping memory as its own evidence stream, never proof about the present sound.
  • /remember: the memory command — situate a sound in its lineage and register the listening into a store.
  • Apparatus declarations: outputs can declare their listening substrate (ASR cascade, audio-token model, speech-native model, DSP toolchain, human ear, hybrid stack), perception sources, and known blind spots, so claim limits derive from the declared apparatus instead of surfacing after the fact.
  • listener (human/agent/hybrid) and memory (akousma links) blocks on outputs, plus akouo_version for contract pinning.
  • Per-claim source (audio, dsp, metadata, model, transcript, context, memory, human) and time_range, so evidence streams never blur and temporal claims stay anchored.
  • Budget-aware routing: plans may carry budget (light/standard/deep) and preset_id; risk always overrides budget.
  • Command permission overrides as data: /forensic suppresses interpreted/speculative claims; /fiction grants declared speculation.

v0.5 Agentic Routing Consolidation

The v0.5 release makes AKOÚŌ more usable as an agent handoff layer. It adds:

  • reference-layer as a portable meta-skill for concepts, methods, traditions, cautions, research routes, and adjacent modes
  • expanded routing plans (routing_plan) for /route and /method, including evidence level, claim permissions, mode chain, stop conditions, and agent handoff notes
  • a router Evidence Ladder that maps available evidence to allowed claim categories
  • a deeper browser-side signal adapter in the reference app: BS.1770-style loudness and loudness range, FFT band-energy and spectral statistics, onset density with a guarded BPM candidate, stereo correlation/width/balance, and clipping-ratio estimates
  • release validation for schema enum alignment, bundled reference schemas, generated artifacts, and public hygiene

v0.4 Agentic Listening Expansion

The v0.4 working set turns AKOÚŌ into a broader listening router for autonomous and cross-app agent workflows. It adds:

  • voice-speech-listening for voice, speech, transcripts, ASR, TTS, voice agents, identity caution, and consent boundaries
  • audiovisual-scenic-listening for video, film, games, UI sound, captions, subtitles, synchronization, and sound-image-scene relations
  • accessibility-normative-listening for hearing norms, captions, transcripts, haptics, assistive paths, sensory variation, fatigue, masking, and implied listener audits
  • material-event-listening for vibration, resonance, duration, flux, material supports, propagation, and event behavior
  • new workflow commands: /voice, /audiovision, /access, /field, /method, and /route

v0.3 Musical/Aesthetic Integration

The v0.3 release promotes musical and aesthetic listening into the public core. musical-aesthetic-listening lets agents describe rhythm, pitch, harmony, timbre, texture, production aesthetics, form, sound-design utility, and poetic usefulness without collapsing into genre labels, cultural overreach, or unsupported claims about tradition, instrument, source, or scene.

v0.2 Conceptual Discipline

The v0.2 skills keep the same public schemas and mode names, but sharpen the conceptual boundaries between:

  • signal, perceptual object, affect, mediation, evidence, ecology, politics, and speculation
  • soundscape, acoustemology, and aurality
  • machine listening, ASR, voice agents, neural audio codecs, and generative audio
  • affect and named emotion
  • testimony, archive, legal evidence, and symbolic witness
  • declared sonic fiction and evidentiary claims

Core Architecture

akoúō is organized as one meta-router, sixteen distinct listening modes, and one conceptual reference layer. These are packaged as portable agent skills (skills/) that can be injected into any LLM agent supporting skill-based system prompts or custom instructions.

  • akouo-router (meta-skill: chooses modes before analysis)
  • signal-inspection-listening
  • acoulogical-object-listening
  • embodied-affective-listening
  • transductive-media-listening
  • forensic-archival-listening
  • ecological-posthuman-listening
  • critical-political-listening
  • musical-aesthetic-listening
  • symbolic-fictional-listening
  • audiovisual-scenic-listening
  • voice-speech-listening
  • accessibility-normative-listening
  • material-event-listening
  • memory-lineage-listening
  • sovereign-listening
  • corpus-listening
  • reference-layer (meta-skill: maps listening to concepts, methods, traditions, and cautions)

Each skill lives in its own folder with a SKILL.md file, following the standard skill format used by OpenCode, Claude Code, and compatible agent frameworks.

Using the System

To give an agent an "akoúō ear":

  1. Supply the agent with the appropriate SKILL.md from the relevant skills/<mode>/ folder as its system prompt or custom instruction.
  2. Supply the agent with the appropriate JSON schema from schemas/, or from the skill's bundled references/ folder when using a skill standalone, to strictly format its output.
  3. Provide the sonic object (as an audio file, a text prompt, a spectrogram, or a transcript).

For agent frameworks that support skill discovery, each SKILL.md includes YAML frontmatter with name and description fields for automatic triggering.

60-Second Reviewer Path

cd app
npm install
npm run dev

Open the local Vite URL, drop a short WAV/AIFF/MP3/OGG file or enter a sound prompt, choose /listen, and run INITIATE LOCAL LISTENING PASS. The deterministic local path runs fully offline: no model provider, benchmark server, credentials, or network connection is required.

Example

User prompt: "Analyze this field recording from a rainforest at night."

Agent loads skills/ecological-posthuman-listening/SKILL.md and produces structured output:

{
  "object_listened_to": "Rainforest night field recording",
  "input_type": "audio_file",
  "listening_mode": "ecological-posthuman-listening",
  "listening_claims": {
    "heard": [],
    "measured": [],
    "inferred": [
      {
        "statement": "An audio model reports layered insect-like textures, distant water movement, and occasional leaf rustle in the supplied recording.",
        "confidence": "medium",
        "basis": "Attributable model observation"
      },
      {
        "statement": "The recording may contain mixed biophony and geophony layers.",
        "confidence": "low",
        "basis": "Ecological layer inference without verified species, location, or spectral inspection"
      }
    ],
    "interpreted": [
      {
        "statement": "The recording presents habitat as an overlap of nonhuman rhythm and weather, not a pure nature scene.",
        "confidence": "medium",
        "basis": "Ecological listening frame"
      }
    ],
    "speculative": [],
    "undetermined": [
      {
        "statement": "Species identity, exact location, season, microphone type, and ecological condition remain unknown.",
        "confidence": "high",
        "basis": "Missing contextual and technical evidence"
      }
    ]
  },
  "what_appears": ["A night field recording can be approached through habitat relations rather than isolated source labels."],
  "what_remains_hidden": ["Species, site, season, recorder position, and ecological condition remain unavailable."],
  "mediations": {
    "technical": ["The microphone and file container mediate access to the habitat."],
    "cultural": [],
    "spatial": ["Distance and microphone position shape what becomes audible."],
    "bodily": [],
    "archival": [],
    "computational": []
  },
  "risks": {
    "hallucination": ["Do not invent species or location from texture alone."],
    "over_identification": ["Do not identify animals without evidence."],
    "cultural_flattening": [],
    "forensic_overreach": [],
    "source_confusion": ["Do not treat the recording as direct unmediated nature."],
    "aesthetic_overstatement": ["Avoid romanticizing the scene as pure wilderness."]
  },
  "main_reading": "The recording should be read as a mediated habitat relation across nonhuman, weather, spatial, and technical layers.",
  "alternative_reading": "A signal-inspection pass could ground this reading in measured frequency, duration, noise, and dynamic traits.",
  "recommended_next_mode": "transductive-media-listening"
}

Commands (located in commands/) combine these skills into reusable listening chains, like /one-sound-many-ears, /forensic, /voice, /audiovision, /access, /field, /corpus, /method, or /route.

For the agent-to-agent handoff format, see examples/routing-plan-example.json: an expanded routing plan (per schemas/routing-plan.schema.json) carrying route_confidence, evidence_level, claim_permissions, a mode_chain, stop_conditions, and an agent_handoff summary that a receiving agent can act on without re-listening.

Repository Layout

akouo/
  README.md
  SYSTEM_GUIDE.md      # Operational guide for commands, workflows, and app contract
  SKILL_INDEX.md       # Quick-reference manifest of all skills
  CHANGELOG.md         # Release history from v0.1 through v0.9.2
  akouo.manifest.json  # Machine-readable system contract (skills, commands, ladder, overrides)
  LICENSE
  .gitignore
  app/                 # Local-first reference app for running AKOÚŌ workflows
  presets/
    presets.json         # Portable listening presets (validated by schemas/preset.schema.json)
  scripts/
    validate-release.sh  # Pre-release validation script
  skills/
    akouo-router/
      SKILL.md
      references/
    signal-inspection-listening/
      SKILL.md
      references/
    acoulogical-object-listening/
      SKILL.md
      references/
    embodied-affective-listening/
      SKILL.md
      references/
    transductive-media-listening/
      SKILL.md
      references/
    forensic-archival-listening/
      SKILL.md
      references/
    ecological-posthuman-listening/
      SKILL.md
      references/
    critical-political-listening/
      SKILL.md
      references/
    musical-aesthetic-listening/
      SKILL.md
      references/
    symbolic-fictional-listening/
      SKILL.md
      references/
    audiovisual-scenic-listening/
      SKILL.md
      references/
    voice-speech-listening/
      SKILL.md
      references/
    accessibility-normative-listening/
      SKILL.md
      references/
    material-event-listening/
      SKILL.md
      references/
    memory-lineage-listening/
      SKILL.md
      references/
    sovereign-listening/
      SKILL.md
      references/
    reference-layer/
      SKILL.md
      references/
  commands/
  schemas/
  examples/

Public-Repo Hygiene

This repository is safe for public release. It contains no private recordings, API keys, personal data, or local system paths. All skills use generic framework-agnostic instructions with no dependency on a specific model provider. node_modules/ and build artifacts are excluded by .gitignore.

Conceptual refinements should be incorporated as public-facing skill language only. Do not bundle private notes, local research directories, unpublished source maps, private file paths, or personal identifiers into the portable skills.

Run ./scripts/validate-release.sh before publishing to verify skill structure, schema consistency, hygiene, and absence of generated build artifacts.

For a full operational guide covering commands, workflows, benchmark ingestion, and data structure, see SYSTEM_GUIDE.md.

The store-connected flow now exists as its own app: the akousmata listening navigator (github.com/sonicfieldlabs/akousmata) is the reference library over the shared store — memory-lineage listening and /remember run against it, and manual human listening events enter the same library as agent listenings (listener human, contract akousmata/v0.1).

License

MIT License. See LICENSE.

About

Operational listening contract and skill library for agentic audio workflows.

Resources

Stars

10 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages