Skip to content

Add Sonic DNA normalization and semantic audio descriptions #397

Description

@szmyty

Outcome

Add the concrete audio-analysis ingestion and semantic-description pipeline for Renderflow's Sonic DNA work (#387).

Given an immutable audio artifact, reviewed lyrics/MIDI when available, and optional time-indexed analysis from Aniflow, Renderflow should emit:

  1. normalized Sonic DNA as renderflow.artifact-dna/v1;
  2. an evidence-backed English description of the sound and production;
  3. sanitized, provider-neutral prompt guidance for generating related—not copied—assets;
  4. source, confidence, hygiene, model/tool, and approval provenance.

A “song upload to useful production description” experience is the reference outcome. Suno and similar products are experience references only; this feature must not depend on, copy, or integrate their proprietary implementation or prompts.

Ownership boundary

Aniflow

Aniflow owns time-dependent extraction and synchronization through:

This includes stems, BPM/beat grids, key estimates, sections, timed lyrics/transcripts, MIDI candidates, loudness timelines, and analyzer provenance.

Renderflow

Renderflow owns:

Flow and products

Flow orchestrates providers and cross-holon recovery. Dreamscape, Antidote, publication repositories, and other products consume the outputs; they do not become the canonical implementation of extraction or DNA.

Do not create a new standalone tool in this issue. If the normalized audio-intelligence domain later develops an independent lifecycle and multiple non-Renderflow consumers, extraction behind the existing contracts can be promoted without changing artifact schemas.

Architecture

audio + reviewed companion artifacts
        ↓
Renderflow intake / Aniflow analysis provider
        ↓
deterministic Sonic DNA normalization
        ↓
optional versioned semantic-description AI skill
        ↓
schema validation + protected-reference hygiene
        ↓
DNA JSON + English description + prompt guidance
        ↓
reviewed downstream consumers

Sonic DNA dimensions

Normalize available evidence into extensible dimensions including:

  • tempo, tempo movement, meter, groove, swing, and rhythmic density;
  • key/tonal center, mode, harmonic movement, chord density, and ambiguity;
  • instrumentation and stem roles;
  • vocal presence, register, delivery, layering, and processing;
  • timbre, spectral balance, texture, distortion, and transient character;
  • dynamics, loudness, energy contour, and contrast;
  • stereo image, depth, ambience, reverb, delay, and spatial motion;
  • arrangement, section order, repetition, transitions, and climax/release;
  • production character, recording cues, mix balance, and mastering traits;
  • lyrical tone, cadence, themes, language, and explicit reviewed/transcribed distinction;
  • MIDI/note-density observations and confidence;
  • mood, energy, narrative function, motifs, and cross-modal anchors;
  • intended-use guidance and exclusions when explicitly reviewed.

Each observation must carry its evidence origin, provider/tool/model version, time or stem scope when applicable, confidence, and deterministic/heuristic/probabilistic classification.

Missing evidence stays unknown. A semantic model must not invent exact BPM, key, instruments, lyrics, or production facts when deterministic/analyzer evidence is absent.

Layered extraction

Layer 1: deterministic technical metadata

Always produce a useful baseline where local inspection is available:

  • format, codec, duration, sample rate, channels, bit depth/sample format;
  • loudness/peak/dynamic-range evidence;
  • embedded metadata under hygiene policy;
  • source and companion-artifact digests.

Layer 2: specialized analysis providers

Consume normalized Aniflow analysis when requested and available. Preserve disagreements and confidence instead of collapsing estimates into unsupported certainty.

Renderflow may expose other replaceable analyzer adapters when they fit the artifact boundary, but it must not reimplement Aniflow's temporal domain logic.

Layer 3: semantic audio understanding

Use a versioned renderflow.ai-skill/v1 skill from #396 to transform structured evidence and, only when policy permits, bounded audio/stem excerpts into schema-valid semantic output.

The skill must declare an audio-capable or multimodal model requirement. A text-only model may interpret structured analyzer output but must not be represented as directly hearing audio.

Layer 4: human review and correction

Allow reviewed overrides/corrections for subjective or incorrect observations. Preserve both the original candidate and the reviewed value with authority and provenance.

English semantic description

Define a versioned structured result that can render to concise and detailed English without making prose the machine contract.

At minimum support:

  • concise summary;
  • detailed production description;
  • rhythm and tempo description;
  • tonal/harmonic description;
  • instrumentation/timbre description;
  • vocal description;
  • arrangement/energy description;
  • spatial/mix/mastering description;
  • lyrical/theme description when permitted;
  • confidence and uncertainty notes;
  • related-asset guidance;
  • explicit “avoid/direct imitation” terms after hygiene.

Descriptions should read naturally enough to support Dreamscape and creative tools while remaining traceable to structured evidence.

Prompt guidance

Generate provider-neutral prompt components from sanitized Sonic DNA rather than copying a platform-specific prompt format.

Support:

  • compact tag list;
  • natural-language creative brief;
  • positive characteristic guidance;
  • negative/avoid guidance;
  • target modality and intended transformation;
  • selected dimensions with weights or importance;
  • omissions caused by policy;
  • prompt skill/version and input DNA digest.

Provider adapters may later render this structure into a particular API request. The canonical output remains provider-neutral.

Copyright, trademark, artist, and privacy hygiene

Before semantic interpretation and again before downstream prompt use:

  • block secrets, PII, unapproved lyrics/audio excerpts, and denied metadata;
  • detect franchise, brand, artist, living-creator, track, album, and protected-work references;
  • support policy-driven replacement of named imitation cues with descriptive musical characteristics;
  • avoid “in the style of [artist]” and direct track-copy instructions when policy forbids them;
  • avoid presenting a genre, common technique, chord progression, or production characteristic as owned merely because it resembles a known work;
  • retain sensitive findings in private evidence without leaking blocked values into public manifests;
  • mark all probabilistic rewrites and generated descriptions as review-required;
  • never claim copyright clearance or commercial suitability.

Companion artifacts

Support optional relationships to:

  • reviewed lyrics;
  • observed transcription;
  • LRC, SRT, WebVTT, TTML, and other timed text;
  • artist-authored MIDI;
  • probabilistic MIDI candidates;
  • stem manifests;
  • cover art or related visual DNA;
  • project/track metadata.

Identity and authority must remain explicit. A transcript is not silently promoted to canonical lyrics; extracted MIDI is not silently promoted to artist-authored MIDI.

Versioned artifacts

Produce inspectable artifacts such as:

  • sonic-dna.json;
  • audio-description.json;
  • optional rendered audio-description.md;
  • prompt-guidance.json;
  • validation and hygiene evidence;
  • provider/model/analyzer provenance.

All JSON artifacts require schemas. Artifact IDs/digests, skill/model/provider identity, analysis inputs, reviewed overrides, and policy versions must participate in cache/resume compatibility.

CLI and profile integration

Expose explicit operations conceptually similar to:

renderflow dna extract --modality "audio"
renderflow dna describe --modality "audio"
renderflow dna prompt --target "audio"

Also allow derivative profiles to request Sonic DNA and descriptions as optional intermediate/terminal artifacts under ordinary network, AI, storage, runtime, and validation budgets.

A deterministic/local-only profile must still emit baseline Sonic DNA when optional AI providers are unavailable. Unavailability should be structured rather than fatal unless the requested target requires semantic output.

Fixtures

Add redistribution-safe synthetic fixtures covering:

  • audio with known technical metadata;
  • known BPM/key analysis evidence with confidence;
  • reviewed lyrics plus a distinct transcript/timed-text candidate;
  • an artist-authored MIDI reference plus a probabilistic MIDI candidate;
  • Aniflow stem/analysis manifest ingestion;
  • deterministic-only Sonic DNA;
  • semantic skill output from a mock/local provider;
  • protected artist/franchise/track references rewritten or blocked by policy;
  • unavailable audio-capable model evidence.

No paid API, proprietary prompt, copyrighted recording, or model-weight download may be required.

Documentation

Document:

  • the Aniflow/Renderflow/Flow/product ownership split;
  • deterministic analysis versus semantic interpretation;
  • model compatibility requirements;
  • local-first and offline behavior;
  • companion artifact authority;
  • prompt sanitization and rights boundaries;
  • how Dreamscape and other products consume Sonic DNA;
  • how to add an analyzer mapping or semantic skill;
  • correction, review, rollback, cache, and reproduction behavior.

Non-goals

  • Depending on Suno or reproducing its implementation/prompts.
  • Building a music generator.
  • Training/fine-tuning an audio model.
  • Reimplementing Aniflow stem, beat, transcript, or alignment logic.
  • Treating model descriptions as objective truth.
  • Automatically approving generated prompts or artifacts.
  • Making Flow or Dreamscape own the extraction implementation.
  • Creating a new standalone repository before the domain boundary proves it needs one.

Acceptance criteria

  • Audio and Aniflow analysis artifacts normalize into versioned Sonic DNA.
  • Deterministic, heuristic, and probabilistic observations remain distinguishable.
  • BPM, key, MIDI, lyrics, transcript, stem, and timeline evidence retain authority/confidence/provenance.
  • Baseline Sonic DNA works locally without an AI model.
  • A versioned Add a model compatibility matrix and spec-driven AI skill runtime #396 AI skill produces schema-valid semantic audio descriptions.
  • Direct-audio interpretation requires an honestly compatible audio/multimodal model.
  • Text-only models can interpret structured evidence without being represented as hearing audio.
  • English summaries and detailed production descriptions render from structured output.
  • Provider-neutral prompt guidance can be derived from sanitized DNA.
  • Protected artist/work/franchise/brand imitation cues can be blocked or rewritten descriptively.
  • Private lyrics/audio/prompts are not leaked through public evidence.
  • Outputs enter the artifact graph with candidate/approval state and complete provenance.
  • Cache/resume identity covers source, analysis, DNA schema, skill, provider/model, configuration, and hygiene policy.
  • Profile/CLI integrations respect AI, network, locality, and budget policy.
  • Synthetic fixtures require no paid API, copyrighted media, or model download.
  • Documentation covers Dreamscape and other downstream consumers without coupling core logic to them.
  • Formatting, Clippy, tests, docs, and CI pass.

Dependencies

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions