You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add the concrete audio-analysis ingestion and semantic-description pipeline for Renderflow's Sonic DNA work (#387).
Given an immutable audio artifact, reviewed lyrics/MIDI when available, and optional time-indexed analysis from Aniflow, Renderflow should emit:
normalized Sonic DNA as renderflow.artifact-dna/v1;
an evidence-backed English description of the sound and production;
sanitized, provider-neutral prompt guidance for generating related—not copied—assets;
source, confidence, hygiene, model/tool, and approval provenance.
A “song upload to useful production description” experience is the reference outcome. Suno and similar products are experience references only; this feature must not depend on, copy, or integrate their proprietary implementation or prompts.
Ownership boundary
Aniflow
Aniflow owns time-dependent extraction and synchronization through:
Flow orchestrates providers and cross-holon recovery. Dreamscape, Antidote, publication repositories, and other products consume the outputs; they do not become the canonical implementation of extraction or DNA.
Do not create a new standalone tool in this issue. If the normalized audio-intelligence domain later develops an independent lifecycle and multiple non-Renderflow consumers, extraction behind the existing contracts can be promoted without changing artifact schemas.
Architecture
audio + reviewed companion artifacts
↓
Renderflow intake / Aniflow analysis provider
↓
deterministic Sonic DNA normalization
↓
optional versioned semantic-description AI skill
↓
schema validation + protected-reference hygiene
↓
DNA JSON + English description + prompt guidance
↓
reviewed downstream consumers
Sonic DNA dimensions
Normalize available evidence into extensible dimensions including:
tempo, tempo movement, meter, groove, swing, and rhythmic density;
key/tonal center, mode, harmonic movement, chord density, and ambiguity;
instrumentation and stem roles;
vocal presence, register, delivery, layering, and processing;
timbre, spectral balance, texture, distortion, and transient character;
dynamics, loudness, energy contour, and contrast;
stereo image, depth, ambience, reverb, delay, and spatial motion;
arrangement, section order, repetition, transitions, and climax/release;
production character, recording cues, mix balance, and mastering traits;
lyrical tone, cadence, themes, language, and explicit reviewed/transcribed distinction;
MIDI/note-density observations and confidence;
mood, energy, narrative function, motifs, and cross-modal anchors;
intended-use guidance and exclusions when explicitly reviewed.
Each observation must carry its evidence origin, provider/tool/model version, time or stem scope when applicable, confidence, and deterministic/heuristic/probabilistic classification.
Missing evidence stays unknown. A semantic model must not invent exact BPM, key, instruments, lyrics, or production facts when deterministic/analyzer evidence is absent.
Layered extraction
Layer 1: deterministic technical metadata
Always produce a useful baseline where local inspection is available:
format, codec, duration, sample rate, channels, bit depth/sample format;
loudness/peak/dynamic-range evidence;
embedded metadata under hygiene policy;
source and companion-artifact digests.
Layer 2: specialized analysis providers
Consume normalized Aniflow analysis when requested and available. Preserve disagreements and confidence instead of collapsing estimates into unsupported certainty.
Renderflow may expose other replaceable analyzer adapters when they fit the artifact boundary, but it must not reimplement Aniflow's temporal domain logic.
Layer 3: semantic audio understanding
Use a versioned renderflow.ai-skill/v1 skill from #396 to transform structured evidence and, only when policy permits, bounded audio/stem excerpts into schema-valid semantic output.
The skill must declare an audio-capable or multimodal model requirement. A text-only model may interpret structured analyzer output but must not be represented as directly hearing audio.
Layer 4: human review and correction
Allow reviewed overrides/corrections for subjective or incorrect observations. Preserve both the original candidate and the reviewed value with authority and provenance.
English semantic description
Define a versioned structured result that can render to concise and detailed English without making prose the machine contract.
At minimum support:
concise summary;
detailed production description;
rhythm and tempo description;
tonal/harmonic description;
instrumentation/timbre description;
vocal description;
arrangement/energy description;
spatial/mix/mastering description;
lyrical/theme description when permitted;
confidence and uncertainty notes;
related-asset guidance;
explicit “avoid/direct imitation” terms after hygiene.
Descriptions should read naturally enough to support Dreamscape and creative tools while remaining traceable to structured evidence.
Prompt guidance
Generate provider-neutral prompt components from sanitized Sonic DNA rather than copying a platform-specific prompt format.
Support:
compact tag list;
natural-language creative brief;
positive characteristic guidance;
negative/avoid guidance;
target modality and intended transformation;
selected dimensions with weights or importance;
omissions caused by policy;
prompt skill/version and input DNA digest.
Provider adapters may later render this structure into a particular API request. The canonical output remains provider-neutral.
Copyright, trademark, artist, and privacy hygiene
Before semantic interpretation and again before downstream prompt use:
block secrets, PII, unapproved lyrics/audio excerpts, and denied metadata;
detect franchise, brand, artist, living-creator, track, album, and protected-work references;
support policy-driven replacement of named imitation cues with descriptive musical characteristics;
avoid “in the style of [artist]” and direct track-copy instructions when policy forbids them;
avoid presenting a genre, common technique, chord progression, or production characteristic as owned merely because it resembles a known work;
retain sensitive findings in private evidence without leaking blocked values into public manifests;
mark all probabilistic rewrites and generated descriptions as review-required;
never claim copyright clearance or commercial suitability.
Companion artifacts
Support optional relationships to:
reviewed lyrics;
observed transcription;
LRC, SRT, WebVTT, TTML, and other timed text;
artist-authored MIDI;
probabilistic MIDI candidates;
stem manifests;
cover art or related visual DNA;
project/track metadata.
Identity and authority must remain explicit. A transcript is not silently promoted to canonical lyrics; extracted MIDI is not silently promoted to artist-authored MIDI.
Versioned artifacts
Produce inspectable artifacts such as:
sonic-dna.json;
audio-description.json;
optional rendered audio-description.md;
prompt-guidance.json;
validation and hygiene evidence;
provider/model/analyzer provenance.
All JSON artifacts require schemas. Artifact IDs/digests, skill/model/provider identity, analysis inputs, reviewed overrides, and policy versions must participate in cache/resume compatibility.
CLI and profile integration
Expose explicit operations conceptually similar to:
renderflow dna extract --modality "audio"
renderflow dna describe --modality "audio"
renderflow dna prompt --target "audio"
Also allow derivative profiles to request Sonic DNA and descriptions as optional intermediate/terminal artifacts under ordinary network, AI, storage, runtime, and validation budgets.
A deterministic/local-only profile must still emit baseline Sonic DNA when optional AI providers are unavailable. Unavailability should be structured rather than fatal unless the requested target requires semantic output.
Outcome
Add the concrete audio-analysis ingestion and semantic-description pipeline for Renderflow's Sonic DNA work (#387).
Given an immutable audio artifact, reviewed lyrics/MIDI when available, and optional time-indexed analysis from Aniflow, Renderflow should emit:
renderflow.artifact-dna/v1;A “song upload to useful production description” experience is the reference outcome. Suno and similar products are experience references only; this feature must not depend on, copy, or integrate their proprietary implementation or prompts.
Ownership boundary
Aniflow
Aniflow owns time-dependent extraction and synchronization through:
This includes stems, BPM/beat grids, key estimates, sections, timed lyrics/transcripts, MIDI candidates, loudness timelines, and analyzer provenance.
Renderflow
Renderflow owns:
Flow and products
Flow orchestrates providers and cross-holon recovery. Dreamscape, Antidote, publication repositories, and other products consume the outputs; they do not become the canonical implementation of extraction or DNA.
Do not create a new standalone tool in this issue. If the normalized audio-intelligence domain later develops an independent lifecycle and multiple non-Renderflow consumers, extraction behind the existing contracts can be promoted without changing artifact schemas.
Architecture
Sonic DNA dimensions
Normalize available evidence into extensible dimensions including:
Each observation must carry its evidence origin, provider/tool/model version, time or stem scope when applicable, confidence, and deterministic/heuristic/probabilistic classification.
Missing evidence stays unknown. A semantic model must not invent exact BPM, key, instruments, lyrics, or production facts when deterministic/analyzer evidence is absent.
Layered extraction
Layer 1: deterministic technical metadata
Always produce a useful baseline where local inspection is available:
Layer 2: specialized analysis providers
Consume normalized Aniflow analysis when requested and available. Preserve disagreements and confidence instead of collapsing estimates into unsupported certainty.
Renderflow may expose other replaceable analyzer adapters when they fit the artifact boundary, but it must not reimplement Aniflow's temporal domain logic.
Layer 3: semantic audio understanding
Use a versioned
renderflow.ai-skill/v1skill from #396 to transform structured evidence and, only when policy permits, bounded audio/stem excerpts into schema-valid semantic output.The skill must declare an audio-capable or multimodal model requirement. A text-only model may interpret structured analyzer output but must not be represented as directly hearing audio.
Layer 4: human review and correction
Allow reviewed overrides/corrections for subjective or incorrect observations. Preserve both the original candidate and the reviewed value with authority and provenance.
English semantic description
Define a versioned structured result that can render to concise and detailed English without making prose the machine contract.
At minimum support:
Descriptions should read naturally enough to support Dreamscape and creative tools while remaining traceable to structured evidence.
Prompt guidance
Generate provider-neutral prompt components from sanitized Sonic DNA rather than copying a platform-specific prompt format.
Support:
Provider adapters may later render this structure into a particular API request. The canonical output remains provider-neutral.
Copyright, trademark, artist, and privacy hygiene
Before semantic interpretation and again before downstream prompt use:
Companion artifacts
Support optional relationships to:
Identity and authority must remain explicit. A transcript is not silently promoted to canonical lyrics; extracted MIDI is not silently promoted to artist-authored MIDI.
Versioned artifacts
Produce inspectable artifacts such as:
sonic-dna.json;audio-description.json;audio-description.md;prompt-guidance.json;All JSON artifacts require schemas. Artifact IDs/digests, skill/model/provider identity, analysis inputs, reviewed overrides, and policy versions must participate in cache/resume compatibility.
CLI and profile integration
Expose explicit operations conceptually similar to:
Also allow derivative profiles to request Sonic DNA and descriptions as optional intermediate/terminal artifacts under ordinary network, AI, storage, runtime, and validation budgets.
A deterministic/local-only profile must still emit baseline Sonic DNA when optional AI providers are unavailable. Unavailability should be structured rather than fatal unless the requested target requires semantic output.
Fixtures
Add redistribution-safe synthetic fixtures covering:
No paid API, proprietary prompt, copyrighted recording, or model-weight download may be required.
Documentation
Document:
Non-goals
Acceptance criteria
Dependencies