Music generation
clavue-music is the official song product id (premium pool). Call POST /v1/audio/speech on api.clavue.com. Flagship: TTT-ready tagged lyrics + three-heading Structured Caption, 30 seconds, seed 11, one shot. A user one-liner is a convenience path — the hop composes lyrics ∥ caption in parallel, then renders. Do not send 180 or 240 as a single engine call (90∥90 / 110∥110). News reads and 朗读 use clavue-tts. Do not ship upstream names in product code.
How to generate
- POST /v1/audio/speech with model=clavue-music
- Flagship: send tagged lyrics + three-heading caption yourself (30s / seed 11 / one shot)
- One-liner: the hop composes lyrics ∥ caption in parallel (thinking off), then renders — weaker than a locked pack
- L1 clients: fire two auto rewrites in parallel (caption ∥ lyrics), then speech
- input = tagged lyrics (tags on their own line) or (instrumental)
- instructions = English Structured Caption (Global Metadata / Vocal Details / Arrangement)
- audio_duration default 30; engine max 110; seed default 11; also send voice as the same seconds string
- 180s: two parallel 90s calls, same caption, reused Chorus — never audio_duration 180
- Cluster live concurrency 12 (do not fill it with 10s jobs; bulk 4–8)
curl -sS https://api.clavue.com/v1/audio/speech \
-H "Authorization: Bearer $CLAVUE_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model":"clavue-music",
"input":"[Verse]\nMorning light filtering through the pine\n[Chorus]\nStay with me in this light",
"instructions":"Global Metadata\nAcoustic pop, 96 BPM, intimate morning.\nVocal Details\nSoft female lead, close-mic.\nArrangement\nFingerpicked guitar, soft piano, brushed snare in the chorus.",
"audio_duration":30,
"voice":"30",
"seed":11,
"response_format":"wav"
}' --output song.wav --max-time 600Compose first (TTT)
TTT is the compose controller above clavue-music, not a fourth engine and not Music3 fine-tuning. Retrieve, critic, and --commit stay in imusic — they do not run on every public speech request. The public hop composes a one-liner in parallel (lyrics ∥ caption, thinking off) then renders. Flagship quality still comes from a TTT-ready payload you store yourself. Public traffic must never default-commit retrieval weights.
Recommended settings
Engine sampling is locked. Send duration, seed, lyrics, and caption only. Timeout ≥600s per clip. Do not send temperature, top_p, speed, or named voices.
- audio_duration: 30 (default) · engine max 110 · product max 240 as 90∥90 or 110∥110
- voice: the same seconds as a string, e.g. "30"
- seed: 11 on the hot path; change seed only when the user asks for another take
- response_format: wav · read the real WAV duration, not just HTTP 200
{
"model": "clavue-music",
"audio_duration": 30,
"voice": "30",
"seed": 11,
"response_format": "wav"
}Prompt demo (30s acoustic)
Flagship: rewrite with auto or clavue-2.1-fast (thinking off) — lyrics ∥ caption in parallel — then send those to /v1/audio/speech. The hop will compose a one-liner if you skip that. Tags must sit on their own line. Weak caption: Genre: pop. BPM: 96.
# User intent
Duration: 30 seconds. Acoustic pop. Soft female, close-mic, no belting.
Theme: pine needles at a morning window, a quiet street.
No full drum kit, no celebrity names.
# Caption rewriter system
You are a song caption engineer. Expand the user's intent into an English Structured Caption.
Output only three headings. No title, no lyrics, no explanation.
### Global Metadata — genre, BPM, mood arc, scene, production
### Vocal Details — gender, register, verse vs chorus, harmony, space
### Arrangement — instrument entries by section; 2–3 lead instruments; 250–450 words
# Lyrics rewriter system
Write lyrics for the target duration. Tags on their own line.
Allowed: [Intro] [Verse] [Pre-Chorus] [Chorus] [Bridge] [Outro]
Chorus short and repeatable. ~2–3 Chinese characters per second.
30s = Verse + Chorus, 8–12 lines. Lyrics only.
# Example input (lyrics)
[Verse]
晨光穿过窗边的松针
这条安静的街像只属于我们
把昨夜的 rumble 轻轻放下
让呼吸自己找到节奏
[Chorus]
世界慢慢醒来
你走在我左边
别说话,先听风
把名字吹成一天
# Example instructions (caption)
Global Metadata
Basic Attributes: bpm is 96. key is C, and scale is major. Contemporary acoustic pop.
Global Emotional Progression: Opens intimate and still, then lifts into a wider, hopeful chorus without turning anthemic.
Application Scenarios & Imagery: Early morning apartment, pine light through a window.
Sonics & Production Profile: Intimate, mid-focused, warm low mids, no harsh cymbals.
Vocal Details
Soft female lead, breathy, close to the microphone. Conversational verse; sustained chorus; no belting.
Light stacked doubles only in the chorus. Short plate, almost dry in the verse.
Arrangement
Intro: fingerpicked steel-string guitar alone.
Verse: guitar plus very soft piano; no full drum kit.
Chorus: brushed snare, upright bass, piano opens; keep the guitar pattern continuous.Limits
Single engine call 10–110s. Timeout ≥600s per clip. Do not send temperature, speed, or named voices. Lyrics shorter than the request end early — read the WAV duration. 180/240 must be two parallel clips.
Not this product
Broadcast, news desk, and scripted 朗读 use clavue-tts on the same /v1/audio/speech endpoint with a preset voice. Do not send a news script as lyrics. Do not send song captions as TTS instructions.