Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 25 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ Agents and external tools should inspect a model declaration before constructing
import { createGenerationClient } from "@neta-art/generation";

const discoveryClient = createGenerationClient();
const declaration = discoveryClient.getModel("qwen-tts");
const declaration = discoveryClient.getModel("cosyvoice-v3.5-plus");
if (!declaration) throw new Error("Model is unavailable");

console.log(discoveryClient.stringifyModelConfig(declaration.model, { format: "json" }));
Expand All @@ -84,7 +84,7 @@ The same declarations can be exported as YAML through the existing CLI:

```bash
neta-generation models list
neta-generation models export qwen-tts --out ./qwen-tts.yaml
neta-generation models export cosyvoice-v3.5-plus --out ./cosyvoice-v3.5-plus.yaml
neta-generation models export-all --out ./models
```

Expand Down Expand Up @@ -170,7 +170,8 @@ const client = createGenerationClient({
- `gpt-image-2`
- `z-image-turbo`
- `qwen-image-edit`
- `qwen-tts`
- `cosyvoice-v3.5-plus`
- `cosyvoice-v3.5-flash`
- `qwen-audio-3.0-tts-plus`
- `qwen-audio-3.0-tts-flash`
- `higgs-tts`
Expand Down Expand Up @@ -268,22 +269,38 @@ Each TTS request accepts exactly one non-empty text block and returns one URL au

| Requirement | Model choice |
| --- | --- |
| Create a voice from a text-only description, without reference audio | Use an explicitly requested Qwen variant; otherwise use `qwen-tts` as the deterministic default |
| Create a voice from a text-only description, without reference audio | Use an explicitly requested cosyvoice variant; otherwise use `cosyvoice-v3.5-flash` as the deterministic default |
| Maximize fidelity to one reference voice | `higgs-tts` |
| Blend 2-16 weighted reference voices | `higgs-tts` |
| Use a default voice, including a delegated choice expressed only as any, random, suitable, or natural | `higgs-tts` |

- Qwen: `voice_prompt` design OR one-reference clone; `qwen-tts` is the unspecified-design default and accepts any text length; Plus / Flash require at least 15 Unicode code points.
- CosyVoice / Qwen-Audio-TTS: `voice_prompt` design OR one-reference clone; `cosyvoice-v3.5-plus`, `cosyvoice-v3.5-flash`, `qwen-audio-3.0-tts-plus`, and `qwen-audio-3.0-tts-flash` all require at least 15 Unicode code points (and at most 200 in design mode); `cosyvoice-v3.5-flash` is the deterministic default when no variant is requested.
- Higgs: delegated default voice, high-fidelity one-reference clone, or weighted 2-16-reference blend.
- Conflict: reference + redesign requires user choice before generation.
- Blend: all references, full text, one request.
- Dependency: clone prior generated audio.
- Ranking: no declared Qwen quality, latency, or cost order.
- Ranking: no declared CosyVoice / Qwen quality, latency, or cost order.

> **Migrating from `qwen-tts` (removed in 0.2.0):** DashScope retires `qwen-tts` on 2026-10-10; this SDK removed
> the declaration in the same release. Switch to `cosyvoice-v3.5-flash` (or `-plus`) — same request shape
> (`voice_prompt` design OR one-reference clone), but `qwen-tts` accepted text of **any length** while
> `cosyvoice-v3.5-*` requires **at least 15 Unicode code points** (and at most 200 in design mode, like the rest
> of this model family). A caller doing a plain model-name swap on short input will start seeing
> `GenerationValidationError` where it previously succeeded — pad short inputs or catch the error.

> **Server-side wire contract:** this SDK talks to the router over HTTP; it does not call the DashScope-facing
> worker (`talesofai/background`'s `qwen_tts_actor`) directly, and as of 2026-09-20 nothing in the pipeline
> between them performs the translation yet. For whoever builds that glue, per the worker's own field names
> (see its module docstring — they do **not** match DashScope's or this SDK's field names one-for-one, which is
> the trap to avoid): the worker's `preview_text` (what's actually spoken, required on every request) is this
> request's primary `text` content block; the worker's `text` field (the voice STYLE description, design mode
> only — unrelated to this SDK's own `input`/text-content naming despite the shared word) is this request's
> `meta.voice_prompt`; `target_model` is this request's `model`, unchanged.

```ts
await client.generate({
model: "qwen-tts",
content: [{ type: "text", text: "欢迎使用语音合成功能。" }],
model: "cosyvoice-v3.5-flash",
content: [{ type: "text", text: "欢迎使用语音合成功能,这是一段示例文本。" }],
meta: {
voice_prompt: "一位沉稳自然的中文播音员,吐字清晰,语速适中",
},
Expand Down
2 changes: 1 addition & 1 deletion examples/text-to-speech.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ if (!apiKey) throw new Error("Set NETA_ROUTER_API_KEY or NETA_API_KEY");

const client = createGenerationClient({ apiKey });
const output = await client.generate({
model: "qwen-tts",
model: "cosyvoice-v3.5-flash",
content: [{ type: "text", text: "欢迎使用语音合成功能,这是一段示例文本。" }],
meta: {
voice_prompt: "一位沉稳自然的中文播音员,吐字清晰,语速适中",
Expand Down
46 changes: 46 additions & 0 deletions models/cosyvoice-v3.5-flash.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
schema: neta.generation.model.v1
model: cosyvoice-v3.5-flash
title: CosyVoice v3.5 Flash
description: "Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points (design mode: <=200).
Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio."
adapter:
type: openai.audioSpeech
content:
input:
- type: text
required: true
min: 1
max: 1
description: Exactly one non-empty text block to speak, with at least 15 Unicode code points (at most 200 in
voice-design mode).
- type: audio
required: false
max: 1
sources:
- url
description: "Clone: one URL; no voice_prompt. Dependency: use prior generated audio."
meta:
fields:
voice_prompt:
type: string
optional: true
description: "Design: custom voice text; no reference audio."
examples:
- title: Voice design
request:
model: cosyvoice-v3.5-flash
content:
- type: text
text: 这是一段长度足够并且表达清晰自然的语音合成测试文本。
meta:
voice_prompt: 一位沉稳干练的男性播音员声音,吐字清晰有力
- title: Voice clone
request:
model: cosyvoice-v3.5-flash
content:
- type: text
text: 这是一段长度足够并且表达清晰自然的语音合成测试文本。
- type: audio
source:
type: url
url: https://example.com/reference.mp3
46 changes: 46 additions & 0 deletions models/cosyvoice-v3.5-plus.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
schema: neta.generation.model.v1
model: cosyvoice-v3.5-plus
title: CosyVoice v3.5 Plus
description: "Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points (design mode: <=200).
Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio."
adapter:
type: openai.audioSpeech
content:
input:
- type: text
required: true
min: 1
max: 1
description: Exactly one non-empty text block to speak, with at least 15 Unicode code points (at most 200 in
voice-design mode).
- type: audio
required: false
max: 1
sources:
- url
description: "Clone: one URL; no voice_prompt. Dependency: use prior generated audio."
meta:
fields:
voice_prompt:
type: string
optional: true
description: "Design: custom voice text; no reference audio."
examples:
- title: Voice design
request:
model: cosyvoice-v3.5-plus
content:
- type: text
text: 这是一段长度足够并且表达清晰自然的语音合成测试文本。
meta:
voice_prompt: 一位沉稳干练的男性播音员声音,吐字清晰有力
- title: Voice clone
request:
model: cosyvoice-v3.5-plus
content:
- type: text
text: 这是一段长度足够并且表达清晰自然的语音合成测试文本。
- type: audio
source:
type: url
url: https://example.com/reference.mp3
10 changes: 6 additions & 4 deletions models/qwen-audio-3.0-tts-flash.yaml
Original file line number Diff line number Diff line change
@@ -1,7 +1,8 @@
schema: neta.generation.model.v1
model: qwen-audio-3.0-tts-flash
title: Qwen Audio 3.0 TTS Flash
description: 'Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points. Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio.'
description: "Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points (design mode: <=200).
Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio."
adapter:
type: openai.audioSpeech
content:
Expand All @@ -10,19 +11,20 @@ content:
required: true
min: 1
max: 1
description: Exactly one non-empty text block to speak, with at least 15 Unicode code points.
description: Exactly one non-empty text block to speak, with at least 15 Unicode code points (at most 200 in
voice-design mode).
- type: audio
required: false
max: 1
sources:
- url
description: 'Clone: one URL; no voice_prompt. Dependency: use prior generated audio.'
description: "Clone: one URL; no voice_prompt. Dependency: use prior generated audio."
meta:
fields:
voice_prompt:
type: string
optional: true
description: 'Design: custom voice text; no reference audio.'
description: "Design: custom voice text; no reference audio."
examples:
- title: Voice design
request:
Expand Down
10 changes: 6 additions & 4 deletions models/qwen-audio-3.0-tts-plus.yaml
Original file line number Diff line number Diff line change
@@ -1,7 +1,8 @@
schema: neta.generation.model.v1
model: qwen-audio-3.0-tts-plus
title: Qwen Audio 3.0 TTS Plus
description: 'Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points. Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio.'
description: "Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points (design mode: <=200).
Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio."
adapter:
type: openai.audioSpeech
content:
Expand All @@ -10,19 +11,20 @@ content:
required: true
min: 1
max: 1
description: Exactly one non-empty text block to speak, with at least 15 Unicode code points.
description: Exactly one non-empty text block to speak, with at least 15 Unicode code points (at most 200 in
voice-design mode).
- type: audio
required: false
max: 1
sources:
- url
description: 'Clone: one URL; no voice_prompt. Dependency: use prior generated audio.'
description: "Clone: one URL; no voice_prompt. Dependency: use prior generated audio."
meta:
fields:
voice_prompt:
type: string
optional: true
description: 'Design: custom voice text; no reference audio.'
description: "Design: custom voice text; no reference audio."
examples:
- title: Voice design
request:
Expand Down
44 changes: 0 additions & 44 deletions models/qwen-tts.yaml

This file was deleted.

2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@neta-art/generation",
"version": "0.1.31",
"version": "0.2.0",
"description": "A lightweight multimodal generation SDK with built-in model presets and adapter-based provider calls.",
"keywords": [
"ai",
Expand Down
45 changes: 39 additions & 6 deletions src/adapters/audio-speech.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,14 @@ import type {
} from "../types.js";

const REQUEST_TIMEOUT_MS = 210_000;
const QWEN_MODELS = new Set(["qwen-tts", "qwen-audio-3.0-tts-plus", "qwen-audio-3.0-tts-flash"]);
const QWEN_AUDIO_3_MODELS = new Set(["qwen-audio-3.0-tts-plus", "qwen-audio-3.0-tts-flash"]);
const VOICE_ENROLLMENT_MODELS = new Set([
"cosyvoice-v3.5-plus",
"cosyvoice-v3.5-flash",
"qwen-audio-3.0-tts-plus",
"qwen-audio-3.0-tts-flash",
]);
const VOICE_ENROLLMENT_DESIGN_MAX_CODE_POINTS = 200;
const VOICE_ENROLLMENT_VOICE_PROMPT_MAX_CODE_POINTS = 500;
const HIGGS_MODEL = "higgs-tts";

type TextBlock = Extract<GenerationContentBlock, { type: "text" }>;
Expand Down Expand Up @@ -71,7 +77,11 @@ function validateQwen(input: ResolvedGenerationRequest, text: TextBlock, audio:
if (audio.length > 1) {
throw new GenerationValidationError(`${input.declaration.model} supports at most one reference audio`);
}
// Keep accepting the retired preview_text key for compatibility, but never use or forward it.
// Keep accepting the retired preview_text key for compatibility, but never use or forward it:
// the text spoken in the preview/output clip is this request's primary `text` content block
// (validated below against the same 15-200 code point window the worker enforces on its own
// `preview_text` field downstream), not a separate meta value. `preview_text` was how older
// callers passed that text directly; treat any value here as a no-op alias, not live input.
const requestMetaKeys = new Set(["voice_prompt", "preview_text"]);
validateMetaKeys("request.metadata", input.request.metadata, requestMetaKeys);
validateMetaKeys("request.meta", input.request.meta, requestMetaKeys);
Expand All @@ -83,6 +93,19 @@ function validateQwen(input: ResolvedGenerationRequest, text: TextBlock, audio:
if (voicePrompt !== undefined && !hasVoicePrompt) {
throw new GenerationValidationError(`${input.declaration.model} meta.voice_prompt must be a non-empty string`);
}
// Upper-bound checks below (here and on `text`) count the RAW string's code
// points, not the trimmed one: buildPayload sends the untrimmed original
// string (deliberate — see the wire-contract preservation tests), and the
// worker counts with plain len() on what it receives, so validating here
// against anything shorter than what's actually sent could let through a
// value the worker then rejects. Lower-bound/non-empty checks stay on the
// trimmed string — trimmed-length >= a minimum only strengthens the
// guarantee, since trimming can only shorten a string.
if (hasVoicePrompt && Array.from(voicePrompt).length > VOICE_ENROLLMENT_VOICE_PROMPT_MAX_CODE_POINTS) {
throw new GenerationValidationError(
`${input.declaration.model} meta.voice_prompt must be at most ${VOICE_ENROLLMENT_VOICE_PROMPT_MAX_CODE_POINTS} Unicode code points`,
);
}
if (audio.length === 0 && !hasVoicePrompt) {
throw new GenerationValidationError(`${input.declaration.model} requires one reference audio or meta.voice_prompt`);
}
Expand All @@ -92,9 +115,19 @@ function validateQwen(input: ResolvedGenerationRequest, text: TextBlock, audio:
);
}

if (QWEN_AUDIO_3_MODELS.has(input.declaration.model) && Array.from(text.text.trim()).length < 15) {
const codePoints = Array.from(text.text.trim()).length;
if (codePoints < 15) {
throw new GenerationValidationError(`${input.declaration.model} requires input of at least 15 Unicode code points`);
}
// Voice design speaks exactly this text as the preview clip, so the model's
// 15-200 character preview window applies; cloning only feeds the separate
// synthesis call, where long-form text is legitimate. Upper bound counts the
// raw (untrimmed) string sent by buildPayload — see the comment above.
if (hasVoicePrompt && Array.from(text.text).length > VOICE_ENROLLMENT_DESIGN_MAX_CODE_POINTS) {
throw new GenerationValidationError(
`${input.declaration.model} voice design requires input of at most ${VOICE_ENROLLMENT_DESIGN_MAX_CODE_POINTS} Unicode code points`,
);
}
}

function hasOwnWeight(block: AudioBlock): boolean {
Expand Down Expand Up @@ -124,7 +157,7 @@ function validateHiggs(input: ResolvedGenerationRequest, text: TextBlock, audio:

function validateAudioSpeechRequest(input: ResolvedGenerationRequest): void {
const { text, audio } = validateCommonContent(input);
if (QWEN_MODELS.has(input.declaration.model)) {
if (VOICE_ENROLLMENT_MODELS.has(input.declaration.model)) {
validateQwen(input, text, audio);
return;
}
Expand All @@ -148,7 +181,7 @@ function buildPayload(input: ResolvedGenerationRequest): Record<string, unknown>
input: text.text,
};

if (QWEN_MODELS.has(input.declaration.model)) {
if (VOICE_ENROLLMENT_MODELS.has(input.declaration.model)) {
if (audio[0]?.source.type === "url") payload.ref_audio = audio[0].source.url.trim();
else payload.metadata = { voice_prompt: input.meta.voice_prompt };
return payload;
Expand Down
Loading