Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 9 additions & 11 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ Agents and external tools should inspect a model declaration before constructing
import { createGenerationClient } from "@neta-art/generation";

const discoveryClient = createGenerationClient();
const declaration = discoveryClient.getModel("qwen-tts");
const declaration = discoveryClient.getModel("qwen-audio-3.1-tts-flash");
if (!declaration) throw new Error("Model is unavailable");

console.log(discoveryClient.stringifyModelConfig(declaration.model, { format: "json" }));
Expand All @@ -84,7 +84,7 @@ The same declarations can be exported as YAML through the existing CLI:

```bash
neta-generation models list
neta-generation models export qwen-tts --out ./qwen-tts.yaml
neta-generation models export qwen-audio-3.1-tts-flash --out ./qwen-audio-3.1-tts-flash.yaml
neta-generation models export-all --out ./models
```

Expand Down Expand Up @@ -170,9 +170,7 @@ const client = createGenerationClient({
- `gpt-image-2`
- `z-image-turbo`
- `qwen-image-edit`
- `qwen-tts`
- `qwen-audio-3.0-tts-plus`
- `qwen-audio-3.0-tts-flash`
- `qwen-audio-3.1-tts-flash`
- `higgs-tts`
- `gemini-3.1-flash-image-preview`
- `kling-text-to-video`
Expand Down Expand Up @@ -268,29 +266,29 @@ Each TTS request accepts exactly one non-empty text block and returns one URL au

| Requirement | Model choice |
| --- | --- |
| Create a voice from a text-only description, without reference audio | Use an explicitly requested Qwen variant; otherwise use `qwen-tts` as the deterministic default |
| Create a voice from a text-only description, without reference audio | `qwen-audio-3.1-tts-flash` |
| Maximize fidelity to one reference voice | `higgs-tts` |
| Blend 2-16 weighted reference voices | `higgs-tts` |
| Use a default voice, including a delegated choice expressed only as any, random, suitable, or natural | `higgs-tts` |

- Qwen: `voice_prompt` design OR one-reference clone; `qwen-tts` is the unspecified-design default and accepts any text length; Plus / Flash require at least 15 Unicode code points.
- Qwen: `voice_prompt` design OR one-reference clone via `qwen-audio-3.1-tts-flash`; requires at least 15 Unicode code points.
- Higgs: delegated default voice, high-fidelity one-reference clone, or weighted 2-16-reference blend.
- Conflict: reference + redesign requires user choice before generation.
- Blend: all references, full text, one request.
- Dependency: clone prior generated audio.
- Ranking: no declared Qwen quality, latency, or cost order.
- Short text: input under 15 Unicode code points has no voice-design path on Qwen; ask the user to lengthen it, or use `higgs-tts` with a default/reference voice instead.

```ts
await client.generate({
model: "qwen-tts",
content: [{ type: "text", text: "欢迎使用语音合成功能。" }],
model: "qwen-audio-3.1-tts-flash",
content: [{ type: "text", text: "欢迎使用语音合成功能,这是一段示例文本。" }],
meta: {
voice_prompt: "一位沉稳自然的中文播音员,吐字清晰,语速适中",
},
});

await client.generate({
model: "qwen-audio-3.0-tts-flash",
model: "qwen-audio-3.1-tts-flash",
content: [
{ type: "text", text: "这是一段长度足够并且表达清晰自然的语音合成文本。" },
{ type: "audio", source: { type: "url", url: "https://example.com/reference.mp3" } },
Expand Down
2 changes: 1 addition & 1 deletion examples/text-to-speech.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ if (!apiKey) throw new Error("Set NETA_ROUTER_API_KEY or NETA_API_KEY");

const client = createGenerationClient({ apiKey });
const output = await client.generate({
model: "qwen-tts",
model: "qwen-audio-3.1-tts-flash",
content: [{ type: "text", text: "欢迎使用语音合成功能,这是一段示例文本。" }],
meta: {
voice_prompt: "一位沉稳自然的中文播音员,吐字清晰,语速适中",
Expand Down
44 changes: 0 additions & 44 deletions models/qwen-audio-3.0-tts-plus.yaml

This file was deleted.

Original file line number Diff line number Diff line change
@@ -1,7 +1,8 @@
schema: neta.generation.model.v1
model: qwen-audio-3.0-tts-flash
title: Qwen Audio 3.0 TTS Flash
description: 'Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points. Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio.'
model: qwen-audio-3.1-tts-flash
title: Qwen Audio 3.1 TTS Flash
description: "Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points. Conflict: ask user;
never combine/reinterpret. Dependency: clone prior generated audio."
adapter:
type: openai.audioSpeech
content:
Expand All @@ -16,25 +17,25 @@ content:
max: 1
sources:
- url
description: 'Clone: one URL; no voice_prompt. Dependency: use prior generated audio.'
description: "Clone: one URL; no voice_prompt. Dependency: use prior generated audio."
meta:
fields:
voice_prompt:
type: string
optional: true
description: 'Design: custom voice text; no reference audio.'
description: "Design: custom voice text; no reference audio."
examples:
- title: Voice design
request:
model: qwen-audio-3.0-tts-flash
model: qwen-audio-3.1-tts-flash
content:
- type: text
text: 这是一段长度足够并且表达清晰自然的语音合成测试文本。
meta:
voice_prompt: 一位沉稳干练的男性播音员声音,吐字清晰有力
- title: Voice clone
request:
model: qwen-audio-3.0-tts-flash
model: qwen-audio-3.1-tts-flash
content:
- type: text
text: 这是一段长度足够并且表达清晰自然的语音合成测试文本。
Expand Down
44 changes: 0 additions & 44 deletions models/qwen-tts.yaml

This file was deleted.

2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@neta-art/generation",
"version": "0.1.31",
"version": "0.2.0",
"description": "A lightweight multimodal generation SDK with built-in model presets and adapter-based provider calls.",
"keywords": [
"ai",
Expand Down
9 changes: 4 additions & 5 deletions src/adapters/audio-speech.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,7 @@ import type {
} from "../types.js";

const REQUEST_TIMEOUT_MS = 210_000;
const QWEN_MODELS = new Set(["qwen-tts", "qwen-audio-3.0-tts-plus", "qwen-audio-3.0-tts-flash"]);
const QWEN_AUDIO_3_MODELS = new Set(["qwen-audio-3.0-tts-plus", "qwen-audio-3.0-tts-flash"]);
const QWEN_MODEL = "qwen-audio-3.1-tts-flash";
const HIGGS_MODEL = "higgs-tts";

type TextBlock = Extract<GenerationContentBlock, { type: "text" }>;
Expand Down Expand Up @@ -92,7 +91,7 @@ function validateQwen(input: ResolvedGenerationRequest, text: TextBlock, audio:
);
}

if (QWEN_AUDIO_3_MODELS.has(input.declaration.model) && Array.from(text.text.trim()).length < 15) {
if (Array.from(text.text.trim()).length < 15) {
throw new GenerationValidationError(`${input.declaration.model} requires input of at least 15 Unicode code points`);
}
}
Expand Down Expand Up @@ -124,7 +123,7 @@ function validateHiggs(input: ResolvedGenerationRequest, text: TextBlock, audio:

function validateAudioSpeechRequest(input: ResolvedGenerationRequest): void {
const { text, audio } = validateCommonContent(input);
if (QWEN_MODELS.has(input.declaration.model)) {
if (input.declaration.model === QWEN_MODEL) {
validateQwen(input, text, audio);
return;
}
Expand All @@ -148,7 +147,7 @@ function buildPayload(input: ResolvedGenerationRequest): Record<string, unknown>
input: text.text,
};

if (QWEN_MODELS.has(input.declaration.model)) {
if (input.declaration.model === QWEN_MODEL) {
if (audio[0]?.source.type === "url") payload.ref_audio = audio[0].source.url.trim();
else payload.metadata = { voice_prompt: input.meta.voice_prompt };
return payload;
Expand Down
53 changes: 12 additions & 41 deletions src/builtins.ts
Original file line number Diff line number Diff line change
Expand Up @@ -708,20 +708,13 @@ function geminiImageModel(
};
}

function qwenTtsModel(
model: string,
title: string,
description: string,
options: { minimumTextCodePoints?: number } = {},
): GenerationModelDeclaration {
const text = options.minimumTextCodePoints
? "这是一段长度足够并且表达清晰自然的语音合成测试文本。"
: "这是一次清晰自然的语音合成测试。";
return {
const audioSpeechModels = [
{
schema: MODEL_SCHEMA,
model,
title,
description,
model: "qwen-audio-3.1-tts-flash",
title: "Qwen Audio 3.1 TTS Flash",
description:
"Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points. Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio.",
adapter: { type: "openai.audioSpeech" },
content: {
input: [
Expand All @@ -730,9 +723,7 @@ function qwenTtsModel(
required: true,
min: 1,
max: 1,
description: options.minimumTextCodePoints
? `Exactly one non-empty text block to speak, with at least ${options.minimumTextCodePoints} Unicode code points.`
: "Exactly one non-empty text block to speak.",
description: "Exactly one non-empty text block to speak, with at least 15 Unicode code points.",
},
{
type: "audio",
Expand All @@ -756,43 +747,23 @@ function qwenTtsModel(
{
title: "Voice design",
request: {
model,
content: [{ type: "text", text }],
model: "qwen-audio-3.1-tts-flash",
content: [{ type: "text", text: "这是一段长度足够并且表达清晰自然的语音合成测试文本。" }],
meta: { voice_prompt: "一位沉稳干练的男性播音员声音,吐字清晰有力" },
},
},
{
title: "Voice clone",
request: {
model,
model: "qwen-audio-3.1-tts-flash",
content: [
{ type: "text", text },
{ type: "text", text: "这是一段长度足够并且表达清晰自然的语音合成测试文本。" },
{ type: "audio", source: { type: "url", url: "https://example.com/reference.mp3" } },
],
},
},
],
};
}

const audioSpeechModels = [
qwenTtsModel(
"qwen-tts",
"Qwen TTS",
"Modes: voice_prompt design OR one-reference clone. Default: unspecified Qwen design. Text: any length. Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio.",
),
qwenTtsModel(
"qwen-audio-3.0-tts-plus",
"Qwen Audio 3.0 TTS Plus",
"Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points. Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio.",
{ minimumTextCodePoints: 15 },
),
qwenTtsModel(
"qwen-audio-3.0-tts-flash",
"Qwen Audio 3.0 TTS Flash",
"Modes: voice_prompt design OR one-reference clone. Text: >=15 Unicode code points. Conflict: ask user; never combine/reinterpret. Dependency: clone prior generated audio.",
{ minimumTextCodePoints: 15 },
),
},
{
schema: MODEL_SCHEMA,
model: "higgs-tts",
Expand Down
Loading