Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
168 changes: 98 additions & 70 deletions .speakeasy/gen.lock

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .speakeasy/gen.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ generation:
documentation: mintlify
preApplyUnionDiscriminators: true
python:
version: 1.3.5
version: 1.3.6
additionalDependencies:
dev: {}
main: {}
Expand Down
141 changes: 133 additions & 8 deletions .speakeasy/out.openapi.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -6252,6 +6252,7 @@ components:
- 'deepseek'
- 'dekallm'
- 'digitalocean'
- 'elevenlabs'
- 'featherless'
- 'fireworks'
- 'fish-audio'
Expand Down Expand Up @@ -17441,6 +17442,7 @@ components:
- 'DeepSeek'
- 'DekaLLM'
- 'DigitalOcean'
- 'ElevenLabs'
- 'Featherless'
- 'Fireworks'
- 'Fish Audio'
Expand Down Expand Up @@ -24972,6 +24974,7 @@ components:
- 'DeepSeek'
- 'DekaLLM'
- 'DigitalOcean'
- 'ElevenLabs'
- 'Featherless'
- 'Fireworks'
- 'Fish Audio'
Expand Down Expand Up @@ -25186,6 +25189,9 @@ components:
digitalocean:
additionalProperties: {}
type: 'object'
elevenlabs:
additionalProperties: {}
type: 'object'
enfer:
additionalProperties: {}
type: 'object'
Expand Down Expand Up @@ -25750,6 +25756,7 @@ components:
- 'DeepSeek'
- 'DekaLLM'
- 'DigitalOcean'
- 'ElevenLabs'
- 'Featherless'
- 'Fireworks'
- 'Fish Audio'
Expand Down Expand Up @@ -28518,8 +28525,38 @@ components:
- 111
logprob: -0.5
token: 'Hello'
STTInputAudio:
description: 'Base64-encoded audio to transcribe'
STTEntity:
description: 'A detected entity, returned when the provider runs entity detection'
example:
end_char: 25
start_char: 15
text: 'John Smith'
type: 'name'
properties:
end_char:
description: 'Zero-based exclusive character offset of the entity end within the response-level text (not seconds)'
example: 25
type: 'integer'
start_char:
description: 'Zero-based character offset of the entity start within the response-level text (not seconds)'
example: 15
type: 'integer'
text:
description: 'Entity text as it appears in the transcript'
type: 'string'
type:
description: 'Provider entity type label'
example: 'name'
type: 'string'
required:
- 'text'
- 'type'
- 'start_char'
- 'end_char'
type: 'object'
STTInlineInputAudio:
additionalProperties: false
description: 'Inline base64 audio input for speech-to-text'
example:
data: 'UklGRiQA...'
format: 'wav'
Expand All @@ -28528,24 +28565,44 @@ components:
description: 'Base64-encoded audio data (raw bytes, not a data URI)'
type: 'string'
format:
description: 'Audio format (e.g., wav, mp3, flac, m4a, ogg, webm, aac). Supported formats vary by provider.'
description: 'Audio format (e.g., wav, mp3, flac, m4a, ogg, webm, aac). Supported formats vary by provider. "pcm" means headerless signed 16-bit little-endian mono audio at 16 kHz.'
pattern: '^[a-zA-Z0-9][a-zA-Z0-9+._-]{0,15}$'
type: 'string'
required:
- 'data'
- 'format'
type: 'object'
STTInputAudio:
anyOf:
- $ref: '#/components/schemas/STTInlineInputAudio'
- $ref: '#/components/schemas/STTUrlInputAudio'
description: 'Audio to transcribe: inline base64 bytes, or a URL the provider downloads directly.'
STTRequest:
description: 'Speech-to-text request input. Accepts a JSON body with input_audio containing base64-encoded audio.'
description: 'Speech-to-text request input. Accepts a JSON body with input_audio containing base64-encoded audio or a URL the provider downloads.'
example:
input_audio:
data: 'UklGRiQA...'
format: 'wav'
language: 'en'
model: 'openai/whisper-large-v3'
properties:
diarize:
description: 'Label each word with the speaker who said it. Speaker labels are returned on the words array (speaker, speaker_label), so response_format must be "verbose_json" (a "json" request is rejected with a 400) and word timestamps are included even when timestamp_granularities omits "word". Only supported by some providers; the request is rejected with a 400 when the selected model cannot diarize. Providers may charge extra.'
example: true
type: 'boolean'
input_audio:
$ref: '#/components/schemas/STTInputAudio'
keyterms:
description: 'Domain terms, names, or phrases to bias recognition toward. Only supported by some providers; the request is rejected with a 400 when the selected model cannot use keyterms. Providers may cap the number of terms or characters per term and may charge extra.'
example:
- 'OpenRouter'
- 'Scribe'
items:
maxLength: 100
minLength: 1
type: 'string'
maxItems: 1000
type: 'array'
language:
description: 'ISO-639-1 language code (e.g., "en", "ja"). Auto-detected if omitted.'
example: 'en'
Expand Down Expand Up @@ -28617,10 +28674,20 @@ components:
example: 9.2
format: 'double'
type: 'number'
entities:
description: 'Detected entities with character offsets into text, present when the provider runs entity detection'
items:
$ref: '#/components/schemas/STTEntity'
type: 'array'
language:
description: 'Detected or forced language, present when response_format is verbose_json'
example: 'english'
type: 'string'
language_confidence:
description: 'Provider confidence in the detected language from 0 to 1, present when response_format is verbose_json and the provider scores language detection'
example: 0.98
format: 'double'
type: 'number'
segments:
description: 'Timestamped transcript segments, present when response_format is verbose_json'
items:
Expand Down Expand Up @@ -28666,6 +28733,10 @@ components:
description: 'Average log probability of the segment'
format: 'double'
type: 'number'
channel:
description: 'Zero-based audio channel index for the segment, present when the provider transcribes channels separately'
example: 0
type: 'integer'
compression_ratio:
description: 'Compression ratio of the segment'
format: 'double'
Expand All @@ -28691,6 +28762,10 @@ components:
description: 'Speaker index for the segment, present when the provider returns diarization data'
example: 0
type: 'integer'
speaker_label:
description: 'Provider speaker label for the segment, present when the provider labels speakers with a string'
example: 'speaker_0'
type: 'string'
start:
description: 'Segment start time in seconds'
example: 0
Expand Down Expand Up @@ -28723,6 +28798,25 @@ components:
example: 'word'
type: 'string'
x-speakeasy-unknown-values: allow
STTUrlInputAudio:
additionalProperties: false
description: 'Audio input fetched by the provider from a URL'
example:
format: 'mp3'
url: 'https://example.com/meeting.mp3'
properties:
format:
description: 'Audio format of the file at the URL. Defaults to the extension of the URL path; required when the path has no extension.'
pattern: '^[a-zA-Z0-9][a-zA-Z0-9+._-]{0,15}$'
type: 'string'
url:
description: 'Publicly reachable http(s) URL of the audio file. The provider downloads it directly, so the inline upload size limit does not apply. Only supported by some providers.'
format: 'uri'
maxLength: 8000
type: 'string'
required:
- 'url'
type: 'object'
STTUsage:
description: 'Aggregated usage statistics for the request'
example:
Expand Down Expand Up @@ -28764,6 +28858,10 @@ components:
start: 0
word: 'Hello'
properties:
channel:
description: 'Zero-based audio channel index for the word, present when the provider transcribes channels separately'
example: 0
type: 'integer'
confidence:
description: 'Provider confidence for the word from 0 to 1, present when the provider returns per-word confidence'
example: 0.98
Expand All @@ -28778,13 +28876,25 @@ components:
description: 'Speaker index for the word, present when the provider returns diarization data'
example: 0
type: 'integer'
speaker_label:
description: 'Provider speaker label for the word, present when the provider labels speakers with a string'
example: 'speaker_0'
type: 'string'
start:
description: 'Word start time in seconds'
example: 0
format: 'double'
type: 'number'
type:
description: 'Kind of entry; omitted or "word" for spoken words, "audio_event" for non-speech sounds the provider tags with timestamps'
enum:
- 'word'
- 'audio_event'
example: 'word'
type: 'string'
x-speakeasy-unknown-values: allow
word:
description: 'The transcribed word'
description: 'The transcribed word, or the event tag such as "(laughter)" when type is audio_event'
example: 'Hello'
type: 'string'
required:
Expand Down Expand Up @@ -33201,7 +33311,7 @@ paths:
- $ref: "#/components/parameters/AppCategories"
/audio/transcriptions:
post:
description: 'Transcribes audio into text. Accepts base64-encoded audio input as JSON or an OpenAI-style multipart/form-data file upload, and returns the transcribed text.'
description: 'Transcribes audio into text. Accepts base64-encoded audio input as JSON, an OpenAI-style multipart/form-data file upload, or a URL the provider downloads directly, and returns the transcribed text.'
operationId: 'createAudioTranscriptions'
requestBody:
content:
Expand All @@ -33221,16 +33331,27 @@ paths:
model: 'openai/whisper-large-v3'
schema:
properties:
diarize:
description: 'Label each word with the speaker who said it (words[].speaker, words[].speaker_label). Requires response_format "verbose_json" (400 otherwise); word timestamps are included even when timestamp_granularities[] omits "word". Only supported by some providers; 400 when the selected model cannot diarize.'
type: 'boolean'
file:
description: 'The audio file to transcribe. The format is derived from the filename extension or the file part content type. Max 25 MB; send larger files as base64 JSON via input_audio.'
description: 'The audio file to transcribe. The format is derived from the filename extension or the file part content type. Max 25 MB; send larger files as base64 JSON via input_audio, or by URL via source_url. Exactly one of file or source_url is required.'
format: 'binary'
type: 'string'
keyterms[]:
description: 'Domain terms, names, or phrases to bias recognition toward; repeat the part once per term (keyterms=... is also accepted). Only supported by some providers; 400 when the selected model cannot use keyterms.'
items:
type: 'string'
type: 'array'
language:
description: 'The language of the input audio (ISO-639-1).'
type: 'string'
model:
description: 'The model to use for transcription.'
type: 'string'
provider:
description: 'JSON-encoded provider preferences object, the same shape as the JSON body field: { "options": { "<provider-slug>": { ... } } }. Only options for the matched provider are forwarded. Must decode to a JSON object.'
type: 'string'
response_format:
description: 'The response format. "json" (default) returns { text, usage }; "verbose_json" additionally returns task, language, duration, and segment-level timestamps (OpenAI-compatible providers only).'
enum:
Expand All @@ -33242,6 +33363,10 @@ paths:
description: 'A unique identifier for grouping related requests (e.g., a conversation or agent workflow). Used for observability grouping in Broadcast and private logging; never sent to the provider. If provided in both the request body and the x-session-id header, the body value takes precedence.'
maxLength: 256
type: 'string'
source_url:
description: 'Publicly reachable http(s) URL of the audio file, downloaded by the provider directly (no size limit on our side). The format is derived from the URL path extension. Only supported by some providers; exactly one of file or source_url is required.'
format: 'uri'
type: 'string'
temperature:
description: 'The sampling temperature.'
type: 'number'
Expand All @@ -33262,7 +33387,6 @@ paths:
maxLength: 256
type: 'string'
required:
- 'file'
- 'model'
type: 'object'
required: true
Expand Down Expand Up @@ -34461,6 +34585,7 @@ paths:
- 'deepseek'
- 'dekallm'
- 'digitalocean'
- 'elevenlabs'
- 'featherless'
- 'fireworks'
- 'fish-audio'
Expand Down
10 changes: 5 additions & 5 deletions .speakeasy/workflow.lock
Original file line number Diff line number Diff line change
Expand Up @@ -2,19 +2,19 @@ speakeasyVersion: 1.787.0
sources:
OpenRouter API:
sourceNamespace: open-router-chat-completions-api
sourceRevisionDigest: sha256:a3666c04ea7f6d0bd1851a3961f51c8fef2cee3f8c6b1ded741777d1f2ebc088
sourceBlobDigest: sha256:007c9b444c48e255ebb0cf0711d3c99659dee78aef07ec978d938635c05bd1cf
sourceRevisionDigest: sha256:fa682f646ac7fac9fd7baeae8f6356547792c4e8f0c14d7b471bb055db9e47e3
sourceBlobDigest: sha256:5caef8a5438f249e60312655dc3aa58171d47772f0a3b24f23a745a84162bc82
tags:
- latest
- 1.0.0
targets:
open-router:
source: OpenRouter API
sourceNamespace: open-router-chat-completions-api
sourceRevisionDigest: sha256:a3666c04ea7f6d0bd1851a3961f51c8fef2cee3f8c6b1ded741777d1f2ebc088
sourceBlobDigest: sha256:007c9b444c48e255ebb0cf0711d3c99659dee78aef07ec978d938635c05bd1cf
sourceRevisionDigest: sha256:fa682f646ac7fac9fd7baeae8f6356547792c4e8f0c14d7b471bb055db9e47e3
sourceBlobDigest: sha256:5caef8a5438f249e60312655dc3aa58171d47772f0a3b24f23a745a84162bc82
codeSamplesNamespace: open-router-python-code-samples
codeSamplesRevisionDigest: sha256:2f4bce7e2782ef7bc1a43785675f95b624c865c6c6cbc4989ea78802b2371e83
codeSamplesRevisionDigest: sha256:fa882c2f6a037169bc675adc4611202adcf873d3290137352cf5de5be043495f
workflow:
workflowVersion: 1.0.0
speakeasyVersion: 1.787.0
Expand Down
4 changes: 2 additions & 2 deletions README-PYPI.md
Original file line number Diff line number Diff line change
Expand Up @@ -276,10 +276,10 @@ with OpenRouter(
api_key=os.getenv("OPENROUTER_API_KEY", ""),
) as open_router:

res = open_router.stt.create_transcription_multipart(file={
res = open_router.stt.create_transcription_multipart(model="openai/whisper-large-v3", file={
"file_name": "example.file",
"content": open("example.file", "rb"),
}, model="openai/whisper-large-v3", language="en")
}, language="en")

# Handle response
print(res)
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -276,10 +276,10 @@ with OpenRouter(
api_key=os.getenv("OPENROUTER_API_KEY", ""),
) as open_router:

res = open_router.stt.create_transcription_multipart(file={
res = open_router.stt.create_transcription_multipart(model="openai/whisper-large-v3", file={
"file_name": "example.file",
"content": open("example.file", "rb"),
}, model="openai/whisper-large-v3", language="en")
}, language="en")

# Handle response
print(res)
Expand Down
12 changes: 11 additions & 1 deletion RELEASES.md
Original file line number Diff line number Diff line change
Expand Up @@ -2819,4 +2819,14 @@ Based on:
### Generated
- [python v1.3.5] .
### Releases
- [PyPI v1.3.5] https://pypi.org/project/openrouter/1.3.5 - .
- [PyPI v1.3.5] https://pypi.org/project/openrouter/1.3.5 - .

## 2026-09-29 19:06:13
### Changes
Based on:
- OpenAPI Doc
- Speakeasy CLI 1.787.0 (2.914.0) https://github.com/speakeasy-api/speakeasy
### Generated
- [python v1.3.6] .
### Releases
- [PyPI v1.3.6] https://pypi.org/project/openrouter/1.3.6 - .
1 change: 1 addition & 0 deletions docs/components/byokproviderslug.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,7 @@ This is an open enum. Unrecognized values will not fail type checks.
- `"deepseek"`
- `"dekallm"`
- `"digitalocean"`
- `"elevenlabs"`
- `"featherless"`
- `"fireworks"`
- `"fish-audio"`
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,7 @@ Provider-specific options keyed by provider slug. Only options for the matched p
| `deepseek` | Dict[str, *Any*] | :heavy_minus_sign: | N/A |
| `dekallm` | Dict[str, *Any*] | :heavy_minus_sign: | N/A |
| `digitalocean` | Dict[str, *Any*] | :heavy_minus_sign: | N/A |
| `elevenlabs` | Dict[str, *Any*] | :heavy_minus_sign: | N/A |
| `enfer` | Dict[str, *Any*] | :heavy_minus_sign: | N/A |
| `fake_provider` | Dict[str, *Any*] | :heavy_minus_sign: | N/A |
| `featherless` | Dict[str, *Any*] | :heavy_minus_sign: | N/A |
Expand Down
Loading
Loading