diff --git a/docs/guides/configuration.md b/docs/guides/configuration.md index d582676e7..86a9fb3a4 100644 --- a/docs/guides/configuration.md +++ b/docs/guides/configuration.md @@ -455,7 +455,7 @@ Use `mode = lines` when: "tools": { "read_file": { "enabled": true, - "mode": "lines", + "mode": "lines", "max_read_file_size": 65536 } } @@ -838,6 +838,9 @@ Legacy Telegram environment variables remain compatible: `PICOCLAW_CHANNELS_TELE Failure behavior is intentionally conservative: if streaming fails before any visible chunk is sent, PicoClaw retries once through the normal `Chat()` path. If a chunk has already been shown to the user, PicoClaw does not send a second non-streaming answer, because that would duplicate visible output. +For model-specific TTS request fields such as custom speech `voice` names or +`response_format: "mp3"`, use `model_list[].extra_body`. + #### Vendor-Specific Examples > **Tip**: You can omit `api_key` fields and store them in `.security.yml` for better security. See [Security Configuration](#-security-configuration-recommended). @@ -856,6 +859,35 @@ Failure behavior is intentionally conservative: if streaming fails before any vi +
+OpenRouter TTS (MAI Voice 2) + +```json +{ + "model_name": "mai-voice-2", + "provider": "openrouter", + "model": "microsoft/mai-voice-2", + "api_base": "https://openrouter.ai/api/v1", + "extra_body": { + "voice": "en-US-Harper:MAI-Voice-2", + "response_format": "mp3" + } + // api_key: set in .security.yml +} +``` + +Pair this with: + +```json +{ + "voice": { + "tts_model_name": "mai-voice-2" + } +} +``` + +
+
VolcEngine (Doubao) diff --git a/docs/guides/providers.md b/docs/guides/providers.md index 40edb87ac..0276a1e52 100644 --- a/docs/guides/providers.md +++ b/docs/guides/providers.md @@ -138,6 +138,9 @@ This design also enables **multi-agent support** with flexible provider selectio When streaming is disabled, omit the `streaming` block. Writing `"streaming": {"enabled": false}` is optional and not needed in generated or hand-written config. +`extra_body` is especially useful for model-specific TTS fields on OpenAI-compatible +speech routes, for example custom `voice` names or `response_format: "mp3"`. + #### Tool Schema Compatibility By default, PicoClaw now forwards tool JSON Schemas unchanged. @@ -205,6 +208,41 @@ If `voice.model_name` is not configured, PicoClaw will continue to fall back to } ``` +#### Voice Synthesis + +You can configure a dedicated text-to-speech model with `voice.tts_model_name`. +When the provider needs model-specific TTS request fields, put them in +`model_list[].extra_body`. + +Example with OpenRouter `microsoft/mai-voice-2`: + +```json +{ + "model_list": [ + { + "model_name": "mai-voice-2", + "provider": "openrouter", + "model": "microsoft/mai-voice-2", + "api_base": "https://openrouter.ai/api/v1", + "extra_body": { + "voice": "en-US-Harper:MAI-Voice-2", + "response_format": "mp3" + }, + "api_keys": ["sk-or-your-openrouter-key"] + } + ], + "voice": { + "tts_model_name": "mai-voice-2" + } +} +``` + +Notes that matter: + +- PicoClaw still uses the OpenAI-compatible `/audio/speech` route for this setup. +- The default TTS request uses `voice: alloy` and `response_format: opus`. +- Override those defaults with `extra_body` when your selected TTS model requires different values. + #### Vendor-Specific Examples **OpenAI** diff --git a/pkg/audio/tts/README.md b/pkg/audio/tts/README.md index ab8491da6..c3fe2d806 100644 --- a/pkg/audio/tts/README.md +++ b/pkg/audio/tts/README.md @@ -25,6 +25,8 @@ Instead: - `voice.tts_model_name` selects a named entry from `model_list`. - That `model_list` entry provides the provider, model ID, API base, and proxy settings. +- For providers that need model-specific TTS parameters, use `model_list[].extra_body` + to pass fields such as `voice` and `response_format`. - `.security.yml` stores the API key for the same named model entry. This is the recommended and supported configuration pattern. @@ -87,6 +89,43 @@ model_list: If you use a custom MiMo endpoint, you can also set `api_base` explicitly. Otherwise PicoClaw will use the provider default. +### Option C: OpenRouter MAI Voice 2 + +Some OpenAI-compatible TTS routes require provider-specific request fields. +OpenRouter's `microsoft/mai-voice-2` is one example: it needs a model-specific +voice name and works best with `response_format: "mp3"`. + +`config.json` + +```json +{ + "voice": { + "tts_model_name": "mai-voice-2" + }, + "model_list": [ + { + "model_name": "mai-voice-2", + "provider": "openrouter", + "model": "microsoft/mai-voice-2", + "api_base": "https://openrouter.ai/api/v1", + "extra_body": { + "voice": "en-US-Harper:MAI-Voice-2", + "response_format": "mp3" + } + } + ] +} +``` + +`.security.yml` + +```yaml +model_list: + mai-voice-2: + api_keys: + - "sk-or-your-openrouter-key" +``` + ## What PicoClaw Sends Today The current TTS runtime uses an OpenAI-compatible speech request with these defaults: @@ -96,11 +135,14 @@ The current TTS runtime uses an OpenAI-compatible speech request with these defa - Voice: `alloy` - Model: taken from the selected `model_list` entry +These defaults can now be overridden per model through `model_list[].extra_body`. + That means: - `openai/tts-1` works naturally. - Other OpenAI-compatible providers can work if they accept the same request format. -- PicoClaw currently does not expose a user-facing config field for changing the TTS voice from `alloy`. +- Provider-specific TTS models may need their own `voice` and `response_format` values. +- If a provider rejects `response_format`, PicoClaw retries once without that field. ## How PicoClaw Chooses a TTS Provider @@ -124,7 +166,8 @@ PicoClaw normalizes the configured base URL for TTS: - Setting `voice.tts_model_name` to a name that does not exist in `model_list`. - Adding a TTS model but forgetting to put its API key in `.security.yml`. -- Assuming PicoClaw will automatically use provider-specific custom voices. +- Assuming PicoClaw will automatically infer provider-specific custom voices. +- Forgetting to set `model_list[].extra_body.voice` or `model_list[].extra_body.response_format` for TTS models that require them. - Using a provider endpoint that is not compatible with the OpenAI `/audio/speech` request format. ## Minimal Checklist diff --git a/pkg/audio/tts/openai_tts.go b/pkg/audio/tts/openai_tts.go index 786414873..651b7b919 100644 --- a/pkg/audio/tts/openai_tts.go +++ b/pkg/audio/tts/openai_tts.go @@ -4,6 +4,7 @@ import ( "bytes" "context" "encoding/json" + "errors" "fmt" "io" "net/http" @@ -16,14 +17,49 @@ import ( ) type OpenAITTSProvider struct { - apiKey string - apiBase string - voice string - model string - httpClient *http.Client + apiKey string + apiBase string + voice string + model string + responseFormat string + httpClient *http.Client +} + +type OpenAITTSOptions struct { + Voice string + ResponseFormat string +} + +type openAITTSAudioStream struct { + io.ReadCloser + fileExt string + contentType string +} + +func (s *openAITTSAudioStream) AudioFileMeta() (string, string) { + return s.fileExt, s.contentType +} + +type openAITTSAPIError struct { + statusCode int + body string +} + +func (e *openAITTSAPIError) Error() string { + return fmt.Sprintf("API error (status %d): %s", e.statusCode, e.body) } func NewOpenAITTSProvider(apiKey string, apiBase string, proxyURL string, model string) *OpenAITTSProvider { + return NewOpenAITTSProviderWithOptions(apiKey, apiBase, proxyURL, model, OpenAITTSOptions{}) +} + +func NewOpenAITTSProviderWithOptions( + apiKey string, + apiBase string, + proxyURL string, + model string, + options OpenAITTSOptions, +) *OpenAITTSProvider { // Normalize apiBase to avoid malformed endpoints like // "https://api.openai.com/audio/speech" when "/v1" is required. if apiBase == "" { @@ -75,12 +111,23 @@ func NewOpenAITTSProvider(apiKey string, apiBase string, proxyURL string, model model = "tts-1" } + voice := strings.TrimSpace(options.Voice) + if voice == "" { + voice = "alloy" + } + + responseFormat := strings.TrimSpace(options.ResponseFormat) + if responseFormat == "" { + responseFormat = "opus" + } + return &OpenAITTSProvider{ - apiKey: apiKey, - apiBase: apiBase, - voice: "alloy", - model: model, - httpClient: client, + apiKey: apiKey, + apiBase: apiBase, + voice: voice, + model: model, + responseFormat: responseFormat, + httpClient: client, } } @@ -91,11 +138,42 @@ func (t *OpenAITTSProvider) Name() string { func (t *OpenAITTSProvider) Synthesize(ctx context.Context, text string) (io.ReadCloser, error) { logger.DebugCF("voice-tts", "Starting TTS synthesis", map[string]any{"text_len": len(text)}) + responseFormat := t.responseFormat + resp, err := t.doSpeechRequest(ctx, text, responseFormat) + if err != nil { + var apiErr *openAITTSAPIError + if errors.As(err, &apiErr) && shouldRetryWithoutResponseFormat(apiErr.body) { + logger.InfoCF("voice-tts", "Retrying TTS without response_format after provider rejection", map[string]any{ + "model": t.model, + }) + responseFormat = "" + resp, err = t.doSpeechRequest(ctx, text, responseFormat) + } + if err != nil { + return nil, err + } + } + + fileExt, contentType := audioFileMetaForResponseFormat(responseFormat) + return &openAITTSAudioStream{ + ReadCloser: resp.Body, + fileExt: fileExt, + contentType: contentType, + }, nil +} + +func (t *OpenAITTSProvider) doSpeechRequest( + ctx context.Context, + text string, + responseFormat string, +) (*http.Response, error) { reqBody := map[string]any{ - "model": t.model, - "input": text, - "voice": t.voice, - "response_format": "opus", + "model": t.model, + "input": text, + "voice": t.voice, + } + if responseFormat != "" { + reqBody["response_format"] = responseFormat } jsonData, err := json.Marshal(reqBody) @@ -119,8 +197,30 @@ func (t *OpenAITTSProvider) Synthesize(ctx context.Context, text string) (io.Rea if resp.StatusCode != http.StatusOK { defer resp.Body.Close() body, _ := io.ReadAll(resp.Body) - return nil, fmt.Errorf("API error (status %d): %s", resp.StatusCode, string(body)) + return nil, &openAITTSAPIError{ + statusCode: resp.StatusCode, + body: string(body), + } } - return resp.Body, nil + return resp, nil +} + +func shouldRetryWithoutResponseFormat(body string) bool { + lower := strings.ToLower(body) + if !strings.Contains(lower, "response_format") { + return false + } + return strings.Contains(lower, "invalid") || strings.Contains(lower, "unsupported") +} + +func audioFileMetaForResponseFormat(responseFormat string) (string, string) { + switch strings.ToLower(strings.TrimSpace(responseFormat)) { + case "", "mp3": + return ".mp3", "audio/mpeg" + case "wav": + return ".wav", "audio/wav" + default: + return ".ogg", "audio/ogg" + } } diff --git a/pkg/audio/tts/tts.go b/pkg/audio/tts/tts.go index 7ae85c8da..202ddc063 100644 --- a/pkg/audio/tts/tts.go +++ b/pkg/audio/tts/tts.go @@ -19,6 +19,10 @@ type TTSProvider interface { Synthesize(ctx context.Context, text string) (io.ReadCloser, error) } +type ttsAudioMetaProvider interface { + AudioFileMeta() (fileExt string, contentType string) +} + func providerFromModelConfig(mc *config.ModelConfig) TTSProvider { if mc == nil || mc.APIKey() == "" { return nil @@ -33,10 +37,31 @@ func providerFromModelConfig(mc *config.ModelConfig) TTSProvider { case "mimo": return NewMimoTTSProvider(mc.APIKey(), providers.ResolveAPIBase(mc), modelID, mc.Proxy) default: - return NewOpenAITTSProvider(mc.APIKey(), providers.ResolveAPIBase(mc), mc.Proxy, modelID) + return NewOpenAITTSProviderWithOptions( + mc.APIKey(), + providers.ResolveAPIBase(mc), + mc.Proxy, + modelID, + openAITTSOptionsFromModelConfig(mc), + ) } } +func openAITTSOptionsFromModelConfig(mc *config.ModelConfig) OpenAITTSOptions { + options := OpenAITTSOptions{} + if mc == nil || mc.ExtraBody == nil { + return options + } + + if voice, ok := mc.ExtraBody["voice"].(string); ok { + options.Voice = strings.TrimSpace(voice) + } + if responseFormat, ok := mc.ExtraBody["response_format"].(string); ok { + options.ResponseFormat = strings.TrimSpace(responseFormat) + } + return options +} + func DetectTTS(cfg *config.Config) TTSProvider { if cfg == nil { return nil @@ -100,6 +125,12 @@ func SynthesizeAndStore( fileExt = ".mp3" contentType = "audio/mpeg" } + if metaProvider, ok := stream.(ttsAudioMetaProvider); ok { + if ext, ct := metaProvider.AudioFileMeta(); ext != "" && ct != "" { + fileExt = ext + contentType = ct + } + } file, err := os.CreateTemp(media.TempDir(), "tts-*"+fileExt) if err != nil { diff --git a/pkg/audio/tts/tts_test.go b/pkg/audio/tts/tts_test.go index 053aa7220..d3e9b8b3c 100644 --- a/pkg/audio/tts/tts_test.go +++ b/pkg/audio/tts/tts_test.go @@ -132,6 +132,64 @@ func TestOpenAITTSProvider_SynthesizeNon200(t *testing.T) { } } +func TestOpenAITTSProvider_SynthesizeRetriesWithoutResponseFormat(t *testing.T) { + var requestBodies []map[string]any + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + bodyBytes, _ := io.ReadAll(r.Body) + _ = r.Body.Close() + + var body map[string]any + _ = json.Unmarshal(bodyBytes, &body) + requestBodies = append(requestBodies, body) + + if len(requestBodies) == 1 { + w.WriteHeader(http.StatusBadRequest) + _, _ = w.Write([]byte(`{"error":"response_format is invalid"}`)) + return + } + + w.WriteHeader(http.StatusOK) + _, _ = w.Write([]byte("audio-bytes")) + })) + defer server.Close() + + provider := NewOpenAITTSProvider("k123", server.URL, "", "x-ai/grok-voice-tts-1.0") + stream, err := provider.Synthesize(context.Background(), "hello") + if err != nil { + t.Fatalf("Synthesize failed: %v", err) + } + defer stream.Close() + + data, err := io.ReadAll(stream) + if err != nil { + t.Fatalf("read stream failed: %v", err) + } + if string(data) != "audio-bytes" { + t.Fatalf("response body mismatch: got %q", string(data)) + } + + if len(requestBodies) != 2 { + t.Fatalf("request count mismatch: got %d, want 2", len(requestBodies)) + } + if requestBodies[0]["response_format"] != "opus" { + t.Fatalf("first request should include opus response_format, got %#v", requestBodies[0]["response_format"]) + } + if _, ok := requestBodies[1]["response_format"]; ok { + t.Fatalf("second request should omit response_format, got %#v", requestBodies[1]["response_format"]) + } + + metaStream, ok := stream.(interface { + AudioFileMeta() (string, string) + }) + if !ok { + t.Fatal("stream does not expose audio metadata") + } + fileExt, contentType := metaStream.AudioFileMeta() + if fileExt != ".mp3" || contentType != "audio/mpeg" { + t.Fatalf("audio metadata mismatch: got (%q, %q)", fileExt, contentType) + } +} + func TestNewOpenAITTSProvider_UsesConfiguredModel(t *testing.T) { t.Parallel() @@ -144,6 +202,27 @@ func TestNewOpenAITTSProvider_UsesConfiguredModel(t *testing.T) { } } +func TestNewOpenAITTSProvider_UsesConfiguredVoiceAndResponseFormat(t *testing.T) { + t.Parallel() + + provider := NewOpenAITTSProviderWithOptions( + "key", + "https://openrouter.ai/api/v1", + "", + "microsoft/mai-voice-2", + OpenAITTSOptions{ + Voice: "en-US-Harper:MAI-Voice-2", + ResponseFormat: "mp3", + }, + ) + if provider.voice != "en-US-Harper:MAI-Voice-2" { + t.Fatalf("voice mismatch: got %q", provider.voice) + } + if provider.responseFormat != "mp3" { + t.Fatalf("responseFormat mismatch: got %q", provider.responseFormat) + } +} + func TestDetectTTS_UsesMimoProviderForMimoModels(t *testing.T) { t.Parallel() @@ -170,6 +249,36 @@ func TestDetectTTS_UsesMimoProviderForMimoModels(t *testing.T) { } } +func TestDetectTTS_UsesOpenAIExtraBodyVoiceAndResponseFormat(t *testing.T) { + t.Parallel() + + provider := DetectTTS(&config.Config{ + Voice: config.VoiceConfig{TTSModelName: "mai-voice"}, + ModelList: []*config.ModelConfig{ + { + ModelName: "mai-voice", + Model: "openrouter/microsoft/mai-voice-2", + APIKeys: config.SimpleSecureStrings("sk-openrouter"), + ExtraBody: map[string]any{ + "voice": "en-US-Harper:MAI-Voice-2", + "response_format": "mp3", + }, + }, + }, + }) + + ttsProvider, ok := provider.(*OpenAITTSProvider) + if !ok { + t.Fatalf("DetectTTS() type = %T, want *OpenAITTSProvider", provider) + } + if ttsProvider.voice != "en-US-Harper:MAI-Voice-2" { + t.Fatalf("voice mismatch: got %q", ttsProvider.voice) + } + if ttsProvider.responseFormat != "mp3" { + t.Fatalf("responseFormat mismatch: got %q", ttsProvider.responseFormat) + } +} + type stubTTSProvider struct { name string } @@ -245,3 +354,51 @@ func TestSynthesizeAndStore_UsesMp3MetadataForMimo(t *testing.T) { t.Fatalf("filename extension = %q, want %q", filepath.Ext(meta.Filename), ".mp3") } } + +func TestSynthesizeAndStore_UsesStreamProvidedAudioMetadata(t *testing.T) { + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + bodyBytes, _ := io.ReadAll(r.Body) + _ = r.Body.Close() + + var body map[string]any + _ = json.Unmarshal(bodyBytes, &body) + if body["response_format"] == "opus" { + w.WriteHeader(http.StatusBadRequest) + _, _ = w.Write([]byte(`{"error":"response_format is invalid"}`)) + return + } + + w.WriteHeader(http.StatusOK) + _, _ = w.Write([]byte("mp3-audio")) + })) + defer server.Close() + + store := media.NewFileMediaStore() + provider := NewOpenAITTSProvider("k123", server.URL, "", "x-ai/grok-voice-tts-1.0") + ref, err := SynthesizeAndStore( + context.Background(), + provider, + store, + "hello", + "", + "telegram", + "chat123", + ) + if err != nil { + t.Fatalf("SynthesizeAndStore failed: %v", err) + } + + path, meta, err := store.ResolveWithMeta(ref) + if err != nil { + t.Fatalf("ResolveWithMeta failed: %v", err) + } + if meta.ContentType != "audio/mpeg" { + t.Fatalf("ContentType = %q, want %q", meta.ContentType, "audio/mpeg") + } + if filepath.Ext(path) != ".mp3" { + t.Fatalf("stored file extension = %q, want %q", filepath.Ext(path), ".mp3") + } + if filepath.Ext(meta.Filename) != ".mp3" { + t.Fatalf("filename extension = %q, want %q", filepath.Ext(meta.Filename), ".mp3") + } +} diff --git a/workspace/skills/picoclaw-agent/SKILL.md b/workspace/skills/picoclaw-agent/SKILL.md index 32797d09d..dd7dc037d 100644 --- a/workspace/skills/picoclaw-agent/SKILL.md +++ b/workspace/skills/picoclaw-agent/SKILL.md @@ -468,7 +468,8 @@ Recommended setup: 1. add a TTS-capable model entry to `model_list` 2. set `voice.tts_model_name` to that entry's `model_name` 3. store the API key in `.security.yml` -4. enable `send_tts` in the tool configuration if you want the agent to emit speech files +4. if the provider needs model-specific TTS fields, add them under `model_list[].extra_body` +5. enable `send_tts` in the tool configuration if you want the agent to emit speech files Example: @@ -493,6 +494,35 @@ model_list: - "sk-openai-your-key" ``` +Example with OpenRouter MAI Voice 2: + +```json +{ + "model_list": [ + { + "model_name": "mai-voice-2", + "provider": "openrouter", + "model": "microsoft/mai-voice-2", + "api_base": "https://openrouter.ai/api/v1", + "extra_body": { + "voice": "en-US-Harper:MAI-Voice-2", + "response_format": "mp3" + } + } + ], + "voice": { + "tts_model_name": "mai-voice-2" + } +} +``` + +```yaml +model_list: + mai-voice-2: + api_keys: + - "sk-or-your-openrouter-key" +``` + ### Current TTS Provider Paths | Provider path | Example model | Notes | @@ -505,6 +535,8 @@ Operational notes: - the preferred selection path is `voice.tts_model_name` - if that is missing, PicoClaw can still scan `model_list` for the first API-backed model whose ID contains `tts` - the current OpenAI-style TTS request defaults to `voice: alloy` and `response_format: opus` +- you can override `voice` and `response_format` for a specific TTS model through `model_list[].extra_body` +- if a provider rejects `response_format`, PicoClaw retries once without that field - `send_tts` is only registered when TTS detection succeeds