Matériel à 10$ · 10 Mo de RAM · Démarrage en 1s · 皮皮虾,我们走!
+
Matériel à $10 · <10 Mo de RAM · Démarrage en <1s · 皮皮虾,我们走!
-
-
+
+
@@ -18,14 +18,17 @@
- [中文](README.zh.md) | [日本語](README.ja.md) | [Português](README.pt-br.md) | [Tiếng Việt](README.vi.md) | [English](README.md) | **Français**
+[中文](README.zh.md) | [日本語](README.ja.md) | [Português](README.pt-br.md) | [Tiếng Việt](README.vi.md) | **Français** | [Italiano](README.it.md) | [Bahasa Indonesia](README.id.md) | [English](README.md)
+
---
-🦐 **PicoClaw** est un assistant personnel IA ultra-léger inspiré de [nanobot](https://github.com/HKUDS/nanobot), entièrement réécrit en **Go** via un processus d'auto-amorçage (self-bootstrapping) — où l'agent IA lui-même a piloté l'intégralité de la migration architecturale et de l'optimisation du code.
+> **PicoClaw** est un projet open-source indépendant initié par [Sipeed](https://sipeed.com). Il est entièrement écrit en **Go** — ce n'est pas un fork d'OpenClaw, de NanoBot ou de tout autre projet.
-⚡️ **Extrêmement léger :** Fonctionne sur du matériel à seulement **10$** avec **<10 Mo** de RAM. C'est 99% de mémoire en moins qu'OpenClaw et 98% moins cher qu'un Mac mini !
+🦐 **PicoClaw** est un assistant personnel IA ultra-léger inspiré de [NanoBot](https://github.com/HKUDS/nanobot), entièrement réécrit en **Go** via un processus d'auto-amorçage (self-bootstrapping) — où l'agent IA lui-même a piloté l'intégralité de la migration architecturale et de l'optimisation du code.
+
+⚡️ **Extrêmement léger :** Fonctionne sur du matériel à seulement **$10** avec **<10 Mo** de RAM. C'est 99% de mémoire en moins qu'OpenClaw et 98% moins cher qu'un Mac mini !
@@ -46,42 +49,64 @@
> **🚨 SÉCURITÉ & CANAUX OFFICIELS**
>
> * **PAS DE CRYPTO :** PicoClaw n'a **AUCUN** token/jeton officiel. Toute annonce sur `pump.fun` ou d'autres plateformes de trading est une **ARNAQUE**.
+>
> * **DOMAINE OFFICIEL :** Le **SEUL** site officiel est **[picoclaw.io](https://picoclaw.io)**, et le site de l'entreprise est **[sipeed.com](https://sipeed.com)**.
-> * **Attention :** De nombreux domaines `.ai/.org/.com/.net/...` sont enregistrés par des tiers et ne nous appartiennent pas.
+> * **Attention :** De nombreux domaines `.ai/.org/.com/.net/...` sont enregistrés par des tiers.
> * **Attention :** PicoClaw est en phase de développement précoce et peut présenter des problèmes de sécurité réseau non résolus. Ne déployez pas en environnement de production avant la version v1.0.
> * **Note :** PicoClaw a récemment fusionné de nombreuses PR, ce qui peut entraîner une empreinte mémoire plus importante (10–20 Mo) dans les dernières versions. Nous prévoyons de prioriser l'optimisation des ressources dès que l'ensemble des fonctionnalités sera stabilisé.
-
## 📢 Actualités
-2026-02-16 🎉 PicoClaw a atteint 12K étoiles en une semaine ! Merci à tous pour votre soutien ! PicoClaw grandit plus vite que nous ne l'avions jamais imaginé. Vu le volume élevé de PR, nous avons un besoin urgent de mainteneurs communautaires. Nos rôles de bénévoles et notre feuille de route sont officiellement publiés [ici](docs/ROADMAP.md) — nous avons hâte de vous accueillir !
+2026-03-17 🚀 **v0.2.3 publié !** Interface système tray (Windows & Linux), suivi de statut des sous-agents (`spawn_status`), rechargement à chaud expérimental du gateway, portes de sécurité cron, et 2 correctifs de sécurité. PicoClaw atteint **25K ⭐** !
-2026-02-13 🎉 PicoClaw a atteint 5000 étoiles en 4 jours ! Merci à la communauté ! Nous finalisons la **Feuille de Route du Projet** et mettons en place le **Groupe de Développeurs** pour accélérer le développement de PicoClaw.
-🚀 **Appel à l'action :** Soumettez vos demandes de fonctionnalités dans les GitHub Discussions. Nous les examinerons et les prioriserons lors de notre prochaine réunion hebdomadaire.
+2026-03-09 🎉 **v0.2.1 — Plus grande mise à jour !** Support du protocole MCP, 4 nouveaux canaux (Matrix/IRC/WeCom/Discord Proxy), 3 nouveaux fournisseurs (Kimi/Minimax/Avian), pipeline de vision, stockage mémoire JSONL, et routage de modèles.
-2026-02-09 🎉 PicoClaw est lancé ! Construit en 1 jour pour apporter les Agents IA au matériel à 10$ avec <10 Mo de RAM. 🦐 PicoClaw, c'est parti !
+2026-02-28 📦 **v0.2.0** publié avec support Docker Compose et lanceur Web UI.
+
+2026-02-26 🎉 PicoClaw a atteint **20K étoiles** en seulement 17 jours ! L'orchestration automatique des canaux et les interfaces de capacités sont arrivées.
+
+
+Actualités précédentes...
+
+2026-02-16 🎉 PicoClaw a atteint 12K étoiles en une semaine ! Les rôles de mainteneurs communautaires et la [feuille de route](ROADMAP.md) sont officiellement publiés.
+
+2026-02-13 🎉 PicoClaw a atteint 5000 étoiles en 4 jours ! La Feuille de Route du Projet et le Groupe de Développeurs sont en cours de mise en place.
+
+2026-02-09 🎉 **PicoClaw est lancé !** Construit en 1 jour pour apporter les Agents IA au matériel à $10 avec <10 Mo de RAM. 🦐 PicoClaw, c'est parti !
+
+
## ✨ Fonctionnalités
-🪶 **Ultra-Léger** : Empreinte mémoire <10 Mo — 99% plus petit que Clawdbot pour les fonctionnalités essentielles.
+🪶 **Ultra-Léger** : Empreinte mémoire <10 Mo — 99% plus petit que les fonctionnalités essentielles d'OpenClaw.*
-💰 **Coût Minimal** : Suffisamment efficace pour fonctionner sur du matériel à 10$ — 98% moins cher qu'un Mac mini.
+💰 **Coût Minimal** : Suffisamment efficace pour fonctionner sur du matériel à $10 — 98% moins cher qu'un Mac mini.
-⚡️ **Démarrage Éclair** : Temps de démarrage 400X plus rapide, boot en 1 seconde même sur un cœur unique à 0,6 GHz.
+⚡️ **Démarrage Éclair** : Temps de démarrage 400X plus rapide, boot en <1 seconde même sur un cœur unique à 0,6 GHz.
🌍 **Véritable Portabilité** : Un seul binaire autonome pour RISC-V, ARM, MIPS et x86. Un clic et c'est parti !
🤖 **Auto-Construit par l'IA** : Implémentation native en Go de manière autonome — 95% du cœur généré par l'Agent avec affinement humain dans la boucle.
+🔌 **Support MCP** : Intégration native du [Model Context Protocol](https://modelcontextprotocol.io/) — connectez n'importe quel serveur MCP pour étendre les capacités de l'agent.
+
+👁️ **Pipeline de Vision** : Envoyez des images et fichiers directement à l'agent — encodage base64 automatique pour les LLM multimodaux.
+
+🧠 **Routage Intelligent** : Routage de modèles basé sur des règles — les requêtes simples vont vers des modèles légers, économisant les coûts API.
+
+_*Les versions récentes peuvent utiliser 10–20 Mo en raison des fusions rapides de fonctionnalités. L'optimisation des ressources est prévue. La comparaison de démarrage est basée sur des benchmarks à cœur unique 0,8 GHz (voir tableau ci-dessous)._
+
| | OpenClaw | NanoBot | **PicoClaw** |
| ----------------------------- | ------------- | ------------------------ | ----------------------------------------- |
| **Langage** | TypeScript | Python | **Go** |
-| **RAM** | >1 Go | >100 Mo | **< 10 Mo** |
+| **RAM** | >1 Go | >100 Mo | **< 10 Mo*** |
| **Démarrage**(cœur 0,8 GHz) | >500s | >30s | **<1s** |
-| **Coût** | Mac Mini 599$ | La plupart des SBC Linux ~50$ | **N'importe quelle carte Linux****À partir de 10$** |
+| **Coût** | Mac Mini $599 | La plupart des SBC Linux ~$50 | **N'importe quelle carte Linux****À partir de $10** |
+> 📋 **[Liste de Compatibilité Matérielle](docs/hardware-compatibility.md)** — Voir toutes les cartes testées, du RISC-V à $5 au Raspberry Pi en passant par les téléphones Android. Votre carte n'est pas listée ? Soumettez une PR !
+
## 🦾 Démonstration
### 🛠️ Flux de Travail Standard de l'Assistant
@@ -108,15 +133,15 @@
Donnez une seconde vie à votre téléphone d'il y a dix ans ! Transformez-le en assistant IA intelligent avec PicoClaw. Démarrage rapide :
-1. **Installez Termux** (disponible sur F-Droid ou Google Play).
+1. **Installez [Termux](https://github.com/termux/termux-app)** (Téléchargez depuis [GitHub Releases](https://github.com/termux/termux-app/releases), ou recherchez sur F-Droid / Google Play).
2. **Exécutez les commandes**
```bash
-# Note : Remplacez v0.1.1 par la dernière version depuis la page des Releases
-wget https://github.com/sipeed/picoclaw/releases/download/v0.1.1/picoclaw-linux-arm64
-chmod +x picoclaw-linux-arm64
+# Téléchargez la dernière version depuis https://github.com/sipeed/picoclaw/releases
+wget https://github.com/sipeed/picoclaw/releases/latest/download/picoclaw_Linux_arm64.tar.gz
+tar xzf picoclaw_Linux_arm64.tar.gz
pkg install proot
-termux-chroot ./picoclaw-linux-arm64 onboard
+termux-chroot ./picoclaw onboard # chroot fournit une disposition standard du système de fichiers Linux
```
Puis suivez les instructions de la section « Démarrage Rapide » pour terminer la configuration !
@@ -128,7 +153,7 @@ Puis suivez les instructions de la section « Démarrage Rapide » pour terminer
PicoClaw peut être déployé sur pratiquement n'importe quel appareil Linux !
- 9,9$ [LicheeRV-Nano](https://www.aliexpress.com/item/1005006519668532.html) version E (Ethernet) ou W (WiFi6), pour un Assistant Domotique Minimaliste
-- 30~50$ [NanoKVM](https://www.aliexpress.com/item/1005007369816019.html), ou 100$ [NanoKVM-Pro](https://www.aliexpress.com/item/1005010048471263.html) pour la Maintenance Automatisée de Serveurs
+- 30~$50 [NanoKVM](https://www.aliexpress.com/item/1005007369816019.html), ou 100$ [NanoKVM-Pro](https://www.aliexpress.com/item/1005010048471263.html) pour la Maintenance Automatisée de Serveurs
- 50$ [MaixCAM](https://www.aliexpress.com/item/1005008053333693.html) ou 100$ [MaixCAM2](https://www.kickstarter.com/projects/zepan/maixcam2-build-your-next-gen-4k-ai-camera) pour la Surveillance Intelligente
@@ -137,11 +162,15 @@ PicoClaw peut être déployé sur pratiquement n'importe quel appareil Linux !
## 📦 Installation
-### Installer avec un binaire précompilé
+### Télécharger depuis picoclaw.io (Recommandé)
-Téléchargez le binaire pour votre plateforme depuis la page des [releases](https://github.com/sipeed/picoclaw/releases).
+Visitez **[picoclaw.io](https://picoclaw.io)** — le site officiel détecte automatiquement votre plateforme et propose un téléchargement en un clic. Pas besoin de choisir manuellement une architecture.
-### Installer depuis les sources (dernières fonctionnalités, recommandé pour le développement)
+### Télécharger le binaire précompilé
+
+Vous pouvez aussi télécharger le binaire pour votre plateforme depuis la page [GitHub Releases](https://github.com/sipeed/picoclaw/releases).
+
+### Compiler depuis les sources (pour le développement)
```bash
git clone https://github.com/sipeed/picoclaw.git
@@ -155,457 +184,29 @@ make build
# Compiler pour plusieurs plateformes
make build-all
+# Compiler pour Raspberry Pi Zero 2 W (32-bit : make build-linux-arm ; 64-bit : make build-linux-arm64)
+make build-pi-zero
+
# Compiler et Installer
make install
```
-## 🐳 Docker Compose
-
-Vous pouvez également exécuter PicoClaw avec Docker Compose sans rien installer localement.
-
-```bash
-# 1. Clonez ce dépôt
-git clone https://github.com/sipeed/picoclaw.git
-cd picoclaw
-
-# 2. Premier lancement — génère docker/data/config.json puis s'arrête
-docker compose -f docker/docker-compose.yml --profile gateway up
-# Le conteneur affiche "First-run setup complete." puis s'arrête.
-
-# 3. Configurez vos clés API
-vim docker/data/config.json # Clés API du fournisseur, tokens de bot, etc.
-
-# 4. Démarrer
-docker compose -f docker/docker-compose.yml --profile gateway up -d
-```
-
-> [!TIP]
-> **Utilisateurs Docker** : Par défaut, le Gateway écoute sur `127.0.0.1`, ce qui n'est pas accessible depuis l'hôte. Si vous avez besoin d'accéder aux endpoints de santé ou d'exposer des ports, définissez `PICOCLAW_GATEWAY_HOST=0.0.0.0` dans votre environnement ou mettez à jour `config.json`.
-
-```bash
-# 5. Voir les logs
-docker compose -f docker/docker-compose.yml logs -f picoclaw-gateway
-
-# 6. Arrêter
-docker compose -f docker/docker-compose.yml --profile gateway down
-```
-
-### Mode Agent (exécution unique)
-
-```bash
-# Poser une question
-docker compose -f docker/docker-compose.yml run --rm picoclaw-agent -m "Combien font 2+2 ?"
-
-# Mode interactif
-docker compose -f docker/docker-compose.yml run --rm picoclaw-agent
-```
-
-### Mettre à jour
-
-```bash
-docker compose -f docker/docker-compose.yml pull
-docker compose -f docker/docker-compose.yml --profile gateway up -d
-```
-
-### 🚀 Démarrage Rapide
-
-> [!TIP]
-> Configurez votre clé API dans `~/.picoclaw/config.json`. Obtenez des clés API : [Volcengine (CodingPlan)](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) (LLM) · [OpenRouter](https://openrouter.ai/keys) (LLM) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) (LLM). La recherche web est optionnelle — obtenez gratuitement l'[API Tavily](https://tavily.com) (1000 requêtes gratuites/mois) ou l'[API Brave Search](https://brave.com/search/api) (2000 requêtes gratuites/mois).
-
-**1. Initialiser**
-
-```bash
-picoclaw onboard
-```
-
-**2. Configurer** (`~/.picoclaw/config.json`)
-
-```json
-{
- "model_list": [
- {
- "model_name": "ark-code-latest",
- "model": "volcengine/ark-code-latest",
- "api_key": "sk-your-api-key",
- "api_base":"https://ark.cn-beijing.volces.com/api/coding/v3"
- },
- {
- "model_name": "gpt-5.4",
- "model": "openai/gpt-5.4",
- "api_key": "sk-your-openai-key",
- "request_timeout": 300,
- "api_base": "https://api.openai.com/v1"
- }
- ],
- "agents": {
- "defaults": {
- "model_name": "gpt-5.4"
- }
- },
- "channels": {
- "telegram": {
- "enabled": true,
- "token": "VOTRE_TOKEN_BOT",
- "allow_from": ["VOTRE_USER_ID"]
- }
- },
- "tools": {
- "web": {
- "brave": {
- "enabled": false,
- "api_key": "VOTRE_CLE_API_BRAVE",
- "max_results": 5
- },
- "duckduckgo": {
- "enabled": true,
- "max_results": 5
- }
- }
- }
-}
-```
-
-> **Nouveau** : Le format de configuration `model_list` permet d'ajouter des fournisseurs sans modifier le code. Voir [Configuration de Modèle](#configuration-de-modèle-model_list) pour plus de détails.
-> `request_timeout` est optionnel et s'exprime en secondes. S'il est omis ou défini à `<= 0`, PicoClaw utilise le délai d'expiration par défaut (120s).
-
-**3. Obtenir des Clés API**
-
-* **Fournisseur LLM** : [OpenRouter](https://openrouter.ai/keys) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) · [Anthropic](https://console.anthropic.com) · [OpenAI](https://platform.openai.com) · [Gemini](https://aistudio.google.com/api-keys)
-* **Recherche Web** (optionnel) : [Brave Search](https://brave.com/search/api) - Offre gratuite disponible (2000 requêtes/mois)
-
-> **Note** : Consultez `config.example.json` pour un modèle de configuration complet.
-
-**4. Discuter**
-
-```bash
-picoclaw agent -m "Combien font 2+2 ?"
-```
-
-Et voilà ! Vous avez un assistant IA fonctionnel en 2 minutes.
-
----
-
-## 💬 Applications de Chat
-
-Discutez avec votre PicoClaw via Telegram, Discord, DingTalk, LINE ou WeCom
-
-| Canal | Configuration |
-| ------------ | -------------------------------------- |
-| **Telegram** | Facile (juste un token) |
-| **Discord** | Facile (token bot + intents) |
-| **QQ** | Facile (AppID + AppSecret) |
-| **DingTalk** | Moyen (identifiants de l'application) |
-| **LINE** | Moyen (identifiants + URL de webhook) |
-| **WeCom AI Bot** | Moyen (Token + clé AES) |
-
-
-Telegram (Recommandé)
-
-**1. Créer un bot**
-
-* Ouvrez Telegram, recherchez `@BotFather`
-* Envoyez `/newbot`, suivez les instructions
-* Copiez le token
-
-**2. Configurer**
-
-```json
-{
- "channels": {
- "telegram": {
- "enabled": true,
- "token": "VOTRE_TOKEN_BOT",
- "allow_from": ["VOTRE_USER_ID"]
- }
- }
-}
-```
-
-> Obtenez votre User ID via `@userinfobot` sur Telegram.
-
-**3. Lancer**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-Discord
-
-**1. Créer un bot**
-
-* Rendez-vous sur
-* Créez une application → Bot → Add Bot
-* Copiez le token du bot
-
-**2. Activer les intents**
-
-* Dans les paramètres du Bot, activez **MESSAGE CONTENT INTENT**
-* (Optionnel) Activez **SERVER MEMBERS INTENT** si vous souhaitez utiliser des listes d'autorisation basées sur les données des membres
-
-**3. Obtenir votre User ID**
-
-* Paramètres Discord → Avancé → activez le **Mode Développeur**
-* Clic droit sur votre avatar → **Copier l'identifiant**
-
-**4. Configurer**
-
-```json
-{
- "channels": {
- "discord": {
- "enabled": true,
- "token": "VOTRE_TOKEN_BOT",
- "allow_from": ["VOTRE_USER_ID"]
- }
- }
-}
-```
-
-**5. Inviter le bot**
-
-* OAuth2 → URL Generator
-* Scopes : `bot`
-* Permissions du Bot : `Send Messages`, `Read Message History`
-* Ouvrez l'URL d'invitation générée et ajoutez le bot à votre serveur
-
-**6. Lancer**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-QQ
-
-**1. Créer un bot**
-
-- Rendez-vous sur la [QQ Open Platform](https://q.qq.com/#)
-- Créez une application → Obtenez l'**AppID** et l'**AppSecret**
-
-**2. Configurer**
-
-```json
-{
- "channels": {
- "qq": {
- "enabled": true,
- "app_id": "VOTRE_APP_ID",
- "app_secret": "VOTRE_APP_SECRET",
- "allow_from": []
- }
- }
-}
-```
-
-> Laissez `allow_from` vide pour autoriser tous les utilisateurs, ou spécifiez des numéros QQ pour restreindre l'accès.
-
-**3. Lancer**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-DingTalk
-
-**1. Créer un bot**
-
-* Rendez-vous sur la [Open Platform](https://open.dingtalk.com/)
-* Créez une application interne
-* Copiez le Client ID et le Client Secret
-
-**2. Configurer**
-
-```json
-{
- "channels": {
- "dingtalk": {
- "enabled": true,
- "client_id": "VOTRE_CLIENT_ID",
- "client_secret": "VOTRE_CLIENT_SECRET",
- "allow_from": []
- }
- }
-}
-```
-
-> Laissez `allow_from` vide pour autoriser tous les utilisateurs, ou spécifiez des identifiants pour restreindre l'accès.
-
-**3. Lancer**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-LINE
-
-**1. Créer un Compte Officiel LINE**
-
-- Rendez-vous sur la [LINE Developers Console](https://developers.line.biz/)
-- Créez un provider → Créez un canal Messaging API
-- Copiez le **Channel Secret** et le **Channel Access Token**
-
-**2. Configurer**
-
-```json
-{
- "channels": {
- "line": {
- "enabled": true,
- "channel_secret": "VOTRE_CHANNEL_SECRET",
- "channel_access_token": "VOTRE_CHANNEL_ACCESS_TOKEN",
- "webhook_path": "/webhook/line",
- "allow_from": []
- }
- }
-}
-```
-
-**3. Configurer l'URL du Webhook**
-
-LINE exige HTTPS pour les webhooks. Utilisez un reverse proxy ou un tunnel :
-
-```bash
-# Exemple avec ngrok (tunnel vers le serveur Gateway partagé)
-ngrok http 18790
-```
-
-Puis configurez l'URL du Webhook dans la LINE Developers Console sur `https://votre-domaine/webhook/line` et activez **Use webhook**.
-
-> **Note** : Le webhook LINE est servi par le serveur Gateway partagé (par défaut `127.0.0.1:18790`). Si vous utilisez ngrok ou un proxy inverse, faites pointer le tunnel vers le port `18790`.
-
-**4. Lancer**
-
-```bash
-picoclaw gateway
-```
-
-> Dans les discussions de groupe, le bot répond uniquement lorsqu'il est mentionné avec @. Les réponses citent le message original.
-
-> **Docker Compose** : Si vous avez besoin d'exposer le webhook LINE via Docker, mappez le port du Gateway partagé (par défaut `18790`) vers l'hôte, par exemple `ports: ["18790:18790"]`. Notez que le serveur Gateway sert les webhooks de tous les canaux à partir de ce port.
-
-
-
-
-WeCom (WeChat Work)
-
-PicoClaw prend en charge trois types d'intégration WeCom :
-
-**Option 1 : WeCom Bot (Robot)** - Configuration plus facile, prend en charge les discussions de groupe
-**Option 2 : WeCom App (Application Personnalisée)** - Plus de fonctionnalités, messagerie proactive, chat privé uniquement
-**Option 3 : WeCom AI Bot (Bot Intelligent)** - Bot IA officiel, réponses en streaming, prend en charge groupe et privé
-
-Voir le [Guide de Configuration WeCom AI Bot](docs/channels/wecom/wecom_aibot/README.zh.md) pour des instructions détaillées.
-
-**Configuration Rapide - WeCom Bot :**
-
-**1. Créer un bot**
-
-* Accédez à la Console d'Administration WeCom → Discussion de Groupe → Ajouter un Bot de Groupe
-* Copiez l'URL du webhook (format : `https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=xxx`)
-
-**2. Configurer**
-
-```json
-{
- "channels": {
- "wecom": {
- "enabled": true,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
- "webhook_url": "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=YOUR_KEY",
- "webhook_path": "/webhook/wecom",
- "allow_from": []
- }
- }
-}
-```
-
-**Configuration Rapide - WeCom App :**
-
-**1. Créer une application**
-
-* Accédez à la Console d'Administration WeCom → Gestion des Applications → Créer une Application
-* Copiez l'**AgentId** et le **Secret**
-* Accédez à la page "Mon Entreprise", copiez le **CorpID**
-
-**2. Configurer la réception des messages**
-
-* Dans les détails de l'application, cliquez sur "Recevoir les Messages" → "Configurer l'API"
-* Définissez l'URL sur `http://your-server:18790/webhook/wecom-app`
-* Générez le **Token** et l'**EncodingAESKey**
-
-**3. Configurer**
-
-```json
-{
- "channels": {
- "wecom_app": {
- "enabled": true,
- "corp_id": "wwxxxxxxxxxxxxxxxx",
- "corp_secret": "YOUR_CORP_SECRET",
- "agent_id": 1000002,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
- "webhook_path": "/webhook/wecom-app",
- "allow_from": []
- }
- }
-}
-```
-
-**4. Lancer**
-
-```bash
-picoclaw gateway
-```
-
-> **Note** : Les callbacks webhook WeCom App sont servis par le serveur Gateway partagé (par défaut `127.0.0.1:18790`). Assurez-vous que le port `18790` est accessible ou utilisez un proxy inverse HTTPS en production.
-
-**Configuration Rapide - WeCom AI Bot :**
-
-**1. Créer un AI Bot**
-
-* Accédez à la Console d'Administration WeCom → Gestion des Applications → AI Bot
-* Configurez l'URL de callback : `http://your-server:18791/webhook/wecom-aibot`
-* Copiez le **Token** et générez l'**EncodingAESKey**
-
-**2. Configurer**
-
-```json
-{
- "channels": {
- "wecom_aibot": {
- "enabled": true,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_43_CHAR_ENCODING_AES_KEY",
- "webhook_path": "/webhook/wecom-aibot",
- "allow_from": [],
- "welcome_message": "Bonjour ! Comment puis-je vous aider ?"
- }
- }
-}
-```
-
-**3. Lancer**
-
-```bash
-picoclaw gateway
-```
-
-> **Note** : WeCom AI Bot utilise le protocole pull en streaming — pas de problème de timeout. Les tâches longues (>5,5 min) basculent automatiquement vers la livraison via `response_url`.
-
-
+**Raspberry Pi Zero 2 W :** Utilisez le binaire correspondant à votre OS : Raspberry Pi OS 32-bit → `make build-linux-arm` ; 64-bit → `make build-linux-arm64`. Ou exécutez `make build-pi-zero` pour compiler les deux.
+
+## 📚 Documentation
+
+Pour des guides détaillés, consultez la documentation ci-dessous. Ce README ne couvre que le démarrage rapide.
+
+| Sujet | Description |
+|-------|-------------|
+| 🐳 [Docker & Démarrage Rapide](docs/fr/docker.md) | Configuration Docker Compose, modes Launcher/Agent, configuration rapide |
+| 💬 [Applications de Chat](docs/fr/chat-apps.md) | Telegram, Discord, WhatsApp, Matrix, QQ, Slack, IRC, DingTalk, LINE, Feishu, WeCom, et plus |
+| ⚙️ [Configuration](docs/fr/configuration.md) | Variables d'environnement, structure du workspace, sources de compétences, bac à sable de sécurité, heartbeat |
+| 🔌 [Fournisseurs & Modèles](docs/fr/providers.md) | 20+ fournisseurs LLM, routage de modèles, configuration model_list, architecture des fournisseurs |
+| 🔄 [Spawn & Tâches Asynchrones](docs/fr/spawn-tasks.md) | Tâches rapides, tâches longues avec spawn, orchestration asynchrone de sous-agents |
+| 🐛 [Dépannage](docs/fr/troubleshooting.md) | Problèmes courants et solutions |
+| 🔧 [Configuration des Outils](docs/fr/tools_configuration.md) | Activation/désactivation par outil, politiques exec |
+| 📋 [Compatibilité Matérielle](docs/hardware-compatibility.md) | Cartes testées, exigences minimales, comment ajouter votre carte |
## Rejoignez le Réseau Social d'Agents
@@ -613,6 +214,27 @@ Connectez PicoClaw au Réseau Social d'Agents simplement en envoyant un seul mes
**Lisez `https://clawdchat.ai/skill.md` et suivez les instructions pour rejoindre [ClawdChat.ai](https://clawdchat.ai)**
+<<<<<<< HEAD
+## 🖥️ Référence CLI
+
+| Commande | Description |
+| ------------------------- | ---------------------------------- |
+| `picoclaw onboard` | Initialiser la config & le workspace |
+| `picoclaw agent -m "..."` | Discuter avec l'agent |
+| `picoclaw agent` | Mode chat interactif |
+| `picoclaw gateway` | Démarrer le gateway |
+| `picoclaw status` | Afficher le statut |
+| `picoclaw version` | Afficher les infos de version |
+| `picoclaw cron list` | Lister les tâches planifiées |
+| `picoclaw cron add ...` | Ajouter une tâche planifiée |
+| `picoclaw cron disable` | Désactiver une tâche planifiée |
+| `picoclaw cron remove` | Supprimer une tâche planifiée |
+| `picoclaw skills list` | Lister les compétences installées |
+| `picoclaw skills install` | Installer une compétence |
+| `picoclaw migrate` | Migrer les données des anciennes versions |
+| `picoclaw auth login` | S'authentifier auprès des fournisseurs |
+| `picoclaw model` | Voir ou changer le modèle par défaut |
+=======
## ⚙️ Configuration
Fichier de configuration : `~/.picoclaw/config.json`
@@ -1153,6 +775,7 @@ Pour le guide de migration détaillé, voir [docs/migration/model-list-migration
| `picoclaw status` | Afficher le statut |
| `picoclaw cron list` | Lister toutes les tâches planifiées |
| `picoclaw cron add ...` | Ajouter une tâche planifiée |
+>>>>>>> refactor/agent
### Tâches Planifiées / Rappels
@@ -1160,78 +783,18 @@ PicoClaw prend en charge les rappels planifiés et les tâches récurrentes via
* **Rappels ponctuels** : « Rappelle-moi dans 10 minutes » → se déclenche une fois après 10 min
* **Tâches récurrentes** : « Rappelle-moi toutes les 2 heures » → se déclenche toutes les 2 heures
-* **Expressions Cron** : « Rappelle-moi à 9h tous les jours » → utilise une expression cron
-
-Les tâches sont stockées dans `~/.picoclaw/workspace/cron/` et traitées automatiquement.
+* **Expressions cron** : « Rappelle-moi à 9h chaque jour » → utilise une expression cron
## 🤝 Contribuer & Feuille de Route
-Les PR sont les bienvenues ! Le code source est volontairement petit et lisible. 🤗
+Les PR sont les bienvenues ! Le code est intentionnellement petit et lisible. 🤗
-Feuille de route à venir...
+Consultez notre [Feuille de Route Communautaire](https://github.com/sipeed/picoclaw/blob/main/ROADMAP.md) complète.
-Groupe de développeurs en construction. Condition d'entrée : au moins 1 PR fusionnée.
+Groupe de développeurs en construction, rejoignez-nous après votre première PR fusionnée !
Groupes d'utilisateurs :
-Discord :
+discord :
-
-## 🐛 Dépannage
-
-### La recherche web affiche « API 配置问题 »
-
-C'est normal si vous n'avez pas encore configuré de clé API de recherche. PicoClaw fournira des liens utiles pour la recherche manuelle.
-
-Pour activer la recherche web :
-
-1. **Option 1 (Recommandé)** : Obtenez une clé API gratuite sur [https://brave.com/search/api](https://brave.com/search/api) (2000 requêtes gratuites/mois) pour les meilleurs résultats.
-2. **Option 2 (Sans carte bancaire)** : Si vous n'avez pas de clé, le système bascule automatiquement sur **DuckDuckGo** (aucune clé requise).
-
-Ajoutez la clé dans `~/.picoclaw/config.json` si vous utilisez Brave :
-
-```json
-{
- "tools": {
- "web": {
- "brave": {
- "enabled": false,
- "api_key": "VOTRE_CLE_API_BRAVE",
- "max_results": 5
- },
- "duckduckgo": {
- "enabled": true,
- "max_results": 5
- }
- }
- }
-}
-```
-
-### Erreurs de filtrage de contenu
-
-Certains fournisseurs (comme Zhipu) disposent d'un filtrage de contenu. Essayez de reformuler votre requête ou utilisez un modèle différent.
-
-### Le bot Telegram affiche « Conflict: terminated by other getUpdates »
-
-Cela se produit lorsqu'une autre instance du bot est en cours d'exécution. Assurez-vous qu'un seul `picoclaw gateway` fonctionne à la fois.
-
----
-
-## 📝 Comparaison des Clés API
-
-| Service | Offre Gratuite | Cas d'Utilisation |
-| ---------------- | -------------------- | ------------------------------------- |
-| **OpenRouter** | 200K tokens/mois | Multiples modèles (Claude, GPT-4, etc.) |
-| **Volcengine CodingPlan** | 9,9¥/premier mois | Idéal pour les utilisateurs chinois, multiples modèles SOTA (Doubao, DeepSeek, etc.) |
-| **Zhipu** | 200K tokens/mois | Convient aux utilisateurs chinois |
-| **Brave Search** | 2000 requêtes/mois | Fonctionnalité de recherche web |
-| **Groq** | Offre gratuite dispo | Inférence ultra-rapide (Llama, Mixtral) |
-| **ModelScope** | 2000 requêtes/jour | Inférence gratuite (Qwen, GLM, DeepSeek, etc.) |
-
----
-
-
-
-
diff --git a/README.id.md b/README.id.md
new file mode 100644
index 000000000..3f462981c
--- /dev/null
+++ b/README.id.md
@@ -0,0 +1,249 @@
+
+
+
+
PicoClaw: Asisten AI Super Ringan berbasis Go
+
+
Perangkat Keras $10 · RAM <10MB · Boot <1 Detik · Ayo, Berangkat!
+
+---
+
+> **PicoClaw** adalah proyek open-source independen yang diinisiasi oleh [Sipeed](https://sipeed.com). Ditulis sepenuhnya dalam **Go** — bukan fork dari OpenClaw, NanoBot, atau proyek lainnya.
+
+🦐 PicoClaw adalah asisten AI pribadi yang super ringan, terinspirasi dari [NanoBot](https://github.com/HKUDS/nanobot), ditulis ulang sepenuhnya dalam Go melalui proses "self-bootstrapping" — di mana AI Agent itu sendiri yang memandu seluruh migrasi arsitektur dan optimasi kode.
+
+⚡️ Berjalan di perangkat keras $10 dengan RAM <10MB: Hemat 99% memori dibanding OpenClaw dan 98% lebih murah dibanding Mac mini!
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+> [!CAUTION]
+> **🚨 KEAMANAN & SALURAN RESMI**
+>
+> * **TANPA KRIPTO:** PicoClaw **TIDAK** memiliki token/koin resmi. Semua klaim di `pump.fun` atau platform trading lainnya adalah **PENIPUAN**.
+>
+> * **DOMAIN RESMI:** Satu-satunya website resmi adalah **[picoclaw.io](https://picoclaw.io)**, dan website perusahaan adalah **[sipeed.com](https://sipeed.com)**
+> * **Peringatan:** Banyak domain `.ai/.org/.com/.net/...` yang didaftarkan oleh pihak ketiga.
+> * **Peringatan:** PicoClaw masih dalam tahap pengembangan awal dan mungkin memiliki masalah keamanan jaringan yang belum teratasi. Jangan deploy ke lingkungan produksi sebelum rilis v1.0.
+> * **Catatan:** PicoClaw baru-baru ini menggabungkan banyak PR, yang mungkin mengakibatkan penggunaan memori lebih besar (10–20MB) pada versi terbaru. Kami berencana untuk memprioritaskan optimasi sumber daya segera setelah fitur saat ini mencapai kondisi stabil.
+
+## 📢 Berita
+
+2026-03-17 🚀 **v0.2.3 Dirilis!** UI system tray (Windows & Linux), pelacakan status sub-agent (`spawn_status`), eksperimental gateway hot-reload, gerbang keamanan cron, dan 2 perbaikan keamanan. PicoClaw kini di **25K ⭐**!
+
+2026-03-09 🎉 **v0.2.1 — Update terbesar!** Dukungan protokol MCP, 4 channel baru (Matrix/IRC/WeCom/Discord Proxy), 3 provider baru (Kimi/Minimax/Avian), pipeline vision, penyimpanan memori JSONL, dan routing model.
+
+2026-02-28 📦 **v0.2.0** dirilis dengan dukungan Docker Compose dan launcher Web UI.
+
+2026-02-26 🎉 PicoClaw mencapai **20K bintang** hanya dalam 17 hari! Orkestrasi channel otomatis dan antarmuka kapabilitas diluncurkan.
+
+
+Berita lama...
+
+2026-02-16 🎉 PicoClaw mencapai 12K bintang dalam satu minggu! Peran maintainer komunitas dan [roadmap](ROADMAP.md) resmi diposting.
+
+2026-02-13 🎉 PicoClaw mencapai 5000 bintang dalam 4 hari! Roadmap Proyek dan pengaturan Grup Pengembang sedang berjalan.
+
+2026-02-09 🎉 **PicoClaw Diluncurkan!** Dibangun dalam 1 hari untuk menghadirkan AI Agent ke perangkat keras $10 dengan RAM <10MB. 🦐 PicoClaw, Ayo Berangkat!
+
+
+
+## ✨ Fitur
+
+🪶 **Super Ringan**: Penggunaan memori <10MB — 99% lebih kecil dari fungsionalitas inti OpenClaw.*
+
+💰 **Biaya Minimal**: Cukup efisien untuk berjalan di perangkat keras $10 — 98% lebih murah dari Mac mini.
+
+⚡️ **Secepat Kilat**: Waktu startup 400X lebih cepat, boot dalam <1 detik bahkan di prosesor single core 0,6GHz.
+
+🌍 **Portabilitas Sejati**: Satu binary mandiri untuk RISC-V, ARM, MIPS, dan x86, Satu Klik Langsung Jalan!
+
+🤖 **AI-Bootstrapped**: Implementasi Go-native secara otonom — 95% kode inti dihasilkan oleh Agent dengan penyempurnaan human-in-the-loop.
+
+🔌 **Dukungan MCP**: Integrasi [Model Context Protocol](https://modelcontextprotocol.io/) native — hubungkan server MCP mana pun untuk memperluas kapabilitas agent.
+
+👁️ **Pipeline Vision**: Kirim gambar dan file langsung ke agent — encoding base64 otomatis untuk LLM multimodal.
+
+🧠 **Routing Cerdas**: Routing model berbasis aturan — kueri sederhana diarahkan ke model ringan, menghemat biaya API.
+
+_*Versi terbaru mungkin menggunakan 10–20MB karena penggabungan fitur yang cepat. Optimasi sumber daya direncanakan. Perbandingan startup berdasarkan benchmark prosesor single-core 0,8GHz (lihat tabel di bawah)._
+
+| | OpenClaw | NanoBot | **PicoClaw** |
+| ----------------------------- | ------------- | ------------------------ | ----------------------------------------- |
+| **Bahasa** | TypeScript | Python | **Go** |
+| **RAM** | >1GB | >100MB | **< 10MB*** |
+| **Startup**(0,8GHz core) | >500d | >30d | **<1d** |
+| **Biaya** | Mac Mini $599 | Kebanyakan Linux SBC ~$50 | **Semua Board Linux****Mulai dari $10** |
+
+
+
+## 🦾 Demonstrasi
+
+### 🛠️ Alur Kerja Asisten Standar
+
+
+
+
🧩 Full-Stack Engineer
+
🗂️ Pencatatan & Manajemen Perencanaan
+
🔎 Pencarian Web & Pembelajaran
+
+
+
+
+
+
+
+
Develop • Deploy • Scale
+
Jadwal • Otomasi • Memori
+
Penemuan • Wawasan • Tren
+
+
+
+### 📱 Jalankan di HP Android Lama
+
+Berikan kehidupan kedua untuk HP lama Anda! Ubah menjadi Asisten AI pintar dengan PicoClaw. Panduan Cepat:
+
+1. **Instal [Termux](https://github.com/termux/termux-app)** (Unduh dari [GitHub Releases](https://github.com/termux/termux-app/releases), atau cari di F-Droid / Google Play).
+2. **Jalankan perintah**
+
+```bash
+# Unduh rilis terbaru dari https://github.com/sipeed/picoclaw/releases
+wget https://github.com/sipeed/picoclaw/releases/latest/download/picoclaw_Linux_arm64.tar.gz
+tar xzf picoclaw_Linux_arm64.tar.gz
+pkg install proot
+termux-chroot ./picoclaw onboard
+```
+
+Kemudian ikuti instruksi di bagian "Panduan Cepat" untuk menyelesaikan konfigurasi!
+
+
+
+### 🐜 Deploy Inovatif dengan Footprint Rendah
+
+PicoClaw dapat di-deploy di hampir semua perangkat Linux!
+
+- $9,9 [LicheeRV-Nano](https://www.aliexpress.com/item/1005006519668532.html) versi E(Ethernet) atau W(WiFi6), untuk Home Assistant Minimal
+- $30~50 [NanoKVM](https://www.aliexpress.com/item/1005007369816019.html), atau $100 [NanoKVM-Pro](https://www.aliexpress.com/item/1005010048471263.html) untuk Pemeliharaan Server Otomatis
+- $50 [MaixCAM](https://www.aliexpress.com/item/1005008053333693.html) atau $100 [MaixCAM2](https://www.kickstarter.com/projects/zepan/maixcam2-build-your-next-gen-4k-ai-camera) untuk Pemantauan Cerdas
+
+
+
+🌟 Lebih Banyak Kasus Deploy Menanti!
+
+## 📦 Instalasi
+
+### Instal dengan binary yang sudah dikompilasi
+
+Unduh binary untuk platform Anda dari halaman [Releases](https://github.com/sipeed/picoclaw/releases).
+
+### Instal dari source (fitur terbaru, disarankan untuk pengembangan)
+
+```bash
+git clone https://github.com/sipeed/picoclaw.git
+
+cd picoclaw
+make deps
+
+# Build, tidak perlu instal
+make build
+
+# Build untuk berbagai platform
+make build-all
+
+# Build untuk Raspberry Pi Zero 2 W (32-bit: make build-linux-arm; 64-bit: make build-linux-arm64)
+make build-pi-zero
+
+# Build dan Instal
+make install
+```
+
+**Raspberry Pi Zero 2 W:** Gunakan binary yang sesuai dengan OS Anda: Raspberry Pi OS 32-bit → `make build-linux-arm`; 64-bit → `make build-linux-arm64`. Atau jalankan `make build-pi-zero` untuk build keduanya.
+
+## 📚 Dokumentasi
+
+Untuk panduan lengkap, lihat dokumen di bawah. README ini hanya berisi panduan cepat.
+
+| Topik | Deskripsi |
+|-------|-----------|
+| 🐳 [Docker & Panduan Cepat](docs/docker.md) | Pengaturan Docker Compose, mode Launcher/Agent, konfigurasi Panduan Cepat |
+| 💬 [Aplikasi Chat](docs/chat-apps.md) | Telegram, Discord, WhatsApp, Matrix, QQ, Slack, IRC, DingTalk, LINE, Feishu, WeCom, dan lainnya |
+| ⚙️ [Konfigurasi](docs/configuration.md) | Variabel environment, tata letak workspace, sumber skill, sandbox keamanan, heartbeat |
+| 🔌 [Provider & Model](docs/providers.md) | 20+ provider LLM, routing model, konfigurasi model_list, arsitektur provider |
+| 🔄 [Spawn & Tugas Async](docs/spawn-tasks.md) | Tugas cepat, tugas panjang dengan spawn, orkestrasi sub-agent async |
+| 🐛 [Pemecahan Masalah](docs/troubleshooting.md) | Masalah umum dan solusinya |
+| 🔧 [Konfigurasi Tools](docs/tools_configuration.md) | Aktifkan/nonaktifkan tool, kebijakan exec |
+
+## Bergabung dengan Jaringan Sosial Agent
+
+Hubungkan Picoclaw ke Jaringan Sosial Agent hanya dengan mengirim satu pesan melalui CLI atau Aplikasi Chat terintegrasi.
+
+**Baca `https://clawdchat.ai/skill.md` dan ikuti instruksi untuk bergabung dengan [ClawdChat.ai](https://clawdchat.ai)**
+
+## 🖥️ Referensi CLI
+
+| Perintah | Deskripsi |
+| ------------------------- | -------------------------------- |
+| `picoclaw onboard` | Inisialisasi konfigurasi & workspace |
+| `picoclaw agent -m "..."` | Chat dengan agent |
+| `picoclaw agent` | Mode chat interaktif |
+| `picoclaw gateway` | Mulai gateway |
+| `picoclaw status` | Tampilkan status |
+| `picoclaw version` | Tampilkan info versi |
+| `picoclaw cron list` | Daftar semua tugas terjadwal |
+| `picoclaw cron add ...` | Tambah tugas terjadwal |
+| `picoclaw cron disable` | Nonaktifkan tugas terjadwal |
+| `picoclaw cron remove` | Hapus tugas terjadwal |
+| `picoclaw skills list` | Daftar skill yang terinstal |
+| `picoclaw skills install` | Instal skill |
+| `picoclaw migrate` | Migrasi data dari versi lama |
+| `picoclaw auth login` | Autentikasi dengan provider |
+
+### Tugas Terjadwal / Pengingat
+
+PicoClaw mendukung pengingat terjadwal dan tugas berulang melalui tool `cron`:
+
+* **Pengingat satu kali**: "Ingatkan saya dalam 10 menit" → terpicu sekali setelah 10 menit
+* **Tugas berulang**: "Ingatkan saya setiap 2 jam" → terpicu setiap 2 jam
+* **Ekspresi cron**: "Ingatkan saya jam 9 pagi setiap hari" → menggunakan ekspresi cron
+
+## 🤝 Kontribusi & Roadmap
+
+PR sangat diterima! Codebase sengaja dibuat kecil dan mudah dibaca. 🤗
+
+Lihat [Roadmap Komunitas](https://github.com/sipeed/picoclaw/blob/main/ROADMAP.md) lengkap kami.
+
+Grup pengembang sedang dibangun, bergabunglah setelah PR pertama Anda di-merge!
+
+Grup Pengguna:
+
+discord:
+
+
diff --git a/README.it.md b/README.it.md
new file mode 100644
index 000000000..27027d95f
--- /dev/null
+++ b/README.it.md
@@ -0,0 +1,249 @@
+
+
+
+
PicoClaw: Assistente IA Ultra-Efficiente in Go
+
+
Hardware da $10 · <10MB RAM · Boot in <1s · 皮皮虾,我们走!
+
+---
+
+> **PicoClaw** è un progetto open-source indipendente avviato da [Sipeed](https://sipeed.com). È scritto interamente in **Go** — non è un fork di OpenClaw, NanoBot o di qualsiasi altro progetto.
+
+🦐 PicoClaw è un assistente IA personale ultra-leggero ispirato a [NanoBot](https://github.com/HKUDS/nanobot), riscritto da zero in Go attraverso un processo di auto-bootstrapping, in cui l'agente IA stesso ha guidato l'intera migrazione architetturale e l'ottimizzazione del codice.
+
+⚡️ Funziona su hardware da $10 con meno di 10MB di RAM: il 99% di memoria in meno rispetto a OpenClaw e il 98% più economico di un Mac mini!
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+> [!CAUTION]
+> **🚨 SICUREZZA & CANALI UFFICIALI**
+>
+> * **NESSUNA CRYPTO:** PicoClaw non ha **NESSUN** token/coin ufficiale. Qualsiasi annuncio su `pump.fun` o altre piattaforme di trading è una **TRUFFA**.
+>
+> * **DOMINIO UFFICIALE:** L'**UNICO** sito ufficiale è **[picoclaw.io](https://picoclaw.io)**, e il sito aziendale è **[sipeed.com](https://sipeed.com)**.
+> * **Attenzione:** Molti domini `.ai/.org/.com/.net/...` sono registrati da terze parti.
+> * **Attenzione:** PicoClaw è in fase di sviluppo iniziale e potrebbe avere problemi di sicurezza di rete non risolti. Non distribuire in ambienti di produzione prima della release v1.0.
+> * **Nota:** PicoClaw ha recentemente unito molte PR, il che potrebbe comportare un'impronta di memoria maggiore (10–20MB) nelle ultime versioni. Prevediamo di dare priorità all'ottimizzazione delle risorse non appena il set di funzionalità corrente raggiungerà uno stato stabile.
+
+## 📢 Novità
+
+2026-03-17 🚀 **v0.2.3 rilasciata!** Interfaccia system tray (Windows & Linux), tracciamento dello stato dei sub-agent (`spawn_status`), hot-reload sperimentale del gateway, gate di sicurezza per cron e 2 correzioni di sicurezza. PicoClaw raggiunge **25K ⭐**!
+
+2026-03-09 🎉 **v0.2.1 — Il più grande aggiornamento di sempre!** Supporto al protocollo MCP, 4 nuovi canali (Matrix/IRC/WeCom/Discord Proxy), 3 nuovi provider (Kimi/Minimax/Avian), pipeline di visione, store di memoria JSONL e routing dei modelli.
+
+2026-02-28 📦 **v0.2.0** rilasciata con supporto Docker Compose e launcher Web UI.
+
+2026-02-26 🎉 PicoClaw ha raggiunto **20K stelle** in soli 17 giorni! Arrivate l'orchestrazione automatica dei canali e le interfacce di capacità.
+
+
+Notizie precedenti...
+
+2026-02-16 🎉 PicoClaw ha raggiunto 12K stelle in una settimana! Ruoli di maintainer della community e [roadmap](ROADMAP.md) pubblicati ufficialmente.
+
+2026-02-13 🎉 PicoClaw ha raggiunto 5000 stelle in 4 giorni! Roadmap del progetto e gruppo sviluppatori in fase di avvio.
+
+2026-02-09 🎉 **PicoClaw lanciato!** Costruito in 1 giorno per portare gli agenti IA su hardware da $10 con <10MB di RAM. 🦐 PicoClaw, andiamo!
+
+
+
+## ✨ Caratteristiche
+
+🪶 **Ultra-Leggero**: Impronta di memoria <10MB — il 99% più piccolo delle funzionalità principali di OpenClaw.*
+
+💰 **Costo Minimo**: Abbastanza efficiente da girare su hardware da $10 — il 98% più economico di un Mac mini.
+
+⚡️ **Avvio Fulmineo**: Tempo di avvio 400 volte più veloce, boot in meno di 1 secondo anche su un singolo core a 0,6 GHz.
+
+🌍 **Vera Portabilità**: Singolo binario autonomo per RISC-V, ARM, MIPS e x86. Un click e si parte!
+
+🤖 **Auto-Costruito dall'IA**: Implementazione nativa in Go in modo autonomo — 95% del core generato dall'Agent con perfezionamento umano nel ciclo.
+
+🔌 **Supporto MCP**: Integrazione nativa del [Model Context Protocol](https://modelcontextprotocol.io/) — connetti qualsiasi server MCP per estendere le capacità dell'agent.
+
+👁️ **Pipeline di Visione**: Invia immagini e file direttamente all'agent — codifica base64 automatica per LLM multimodali.
+
+🧠 **Routing Intelligente**: Routing dei modelli basato su regole — le query semplici vanno verso modelli leggeri, risparmiando sui costi API.
+
+_*Le versioni recenti potrebbero usare 10–20MB a causa delle fusioni rapide di funzionalità. L'ottimizzazione delle risorse è pianificata. Il confronto dell'avvio è basato su benchmark con singolo core a 0,8 GHz (vedi tabella sotto)._
+
+| | OpenClaw | NanoBot | **PicoClaw** |
+| ----------------------------- | ------------- | ------------------------ | ----------------------------------------- |
+| **Linguaggio** | TypeScript | Python | **Go** |
+| **RAM** | >1GB | >100MB | **< 10MB*** |
+| **Avvio**(core 0,8 GHz) | >500s | >30s | **<1s** |
+| **Costo** | Mac Mini $599 | La maggior parte degli SBC Linux ~$50 | **Qualsiasi scheda Linux****A partire da $10** |
+
+
+
+## 🦾 Dimostrazione
+
+### 🛠️ Flussi di Lavoro Standard dell'Assistente
+
+
+
+
🧩 Ingegnere Full-Stack
+
🗂️ Gestione Log & Pianificazione
+
🔎 Ricerca Web & Apprendimento
+
+
+
+
+
+
+
+
Sviluppa • Distribuisci • Scala
+
Pianifica • Automatizza • Memorizza
+
Scopri • Analizza • Tendenze
+
+
+
+### 📱 Usa su vecchi telefoni Android
+
+Dai una seconda vita al tuo telefono di dieci anni fa! Trasformalo in un assistente IA intelligente con PicoClaw. Avvio rapido:
+
+1. **Installa [Termux](https://github.com/termux/termux-app)** (Scarica da [GitHub Releases](https://github.com/termux/termux-app/releases), o cerca su F-Droid / Google Play).
+2. **Esegui i comandi**
+
+```bash
+# Scarica l'ultima release da https://github.com/sipeed/picoclaw/releases
+wget https://github.com/sipeed/picoclaw/releases/latest/download/picoclaw_Linux_arm64.tar.gz
+tar xzf picoclaw_Linux_arm64.tar.gz
+pkg install proot
+termux-chroot ./picoclaw onboard
+```
+
+Poi segui le istruzioni nella sezione "Avvio Rapido" per completare la configurazione!
+
+
+
+### 🐜 Deploy Innovativo a Bassa Impronta
+
+PicoClaw può essere distribuito su quasi qualsiasi dispositivo Linux!
+
+- $9,9 [LicheeRV-Nano](https://www.aliexpress.com/item/1005006519668532.html) versione E (Ethernet) o W (WiFi6), per un Assistente Domotico Minimale
+- $30~50 [NanoKVM](https://www.aliexpress.com/item/1005007369816019.html), o $100 [NanoKVM-Pro](https://www.aliexpress.com/item/1005010048471263.html) per la Manutenzione Automatizzata dei Server
+- $50 [MaixCAM](https://www.aliexpress.com/item/1005008053333693.html) o $100 [MaixCAM2](https://www.kickstarter.com/projects/zepan/maixcam2-build-your-next-gen-4k-ai-camera) per il Monitoraggio Intelligente
+
+
+
+🌟 Molti altri scenari di deploy ti aspettano!
+
+## 📦 Installazione
+
+### Installa con binario precompilato
+
+Scarica il binario per la tua piattaforma dalla pagina delle [Releases](https://github.com/sipeed/picoclaw/releases).
+
+### Installa dai sorgenti (ultime funzionalità, consigliato per lo sviluppo)
+
+```bash
+git clone https://github.com/sipeed/picoclaw.git
+
+cd picoclaw
+make deps
+
+# Compila, senza installare
+make build
+
+# Compila per più piattaforme
+make build-all
+
+# Compila per Raspberry Pi Zero 2 W (32-bit: make build-linux-arm; 64-bit: make build-linux-arm64)
+make build-pi-zero
+
+# Compila e Installa
+make install
+```
+
+**Raspberry Pi Zero 2 W:** Usa il binario che corrisponde al tuo OS: Raspberry Pi OS 32-bit → `make build-linux-arm`; 64-bit → `make build-linux-arm64`. Oppure esegui `make build-pi-zero` per compilare entrambi.
+
+## 📚 Documentazione
+
+Per guide dettagliate, consulta la documentazione qui sotto. Il README copre solo l'avvio rapido.
+
+| Argomento | Descrizione |
+|-----------|-------------|
+| 🐳 [Docker & Avvio Rapido](docs/docker.md) | Configurazione Docker Compose, modalità Launcher/Agent, configurazione rapida |
+| 💬 [App di Chat](docs/chat-apps.md) | Telegram, Discord, WhatsApp, Matrix, QQ, Slack, IRC, DingTalk, LINE, Feishu, WeCom e altro |
+| ⚙️ [Configurazione](docs/it/configuration.md) | Variabili d'ambiente, struttura del workspace, sorgenti delle skill, sandbox di sicurezza, heartbeat |
+| 🔌 [Provider & Modelli](docs/providers.md) | 20+ provider LLM, routing dei modelli, configurazione model_list, architettura dei provider |
+| 🔄 [Spawn & Task Asincroni](docs/spawn-tasks.md) | Task veloci, task lunghi con spawn, orchestrazione asincrona di sub-agent |
+| 🐛 [Risoluzione Problemi](docs/troubleshooting.md) | Problemi comuni e soluzioni |
+| 🔧 [Configurazione degli Strumenti](docs/tools_configuration.md) | Abilitazione/disabilitazione per strumento, politiche exec |
+
+## Unisciti al Social Network degli Agent
+
+Connetti PicoClaw al Social Network degli Agent semplicemente inviando un singolo messaggio tramite CLI o qualsiasi app di chat integrata.
+
+**Leggi `https://clawdchat.ai/skill.md` e segui le istruzioni per unirti a [ClawdChat.ai](https://clawdchat.ai)**
+
+## 🖥️ Riferimento CLI
+
+| Comando | Descrizione |
+| ------------------------- | ---------------------------------- |
+| `picoclaw onboard` | Inizializza config & workspace |
+| `picoclaw agent -m "..."` | Chatta con l'agent |
+| `picoclaw agent` | Modalità chat interattiva |
+| `picoclaw gateway` | Avvia il gateway |
+| `picoclaw status` | Mostra lo stato |
+| `picoclaw version` | Mostra le info sulla versione |
+| `picoclaw cron list` | Elenca tutti i job pianificati |
+| `picoclaw cron add ...` | Aggiunge un job pianificato |
+| `picoclaw cron disable` | Disabilita un job pianificato |
+| `picoclaw cron remove` | Rimuove un job pianificato |
+| `picoclaw skills list` | Elenca le skill installate |
+| `picoclaw skills install` | Installa una skill |
+| `picoclaw migrate` | Migra i dati dalle versioni precedenti |
+| `picoclaw auth login` | Autenticazione con i provider |
+
+### Task Pianificati / Promemoria
+
+PicoClaw supporta promemoria pianificati e task ricorrenti tramite lo strumento `cron`:
+
+* **Promemoria una tantum**: "Ricordami tra 10 minuti" → si attiva una volta dopo 10 min
+* **Task ricorrenti**: "Ricordami ogni 2 ore" → si attiva ogni 2 ore
+* **Espressioni cron**: "Ricordami alle 9 ogni giorno" → usa un'espressione cron
+
+## 🤝 Contribuisci & Roadmap
+
+Le PR sono benvenute! Il codice è volutamente piccolo e leggibile. 🤗
+
+Consulta la nostra [Roadmap della Community](https://github.com/sipeed/picoclaw/blob/main/ROADMAP.md) completa.
+
+Gruppo sviluppatori in costruzione, unisciti dopo la tua prima PR accettata!
+
+Gruppi utenti:
+
+discord:
+
+
diff --git a/README.ja.md b/README.ja.md
index 3f43e29ad..a2265d6be 100644
--- a/README.ja.md
+++ b/README.ja.md
@@ -1,13 +1,12 @@
-[中文](README.zh.md) | [日本語](README.ja.md) | [Português](README.pt-br.md) | [Tiếng Việt](README.vi.md) | [Français](README.fr.md) | **English**
+[中文](README.zh.md) | [日本語](README.ja.md) | [Português](README.pt-br.md) | [Tiếng Việt](README.vi.md) | [Français](README.fr.md) | [Italiano](README.it.md) | [Bahasa Indonesia](README.id.md) | **English**
---
-🦐 PicoClaw is an ultra-lightweight personal AI Assistant inspired by [nanobot](https://github.com/HKUDS/nanobot), refactored from the ground up in Go through a self-bootstrapping process, where the AI agent itself drove the entire architectural migration and code optimization.
+> **PicoClaw** is an independent open-source project initiated by [Sipeed](https://sipeed.com). It is written entirely in **Go** — not a fork of OpenClaw, NanoBot, or any other project.
+
+🦐 PicoClaw is an ultra-lightweight personal AI Assistant inspired by [NanoBot](https://github.com/HKUDS/nanobot), refactored from the ground up in Go through a self-bootstrapping process, where the AI agent itself drove the entire architectural migration and code optimization.
⚡️ Runs on $10 hardware with <10MB RAM: That's 99% less memory than OpenClaw and 98% cheaper than a Mac mini!
@@ -55,34 +57,56 @@
## 📢 News
-2026-02-16 🎉 PicoClaw hit 12K stars in one week! Thank you all for your support! PicoClaw is growing faster than we ever imagined. Given the high volume of PRs, we urgently need community maintainers. Our volunteer roles and roadmap are officially posted [here](ROADMAP.md) —we can’t wait to have you on board!
+2026-03-17 🚀 **v0.2.3 Released!** System tray UI (Windows & Linux), sub-agent status tracking (`spawn_status`), experimental gateway hot-reload, cron security gates, and 2 security fixes. PicoClaw now at **25K ⭐**!
-2026-02-13 🎉 PicoClaw hit 5000 stars in 4days! Thank you for the community! There are so many PRs & issues coming in (during Chinese New Year holidays), we are finalizing the Project Roadmap and setting up the Developer Group to accelerate PicoClaw's development.
-🚀 Call to Action: Please submit your feature requests in GitHub Discussions. We will review and prioritize them during our upcoming weekly meeting.
+2026-03-09 🎉 **v0.2.1 — Biggest update yet!** MCP protocol support, 4 new channels (Matrix/IRC/WeCom/Discord Proxy), 3 new providers (Kimi/Minimax/Avian), vision pipeline, JSONL memory store, and model routing.
-2026-02-09 🎉 PicoClaw Launched! Built in 1 day to bring AI Agents to $10 hardware with <10MB RAM. 🦐 PicoClaw,Let's Go!
+2026-02-28 📦 **v0.2.0** released with Docker Compose support and Web UI launcher.
+
+2026-02-26 🎉 PicoClaw hit **20K stars** in just 17 days! Channel auto-orchestration and capability interfaces landed.
+
+
+Older news...
+
+2026-02-16 🎉 PicoClaw hit 12K stars in one week! Community maintainer roles and [roadmap](ROADMAP.md) officially posted.
+
+2026-02-13 🎉 PicoClaw hit 5000 stars in 4 days! Project Roadmap and Developer Group setup underway.
+
+2026-02-09 🎉 **PicoClaw Launched!** Built in 1 day to bring AI Agents to $10 hardware with <10MB RAM. 🦐 PicoClaw,Let's Go!
+
+
## ✨ Features
-🪶 **Ultra-Lightweight**: <10MB Memory footprint — 99% smaller than Clawdbot - core functionality.
+🪶 **Ultra-Lightweight**: <10MB Memory footprint — 99% smaller than OpenClaw core functionality.*
💰 **Minimal Cost**: Efficient enough to run on $10 Hardware — 98% cheaper than a Mac mini.
-⚡️ **Lightning Fast**: 400X Faster startup time, boot in 1 second even in 0.6GHz single core.
+⚡️ **Lightning Fast**: 400X Faster startup time, boot in <1 second even on 0.6GHz single core.
🌍 **True Portability**: Single self-contained binary across RISC-V, ARM, MIPS, and x86, One-click to Go!
🤖 **AI-Bootstrapped**: Autonomous Go-native implementation — 95% Agent-generated core with human-in-the-loop refinement.
+🔌 **MCP Support**: Native [Model Context Protocol](https://modelcontextprotocol.io/) integration — connect any MCP server to extend agent capabilities.
+
+👁️ **Vision Pipeline**: Send images and files directly to the agent — automatic base64 encoding for multimodal LLMs.
+
+🧠 **Smart Routing**: Rule-based model routing — simple queries go to lightweight models, saving API costs.
+
+_*Recent versions may use 10–20MB due to rapid feature merges. Resource optimization is planned. Startup comparison based on 0.8GHz single-core benchmarks (see table below)._
+
| | OpenClaw | NanoBot | **PicoClaw** |
| ----------------------------- | ------------- | ------------------------ | ----------------------------------------- |
| **Language** | TypeScript | Python | **Go** |
-| **RAM** | >1GB | >100MB | **< 10MB** |
+| **RAM** | >1GB | >100MB | **< 10MB*** |
| **Startup**(0.8GHz core) | >500s | >30s | **<1s** |
-| **Cost** | Mac Mini 599$ | Most Linux SBC ~50$ | **Any Linux Board****As low as 10$** |
+| **Cost** | Mac Mini $599 | Most Linux SBC ~$50 | **Any Linux Board****As low as $10** |
+> 📋 **[Hardware Compatibility List](docs/hardware-compatibility.md)** — See all tested boards, from $5 RISC-V to Raspberry Pi to Android phones. Your board not listed? Submit a PR!
+
## 🦾 Demonstration
### 🛠️ Standard Assistant Workflows
@@ -109,18 +133,19 @@
Give your decade-old phone a second life! Turn it into a smart AI Assistant with PicoClaw. Quick Start:
-1. **Install Termux** (Available on F-Droid or Google Play).
+1. **Install [Termux](https://github.com/termux/termux-app)** (Download from [GitHub Releases](https://github.com/termux/termux-app/releases), or search in F-Droid / Google Play).
2. **Execute cmds**
```bash
-# Note: Replace v0.1.1 with the latest version from the Releases page
-wget https://github.com/sipeed/picoclaw/releases/download/v0.1.1/picoclaw-linux-arm64
-chmod +x picoclaw-linux-arm64
+# Download the latest release from https://github.com/sipeed/picoclaw/releases
+wget https://github.com/sipeed/picoclaw/releases/latest/download/picoclaw_Linux_arm64.tar.gz
+tar xzf picoclaw_Linux_arm64.tar.gz
pkg install proot
-termux-chroot ./picoclaw-linux-arm64 onboard
+termux-chroot ./picoclaw onboard # chroot provides a standard Linux filesystem layout
```
And then follow the instructions in the "Quick Start" section to complete the configuration!
+
### 🐜 Innovative Low-Footprint Deploy
@@ -137,11 +162,15 @@ PicoClaw can be deployed on almost any Linux device!
## 📦 Install
-### Install with precompiled binary
+### Download from picoclaw.io (Recommended)
-Download the firmware for your platform from the [release](https://github.com/sipeed/picoclaw/releases) page.
+Visit **[picoclaw.io](https://picoclaw.io)** — the official website auto-detects your platform and provides one-click download. No need to manually pick an architecture.
-### Install from source (latest features, recommended for development)
+### Download precompiled binary
+
+Alternatively, download the binary for your platform from the [GitHub Releases](https://github.com/sipeed/picoclaw/releases) page.
+
+### Build from source (for development)
```bash
git clone https://github.com/sipeed/picoclaw.git
@@ -162,11 +191,11 @@ make build-pi-zero
make install
```
-**Raspberry Pi Zero 2 W:** Use the binary that matches your OS: 32-bit Raspberry Pi OS → `make build-linux-arm` (output: `build/picoclaw-linux-arm`); 64-bit → `make build-linux-arm64` (output: `build/picoclaw-linux-arm64`). Or run `make build-pi-zero` to build both.
+**Raspberry Pi Zero 2 W:** Use the binary that matches your OS: 32-bit Raspberry Pi OS → `make build-linux-arm`; 64-bit → `make build-linux-arm64`. Or run `make build-pi-zero` to build both.
-## 🐳 Docker Compose
+## 📚 Documentation
-You can also run PicoClaw using Docker Compose without installing anything locally.
+For detailed guides, see the docs below. The README covers quick start only.
```bash
# 1. Clone this repo
@@ -228,7 +257,7 @@ docker compose -f docker/docker-compose.yml --profile gateway up -d
### 🚀 Quick Start
> [!TIP]
-> Set your API Key in `~/.picoclaw/config.json`. Get API Keys: [Volcengine (CodingPlan)](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) (LLM) · [OpenRouter](https://openrouter.ai/keys) (LLM) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) (LLM). Web search is optional — get a free [Tavily API](https://tavily.com) (1000 free queries/month) or [Brave Search API](https://brave.com/search/api) (2000 free queries/month).
+> Set your API Key in `~/.picoclaw/config.json`. Get API Keys: [Volcengine (CodingPlan)](https://console.volcengine.com) (LLM) · [OpenRouter](https://openrouter.ai/keys) (LLM) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) (LLM). Web search is optional — get a free [Tavily API](https://tavily.com) (1000 free queries/month) or [Brave Search API](https://brave.com/search/api) (2000 free queries/month).
**1. Initialize**
@@ -253,8 +282,7 @@ picoclaw onboard
{
"model_name": "ark-code-latest",
"model": "volcengine/ark-code-latest",
- "api_key": "sk-your-api-key",
- "api_base":"https://ark.cn-beijing.volces.com/api/coding/v3"
+ "api_key": "sk-your-api-key"
},
{
"model_name": "gpt-5.4",
@@ -640,80 +668,13 @@ PicoClaw supports three types of WeCom integration:
See [WeCom AI Bot Configuration Guide](docs/channels/wecom/wecom_aibot/README.zh.md) for detailed setup instructions.
-**Quick Setup - WeCom Bot:**
-
-**1. Create a bot**
-
-* Go to WeCom Admin Console → Group Chat → Add Group Bot
-* Copy the webhook URL (format: `https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=xxx`)
-
-**2. Configure**
-
-```json
-{
- "channels": {
- "wecom": {
- "enabled": true,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
- "webhook_url": "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=YOUR_KEY",
- "webhook_path": "/webhook/wecom",
- "allow_from": []
- }
- }
-}
-```
-
-> WeCom webhook is served on the shared Gateway server (`gateway.host`:`gateway.port`, default `127.0.0.1:18790`).
-
-**Quick Setup - WeCom App:**
-
-**1. Create an app**
-
-* Go to WeCom Admin Console → App Management → Create App
-* Copy **AgentId** and **Secret**
-* Go to "My Company" page, copy **CorpID**
-
-**2. Configure receive message**
-
-* In App details, click "Receive Message" → "Set API"
-* Set URL to `http://your-server:18790/webhook/wecom-app`
-* Generate **Token** and **EncodingAESKey**
-
-**3. Configure**
-
-```json
-{
- "channels": {
- "wecom_app": {
- "enabled": true,
- "corp_id": "wwxxxxxxxxxxxxxxxx",
- "corp_secret": "YOUR_CORP_SECRET",
- "agent_id": 1000002,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
- "webhook_path": "/webhook/wecom-app",
- "allow_from": []
- }
- }
-}
-```
-
-**4. Run**
-
-```bash
-picoclaw gateway
-```
-
-> **Note**: WeCom webhook callbacks are served on the Gateway port (default 18790). Use a reverse proxy for HTTPS.
-
**Quick Setup - WeCom AI Bot:**
**1. Create an AI Bot**
-* Go to WeCom Admin Console → App Management → AI Bot
-* In the AI Bot settings, configure callback URL: `http://your-server:18791/webhook/wecom-aibot`
-* Copy **Token** and click "Random Generate" for **EncodingAESKey**
+* Go to WeCom Admin Console → AI Bot
+* Create a new AI Bot → Set name, avatar, etc.
+* Copy **Bot ID** and **Secret**
**2. Configure**
@@ -722,9 +683,8 @@ picoclaw gateway
"channels": {
"wecom_aibot": {
"enabled": true,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_43_CHAR_ENCODING_AES_KEY",
- "webhook_path": "/webhook/wecom-aibot",
+ "bot_id": "YOUR_BOT_ID",
+ "secret": "YOUR_SECRET",
"allow_from": [],
"welcome_message": "Hello! How can I help you?"
}
@@ -1391,7 +1351,7 @@ picoclaw agent -m "Hello"
-## CLI Reference
+## 🖥️ CLI Reference
| Command | Description |
| ------------------------- | ----------------------------- |
@@ -1400,8 +1360,15 @@ picoclaw agent -m "Hello"
| `picoclaw agent` | Interactive chat mode |
| `picoclaw gateway` | Start the gateway |
| `picoclaw status` | Show status |
+| `picoclaw version` | Show version info |
| `picoclaw cron list` | List all scheduled jobs |
| `picoclaw cron add ...` | Add a scheduled job |
+| `picoclaw cron disable` | Disable a scheduled job |
+| `picoclaw cron remove` | Remove a scheduled job |
+| `picoclaw skills list` | List installed skills |
+| `picoclaw skills install` | Install a skill |
+| `picoclaw migrate` | Migrate data from older versions |
+| `picoclaw auth login` | Authenticate with providers |
### Scheduled Tasks / Reminders
@@ -1411,8 +1378,6 @@ PicoClaw supports scheduled reminders and recurring tasks through the `cron` too
* **Recurring tasks**: "Remind me every 2 hours" → triggers every 2 hours
* **Cron expressions**: "Remind me at 9am daily" → uses cron expression
-Jobs are stored in `~/.picoclaw/workspace/cron/` and processed automatically.
-
## 🤝 Contribute & Roadmap
PRs welcome! The codebase is intentionally small and readable. 🤗
@@ -1425,134 +1390,5 @@ User Groups:
discord:
-
-
-## 🐛 Troubleshooting
-
-### Web search says "API key configuration issue"
-
-This is normal if you haven't configured a search API key yet. PicoClaw will provide helpful links for manual searching.
-
-#### Search Provider Priority
-
-PicoClaw automatically selects the best available search provider in this order:
-1. **Perplexity** (if enabled and API key configured) - AI-powered search with citations
-2. **Brave Search** (if enabled and API key configured) - Privacy-focused paid API ($5/1000 queries)
-3. **SearXNG** (if enabled and base_url configured) - Self-hosted metasearch aggregating 70+ engines (free)
-4. **DuckDuckGo** (if enabled, default fallback) - No API key required (free)
-
-#### Web Search Configuration Options
-
-**Option 1 (Best Results)**: Perplexity AI Search
-```json
-{
- "tools": {
- "web": {
- "perplexity": {
- "enabled": true,
- "api_key": "YOUR_PERPLEXITY_API_KEY",
- "max_results": 5
- }
- }
- }
-}
-```
-
-**Option 2 (Paid API)**: Get an API key at [https://brave.com/search/api](https://brave.com/search/api) ($5/1000 queries, ~$5-6/month)
-```json
-{
- "tools": {
- "web": {
- "brave": {
- "enabled": true,
- "api_key": "YOUR_BRAVE_API_KEY",
- "max_results": 5
- }
- }
- }
-}
-```
-
-**Option 3 (Self-Hosted)**: Deploy your own [SearXNG](https://github.com/searxng/searxng) instance
-```json
-{
- "tools": {
- "web": {
- "searxng": {
- "enabled": true,
- "base_url": "http://your-server:8888",
- "max_results": 5
- }
- }
- }
-}
-```
-
-Benefits of SearXNG:
-- **Zero cost**: No API fees or rate limits
-- **Privacy-focused**: Self-hosted, no tracking
-- **Aggregate results**: Queries 70+ search engines simultaneously
-- **Perfect for cloud VMs**: Solves datacenter IP blocking issues (Oracle Cloud, GCP, AWS, Azure)
-- **No API key needed**: Just deploy and configure the base URL
-
-**Option 4 (No Setup Required)**: DuckDuckGo is enabled by default as fallback (no API key needed)
-
-Add the key to `~/.picoclaw/config.json` if using Brave:
-
-```json
-{
- "tools": {
- "web": {
- "brave": {
- "enabled": false,
- "api_key": "YOUR_BRAVE_API_KEY",
- "max_results": 5
- },
- "duckduckgo": {
- "enabled": true,
- "max_results": 5
- },
- "perplexity": {
- "enabled": false,
- "api_key": "YOUR_PERPLEXITY_API_KEY",
- "max_results": 5
- },
- "searxng": {
- "enabled": false,
- "base_url": "http://your-searxng-instance:8888",
- "max_results": 5
- }
- }
- }
-}
-```
-
-### Getting content filtering errors
-
-Some providers (like Zhipu) have content filtering. Try rephrasing your query or use a different model.
-
-### Telegram bot says "Conflict: terminated by other getUpdates"
-
-This happens when another instance of the bot is running. Make sure only one `picoclaw gateway` is running at a time.
-
----
-
-## 📝 API Key Comparison
-
-| Service | Free Tier | Use Case |
-| ---------------- | ------------------------ | ------------------------------------- |
-| **OpenRouter** | 200K tokens/month | Multiple models (Claude, GPT-4, etc.) |
-| **Volcengine CodingPlan** | ¥9.9/first month | Best for Chinese users, multiple SOTA models (Doubao, DeepSeek, etc.) |
-| **Zhipu** | 200K tokens/month | Suitable for Chinese users |
-| **Brave Search** | Paid ($5/1000 queries) | Web search functionality |
-| **SearXNG** | Unlimited (self-hosted) | Privacy-focused metasearch (70+ engines) |
-| **Groq** | Free tier available | Fast inference (Llama, Mixtral) |
-| **Cerebras** | Free tier available | Fast inference (Llama, Qwen, etc.) |
-| **LongCat** | Up to 5M tokens/day | Fast inference (free tier) |
-| **ModelScope** | 2000 requests/day | Free inference (Qwen, GLM, DeepSeek, etc.) |
-
----
-
-
---
-🦐 **PicoClaw** é um assistente pessoal de IA ultra-leve inspirado no [nanobot](https://github.com/HKUDS/nanobot), reescrito do zero em **Go** por meio de um processo de "auto-inicialização" (self-bootstrapping) — onde o próprio agente de IA conduziu toda a migração de arquitetura e otimização de código.
+> **PicoClaw** é um projeto open-source independente iniciado pela [Sipeed](https://sipeed.com). É escrito inteiramente em **Go** — não é um fork do OpenClaw, NanoBot ou qualquer outro projeto.
-⚡️ **Extremamente leve:** Roda em hardware de apenas **$10** com **<10MB** de RAM. Isso é 99% menos memória que o OpenClaw e 98% mais barato que um Mac mini!
+🦐 PicoClaw é um assistente pessoal de IA ultra-leve inspirado no [NanoBot](https://github.com/HKUDS/nanobot), reescrito do zero em Go por meio de um processo de auto-inicialização (self-bootstrapping), onde o próprio agente de IA conduziu toda a migração de arquitetura e otimização de código.
+
+⚡️ Roda em hardware de $10 com <10MB de RAM: Isso é 99% menos memória que o OpenClaw e 98% mais barato que um Mac mini!
-
-
-
-
-
-
-
-
-
-
-
-
+
+
+
+
+
+
+
+
+
+
+
+
> [!CAUTION]
> **🚨 DECLARAÇÃO DE SEGURANÇA & CANAIS OFICIAIS**
>
> * **SEM CRIPTOMOEDAS:** O PicoClaw **NÃO** possui nenhum token/moeda oficial. Todas as alegações no `pump.fun` ou outras plataformas de negociação são **GOLPES**.
-> * **DOMÍNIO OFICIAL:** O **ÚNICO** site oficial é o **[picoclaw.io](https://picoclaw.io)**, e o site da empresa é o **[sipeed.com](https://sipeed.com)**.
-> * **Aviso:** Muitos domínios `.ai/.org/.com/.net/...` foram registrados por terceiros, não são nossos.
+>
+> * **DOMÍNIO OFICIAL:** O **ÚNICO** site oficial é o **[picoclaw.io](https://picoclaw.io)**, e o site da empresa é o **[sipeed.com](https://sipeed.com)**
+> * **Aviso:** Muitos domínios `.ai/.org/.com/.net/...` foram registrados por terceiros.
> * **Aviso:** O PicoClaw está em fase inicial de desenvolvimento e pode ter problemas de segurança de rede não resolvidos. Não implante em ambientes de produção antes da versão v1.0.
-> * **Nota:** O PicoClaw recentemente fez merge de muitos PRs, o que pode resultar em maior consumo de memória (10-20MB) nas versões mais recentes. Planejamos priorizar a otimização de recursos assim que o conjunto de funcionalidades estiver estável.
-
+> * **Nota:** O PicoClaw recentemente fez merge de muitos PRs, o que pode resultar em maior consumo de memória (10–20MB) nas versões mais recentes. Planejamos priorizar a otimização de recursos assim que o conjunto de funcionalidades estiver estável.
## 📢 Novidades
-2026-02-16 🎉 PicoClaw atingiu 12K stars em uma semana! Obrigado a todos pelo apoio! O PicoClaw está crescendo mais rápido do que jamais imaginamos. Dado o alto volume de PRs, precisamos urgentemente de maintainers da comunidade. Nossos papéis de voluntários e roadmap foram publicados oficialmente [aqui](docs/ROADMAP.md) — estamos ansiosos para ter você a bordo!
+2026-03-17 🚀 **v0.2.3 Lançado!** Interface de bandeja do sistema (Windows & Linux), rastreamento de status de sub-agentes (`spawn_status`), hot-reload experimental do gateway, portões de segurança para cron e 2 correções de segurança. PicoClaw agora com **25K ⭐**!
-2026-02-13 🎉 PicoClaw atingiu 5000 stars em 4 dias! Obrigado à comunidade! Estamos finalizando o **Roadmap do Projeto** e configurando o **Grupo de Desenvolvedores** para acelerar o desenvolvimento do PicoClaw.
+2026-03-09 🎉 **v0.2.1 — Maior atualização até agora!** Suporte ao protocolo MCP, 4 novos canais (Matrix/IRC/WeCom/Discord Proxy), 3 novos provedores (Kimi/Minimax/Avian), pipeline de visão, armazenamento de memória JSONL e roteamento de modelos.
-🚀 **Chamada para Ação:** Envie suas solicitações de funcionalidades nas GitHub Discussions. Revisaremos e priorizaremos na próxima reunião semanal.
+2026-02-28 📦 **v0.2.0** lançado com suporte a Docker Compose e launcher Web UI.
-2026-02-09 🎉 PicoClaw lançado oficialmente! Construído em 1 dia para trazer Agentes de IA para hardware de $10 com <10MB de RAM. 🦐 PicoClaw, Partiu!
+2026-02-26 🎉 PicoClaw atingiu **20K stars** em apenas 17 dias! Orquestração automática de canais e interfaces de capacidade implementadas.
+
+
+Novidades anteriores...
+
+2026-02-16 🎉 PicoClaw atingiu 12K stars em uma semana! Papéis de maintainers da comunidade e [roadmap](ROADMAP.md) publicados oficialmente.
+
+2026-02-13 🎉 PicoClaw atingiu 5000 stars em 4 dias! Roadmap do Projeto e Grupo de Desenvolvedores em preparação.
+
+2026-02-09 🎉 **PicoClaw Lançado!** Construído em 1 dia para trazer Agentes de IA para hardware de $10 com <10MB de RAM. 🦐 PicoClaw, Partiu!
+
+
## ✨ Funcionalidades
-🪶 **Ultra-Leve**: Consumo de memória <10MB — 99% menor que o Clawdbot para funcionalidades essenciais.
+🪶 **Ultra-Leve**: Consumo de memória <10MB — 99% menor que o OpenClaw para funcionalidades essenciais.*
💰 **Custo Mínimo**: Eficiente o suficiente para rodar em hardware de $10 — 98% mais barato que um Mac mini.
-⚡️ **Inicialização Relámpago**: Tempo de inicialização 400X mais rápido, boot em 1 segundo mesmo em CPU single-core de 0.6GHz.
+⚡️ **Inicialização Relâmpago**: Tempo de inicialização 400X mais rápido, boot em <1 segundo mesmo em CPU single-core de 0.6GHz.
🌍 **Portabilidade Real**: Um único binário auto-contido para RISC-V, ARM, MIPS e x86. Um clique e já era!
🤖 **Auto-Construído por IA**: Implementação nativa em Go de forma autônoma — 95% do núcleo gerado pelo Agente com refinamento humano no loop.
+🔌 **Suporte MCP**: Integração nativa com o [Model Context Protocol](https://modelcontextprotocol.io/) — conecte qualquer servidor MCP para estender as capacidades do agente.
+
+👁️ **Pipeline de Visão**: Envie imagens e arquivos diretamente ao agente — codificação base64 automática para LLMs multimodais.
+
+🧠 **Roteamento Inteligente**: Roteamento de modelos baseado em regras — consultas simples vão para modelos leves, economizando custos de API.
+
+_*Versões recentes podem usar 10–20MB devido a merges rápidos de funcionalidades. Otimização de recursos está planejada. Comparação de inicialização baseada em benchmarks de single-core a 0.8GHz (veja tabela abaixo)._
+
| | OpenClaw | NanoBot | **PicoClaw** |
| ----------------------------- | ------------- | ------------------------ | ----------------------------------------- |
| **Linguagem** | TypeScript | Python | **Go** |
-| **RAM** | >1GB | >100MB | **< 10MB** |
+| **RAM** | >1GB | >100MB | **< 10MB*** |
| **Inicialização**(CPU 0.8GHz) | >500s | >30s | **<1s** |
| **Custo** | Mac Mini $599 | Maioria dos SBC Linux ~$50 | **Qualquer placa Linux****A partir de $10** |
+> 📋 **[Lista de Compatibilidade de Hardware](docs/hardware-compatibility.md)** — Veja todas as placas testadas, de RISC-V de $5 a Raspberry Pi e telefones Android. Sua placa não está listada? Envie um PR!
+
## 🦾 Demonstração
### 🛠️ Fluxos de Trabalho Padrão do Assistente
-
-
🧩 Engenharia Full-Stack
-
🗂️ Gerenciamento de Logs & Planejamento
-
🔎 Busca Web & Aprendizado
-
-
-
-
-
-
-
-
Desenvolver • Implantar • Escalar
-
Agendar • Automatizar • Memorizar
-
Descobrir • Analisar • Tendências
-
+
+
🧩 Engenharia Full-Stack
+
🗂️ Gerenciamento de Logs & Planejamento
+
🔎 Busca Web & Aprendizado
+
+
+
+
+
+
+
+
Desenvolver • Implantar • Escalar
+
Agendar • Automatizar • Memorizar
+
Descobrir • Analisar • Tendências
+
### 📱 Rode em celulares Android antigos
Dê uma segunda vida ao seu celular de dez anos atrás! Transforme-o em um assistente de IA inteligente com o PicoClaw. Início rápido:
-1. **Instale o Termux** (Disponível no F-Droid ou Google Play).
+1. **Instale o [Termux](https://github.com/termux/termux-app)** (Baixe em [GitHub Releases](https://github.com/termux/termux-app/releases), ou busque no F-Droid / Google Play).
2. **Execute os comandos**
```bash
-# Nota: Substitua v0.1.1 pela versao mais recente da pagina de Releases
-wget https://github.com/sipeed/picoclaw/releases/download/v0.1.1/picoclaw-linux-arm64
-chmod +x picoclaw-linux-arm64
+# Baixe a versão mais recente em https://github.com/sipeed/picoclaw/releases
+wget https://github.com/sipeed/picoclaw/releases/latest/download/picoclaw_Linux_arm64.tar.gz
+tar xzf picoclaw_Linux_arm64.tar.gz
pkg install proot
-termux-chroot ./picoclaw-linux-arm64 onboard
+termux-chroot ./picoclaw onboard # chroot fornece um layout padrão do sistema de arquivos Linux
```
Depois siga as instruções na seção "Início Rápido" para completar a configuração!
@@ -128,21 +152,25 @@ Depois siga as instruções na seção "Início Rápido" para completar a config
O PicoClaw pode ser implantado em praticamente qualquer dispositivo Linux!
-- $9.9 [LicheeRV-Nano](https://www.aliexpress.com/item/1005006519668532.html) versão E (Ethernet) ou W (WiFi6), para Assistente Doméstico Minimalista
+- $9.9 [LicheeRV-Nano](https://www.aliexpress.com/item/1005006519668532.html) versão E(Ethernet) ou W(WiFi6), para Assistente Doméstico Minimalista
- $30~50 [NanoKVM](https://www.aliexpress.com/item/1005007369816019.html), ou $100 [NanoKVM-Pro](https://www.aliexpress.com/item/1005010048471263.html) para Manutenção Automatizada de Servidores
- $50 [MaixCAM](https://www.aliexpress.com/item/1005008053333693.html) ou $100 [MaixCAM2](https://www.kickstarter.com/projects/zepan/maixcam2-build-your-next-gen-4k-ai-camera) para Monitoramento Inteligente
-https://private-user-images.githubusercontent.com/83055338/547056448-e7b031ff-d6f5-4468-bcca-5726b6fecb5c.mp4
+
🌟 Mais cenários de implantação aguardam você!
## 📦 Instalação
-### Instalar com binário pré-compilado
+### Baixar de picoclaw.io (Recomendado)
-Baixe o binário para sua plataforma na página de [releases](https://github.com/sipeed/picoclaw/releases).
+Visite **[picoclaw.io](https://picoclaw.io)** — o site oficial detecta automaticamente sua plataforma e oferece download com um clique. Sem necessidade de escolher manualmente a arquitetura.
-### Instalar a partir do código-fonte (funcionalidades mais recentes, recomendado para desenvolvimento)
+### Baixar binário pré-compilado
+
+Alternativamente, baixe o binário para sua plataforma na página de [GitHub Releases](https://github.com/sipeed/picoclaw/releases).
+
+### Compilar a partir do código-fonte (para desenvolvimento)
```bash
git clone https://github.com/sipeed/picoclaw.git
@@ -153,462 +181,60 @@ make deps
# Build, sem necessidade de instalar
make build
-# Build para multiplas plataformas
+# Build para múltiplas plataformas
make build-all
+# Build para Raspberry Pi Zero 2 W (32-bit: make build-linux-arm; 64-bit: make build-linux-arm64)
+make build-pi-zero
+
# Build e Instalar
make install
```
-## 🐳 Docker Compose
+**Raspberry Pi Zero 2 W:** Use o binário correspondente ao seu SO: Raspberry Pi OS 32-bit → `make build-linux-arm`; 64-bit → `make build-linux-arm64`. Ou execute `make build-pi-zero` para compilar ambos.
-Você tambêm pode rodar o PicoClaw usando Docker Compose sem instalar nada localmente.
+## 📚 Documentação
-```bash
-# 1. Clone este repositorio
-git clone https://github.com/sipeed/picoclaw.git
-cd picoclaw
+Para guias detalhados, consulte a documentação abaixo. Este README cobre apenas o início rápido.
-# 2. Primeiro uso — gera docker/data/config.json automaticamente e para
-docker compose -f docker/docker-compose.yml --profile gateway up
-# O contêiner exibe "First-run setup complete." e para.
+| Tópico | Descrição |
+|--------|-----------|
+| 🐳 [Docker & Início Rápido](docs/pt-br/docker.md) | Configuração Docker Compose, modos Launcher/Agent, configuração de Início Rápido |
+| 💬 [Apps de Chat](docs/pt-br/chat-apps.md) | Telegram, Discord, WhatsApp, Matrix, QQ, Slack, IRC, DingTalk, LINE, Feishu, WeCom e mais |
+| ⚙️ [Configuração](docs/pt-br/configuration.md) | Variáveis de ambiente, estrutura do workspace, fontes de skills, sandbox de segurança, heartbeat |
+| 🔌 [Provedores & Modelos](docs/pt-br/providers.md) | 20+ provedores LLM, roteamento de modelos, configuração model_list, arquitetura de provedores |
+| 🔄 [Spawn & Tarefas Assíncronas](docs/pt-br/spawn-tasks.md) | Tarefas rápidas, tarefas longas com spawn, orquestração assíncrona de sub-agentes |
+| 🐛 [Solução de Problemas](docs/pt-br/troubleshooting.md) | Problemas comuns e soluções |
+| 🔧 [Configuração de Ferramentas](docs/pt-br/tools_configuration.md) | Habilitar/desabilitar por ferramenta, políticas de execução |
+| 📋 [Compatibilidade de Hardware](docs/hardware-compatibility.md) | Placas testadas, requisitos mínimos, como adicionar sua placa |
-# 3. Configure suas API keys
-vim docker/data/config.json # Chaves de API do provedor, tokens de bot, etc.
+## Junte-se à Rede Social de Agentes
-# 4. Iniciar
-docker compose -f docker/docker-compose.yml --profile gateway up -d
-```
-
-> [!TIP]
-> **Usuários Docker**: Por padrão, o Gateway ouve em `127.0.0.1`, o que não é acessível a partir do host. Se você precisar acessar os endpoints de integridade ou expor portas, defina `PICOCLAW_GATEWAY_HOST=0.0.0.0` em seu ambiente ou atualize o `config.json`.
-
-```bash
-# 5. Ver logs
-docker compose -f docker/docker-compose.yml logs -f picoclaw-gateway
-
-# 6. Parar
-docker compose -f docker/docker-compose.yml --profile gateway down
-```
-
-### Modo Agente (Execução única)
-
-```bash
-# Fazer uma pergunta
-docker compose -f docker/docker-compose.yml run --rm picoclaw-agent -m "Quanto e 2+2?"
-
-# Modo interativo
-docker compose -f docker/docker-compose.yml run --rm picoclaw-agent
-```
-
-### Atualizar
-
-```bash
-docker compose -f docker/docker-compose.yml pull
-docker compose -f docker/docker-compose.yml --profile gateway up -d
-```
-
-### 🚀 Início Rápido
-
-> [!TIP]
-> Configure sua API key em `~/.picoclaw/config.json`. Obtenha API keys: [Volcengine (CodingPlan)](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) (LLM) · [OpenRouter](https://openrouter.ai/keys) (LLM) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) (LLM). Busca web é **opcional** — obtenha a [API Tavily](https://tavily.com) gratuita (1000 consultas grátis/mês) ou a [Brave Search API](https://brave.com/search/api) (2000 consultas grátis/mês).
-
-**1. Inicializar**
-
-```bash
-picoclaw onboard
-```
-
-**2. Configurar** (`~/.picoclaw/config.json`)
-
-```json
-{
- "model_list": [
- {
- "model_name": "ark-code-latest",
- "model": "volcengine/ark-code-latest",
- "api_key": "sk-your-api-key",
- "api_base":"https://ark.cn-beijing.volces.com/api/coding/v3"
- },
- {
- "model_name": "gpt-5.4",
- "model": "openai/gpt-5.4",
- "api_key": "sk-your-openai-key",
- "request_timeout": 300,
- "api_base": "https://api.openai.com/v1"
- }
- ],
- "agents": {
- "defaults": {
- "model_name": "gpt-5.4"
- }
- },
- "tools": {
- "web": {
- "brave": {
- "enabled": false,
- "api_key": "YOUR_BRAVE_API_KEY",
- "max_results": 5
- },
- "duckduckgo": {
- "enabled": true,
- "max_results": 5
- }
- }
- }
-}
-```
-
-> **Novo**: O formato de configuração `model_list` permite adicionar provedores sem alterar código. Veja [Configuração de Modelo](#configuração-de-modelo-model_list) para detalhes.
-> `request_timeout` é opcional e usa segundos. Se omitido ou definido como `<= 0`, o PicoClaw usa o timeout padrão (120s).
-
-**3. Obter API Keys**
-
-* **Provedor de LLM**: [OpenRouter](https://openrouter.ai/keys) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) · [Anthropic](https://console.anthropic.com) · [OpenAI](https://platform.openai.com) · [Gemini](https://aistudio.google.com/api-keys)
-* **Busca Web** (opcional): [Brave Search](https://brave.com/search/api) - Plano gratuito disponível (2000 consultas/mês)
-
-> **Nota**: Veja `config.example.json` para um modelo de configuração completo.
-
-**4. Conversar**
-
-```bash
-picoclaw agent -m "Quanto e 2+2?"
-```
-
-Pronto! Você tem um assistente de IA funcionando em 2 minutos.
-
----
-
-## 💬 Integração com Apps de Chat
-
-Converse com seu PicoClaw via Telegram, Discord, DingTalk, LINE ou WeCom.
-
-| Canal | Nível de Configuração |
-| --- | --- |
-| **Telegram** | Fácil (apenas um token) |
-| **Discord** | Fácil (bot token + intents) |
-| **QQ** | Fácil (AppID + AppSecret) |
-| **DingTalk** | Médio (credenciais do app) |
-| **LINE** | Médio (credenciais + webhook URL) |
-| **WeCom AI Bot** | Médio (Token + chave AES) |
-
-
-Telegram (Recomendado)
-
-**1. Criar o bot**
-
-* Abra o Telegram, busque `@BotFather`
-* Envie `/newbot`, siga as instruções
-* Copie o token
-
-**2. Configurar**
-
-```json
-{
- "channels": {
- "telegram": {
- "enabled": true,
- "token": "YOUR_BOT_TOKEN",
- "allow_from": ["YOUR_USER_ID"]
- }
- }
-}
-```
-
-> Obtenha seu User ID pelo `@userinfobot` no Telegram.
-
-**3. Executar**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-Discord
-
-**1. Criar o bot**
-
-* Acesse
-* Crie um aplicativo → Bot → Add Bot
-* Copie o token do bot
-
-**2. Habilitar Intents**
-
-* Nas configurações do Bot, habilite **MESSAGE CONTENT INTENT**
-* (Opcional) Habilite **SERVER MEMBERS INTENT** se quiser usar lista de permissões baseada em dados dos membros
-
-**3. Obter seu User ID**
-
-* Configurações do Discord → Avançado → habilite **Modo Desenvolvedor**
-* Clique com botão direito no seu avatar → **Copiar ID do Usuário**
-
-**4. Configurar**
-
-```json
-{
- "channels": {
- "discord": {
- "enabled": true,
- "token": "YOUR_BOT_TOKEN",
- "allow_from": ["YOUR_USER_ID"]
- }
- }
-}
-```
-
-**5. Convidar o bot**
-
-* OAuth2 → URL Generator
-* Scopes: `bot`
-* Bot Permissions: `Send Messages`, `Read Message History`
-* Abra a URL de convite gerada e adicione o bot ao seu servidor
-
-**6. Executar**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-QQ
-
-**1. Criar o bot**
-
-- Acesse a [QQ Open Platform](https://q.qq.com/#)
-- Crie um aplicativo → Obtenha **AppID** e **AppSecret**
-
-**2. Configurar**
-
-```json
-{
- "channels": {
- "qq": {
- "enabled": true,
- "app_id": "YOUR_APP_ID",
- "app_secret": "YOUR_APP_SECRET",
- "allow_from": []
- }
- }
-}
-```
-
-> Deixe `allow_from` vazio para permitir todos os usuários, ou especifique números QQ para restringir o acesso.
-
-**3. Executar**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-DingTalk
-
-**1. Criar o bot**
-
-* Acesse a [Open Platform](https://open.dingtalk.com/)
-* Crie um app interno
-* Copie o Client ID e Client Secret
-
-**2. Configurar**
-
-```json
-{
- "channels": {
- "dingtalk": {
- "enabled": true,
- "client_id": "YOUR_CLIENT_ID",
- "client_secret": "YOUR_CLIENT_SECRET",
- "allow_from": []
- }
- }
-}
-```
-
-> Deixe `allow_from` vazio para permitir todos os usuários, ou especifique IDs para restringir o acesso.
-
-**3. Executar**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-LINE
-
-**1. Criar uma Conta Oficial LINE**
-
-- Acesse o [LINE Developers Console](https://developers.line.biz/)
-- Crie um provider → Crie um canal Messaging API
-- Copie o **Channel Secret** e o **Channel Access Token**
-
-**2. Configurar**
-
-```json
-{
- "channels": {
- "line": {
- "enabled": true,
- "channel_secret": "YOUR_CHANNEL_SECRET",
- "channel_access_token": "YOUR_CHANNEL_ACCESS_TOKEN",
- "webhook_path": "/webhook/line",
- "allow_from": []
- }
- }
-}
-```
-
-**3. Configurar URL do Webhook**
-
-O LINE requer HTTPS para webhooks. Use um reverse proxy ou tunnel:
-
-```bash
-# Exemplo com ngrok
-ngrok http 18790
-```
-
-Em seguida, configure a Webhook URL no LINE Developers Console para `https://seu-dominio/webhook/line` e habilite **Use webhook**.
-
-> **Nota**: O webhook do LINE é servido pelo Gateway compartilhado (padrão 127.0.0.1:18790). Use um proxy reverso/HTTPS ou túnel (como ngrok) para expor o Gateway de forma segura quando necessário.
-
-**4. Executar**
-
-```bash
-picoclaw gateway
-```
-
-> Em chats de grupo, o bot responde apenas quando mencionado com @. As respostas citam a mensagem original.
-
-> **Docker Compose**: Se você usa Docker Compose, exponha o Gateway (padrão 127.0.0.1:18790) se precisar acessar o webhook LINE externamente, por exemplo `ports: ["18790:18790"]`.
-
-
-
-
-WeCom (WeChat Work)
-
-O PicoClaw suporta três tipos de integração WeCom:
-
-**Opção 1: WeCom Bot (Robô)** - Configuração mais fácil, suporta chats em grupo
-**Opção 2: WeCom App (Aplicativo Personalizado)** - Mais recursos, mensagens proativas, somente chat privado
-**Opção 3: WeCom AI Bot (Robô Inteligente)** - Bot IA oficial, respostas em streaming, suporta grupo e privado
-
-Veja o [Guia de Configuração WeCom AI Bot](docs/channels/wecom/wecom_aibot/README.zh.md) para instruções detalhadas.
-
-**Configuração Rápida - WeCom Bot:**
-
-**1. Criar um bot**
-
-* Acesse o Console de Administração WeCom → Chat em Grupo → Adicionar Bot de Grupo
-* Copie a URL do webhook (formato: `https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=xxx`)
-
-**2. Configurar**
-
-```json
-{
- "channels": {
- "wecom": {
- "enabled": true,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
- "webhook_url": "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=YOUR_KEY",
- "webhook_path": "/webhook/wecom",
- "allow_from": []
- }
- }
-}
-```
-
-> **Nota**: O webhook do WeCom Bot é atendido pelo Gateway compartilhado (padrão 127.0.0.1:18790). Use um proxy reverso/HTTPS ou túnel para expor o Gateway em produção.
-
-**Configuração Rápida - WeCom App:**
-
-**1. Criar um aplicativo**
-
-* Acesse o Console de Administração WeCom → Gerenciamento de Aplicativos → Criar Aplicativo
-* Copie o **AgentId** e o **Secret**
-* Acesse a página "Minha Empresa", copie o **CorpID**
-
-**2. Configurar recebimento de mensagens**
-
-* Nos detalhes do aplicativo, clique em "Receber Mensagens" → "Configurar API"
-* Defina a URL como `http://your-server:18790/webhook/wecom-app`
-* Gere o **Token** e o **EncodingAESKey**
-
-**3. Configurar**
-
-```json
-{
- "channels": {
- "wecom_app": {
- "enabled": true,
- "corp_id": "wwxxxxxxxxxxxxxxxx",
- "corp_secret": "YOUR_CORP_SECRET",
- "agent_id": 1000002,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
- "webhook_path": "/webhook/wecom-app",
- "allow_from": []
- }
- }
-}
-```
-
-**4. Executar**
-
-```bash
-picoclaw gateway
-```
-
-> **Nota**: O WeCom App (callbacks de webhook) é servido pelo Gateway compartilhado (padrão 127.0.0.1:18790). Em produção use um proxy reverso HTTPS para expor a porta do Gateway, ou atualize `PICOCLAW_GATEWAY_HOST` para `0.0.0.0` se necessário.
-
-**Configuração Rápida - WeCom AI Bot:**
-
-**1. Criar um AI Bot**
-
-* Acesse o Console de Administração WeCom → Gerenciamento de Aplicativos → AI Bot
-* Configure a URL de callback: `http://your-server:18791/webhook/wecom-aibot`
-* Copie o **Token** e gere o **EncodingAESKey**
-
-**2. Configurar**
-
-```json
-{
- "channels": {
- "wecom_aibot": {
- "enabled": true,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_43_CHAR_ENCODING_AES_KEY",
- "webhook_path": "/webhook/wecom-aibot",
- "allow_from": [],
- "welcome_message": "Olá! Como posso ajudá-lo?"
- }
- }
-}
-```
-
-**3. Executar**
-
-```bash
-picoclaw gateway
-```
-
-> **Nota**: O WeCom AI Bot usa protocolo de pull em streaming — sem preocupações com timeout de resposta. Tarefas longas (>5,5 min) alternam automaticamente para entrega via `response_url`.
-
-
-
-## Junte-se a Rede Social de Agentes
-
-Conecte o PicoClaw a Rede Social de Agentes simplesmente enviando uma única mensagem via CLI ou qualquer App de Chat integrado.
+Conecte o PicoClaw à Rede Social de Agentes simplesmente enviando uma única mensagem via CLI ou qualquer App de Chat integrado.
**Leia `https://clawdchat.ai/skill.md` e siga as instruções para entrar no [ClawdChat.ai](https://clawdchat.ai)**
+<<<<<<< HEAD
+## 🖥️ Referência CLI
+
+| Comando | Descrição |
+| ------------------------- | ----------------------------- |
+| `picoclaw onboard` | Inicializar configuração & workspace |
+| `picoclaw agent -m "..."` | Conversar com o agente |
+| `picoclaw agent` | Modo de chat interativo |
+| `picoclaw gateway` | Iniciar o gateway |
+| `picoclaw status` | Mostrar status |
+| `picoclaw version` | Mostrar informações de versão |
+| `picoclaw cron list` | Listar todas as tarefas agendadas |
+| `picoclaw cron add ...` | Adicionar uma tarefa agendada |
+| `picoclaw cron disable` | Desabilitar uma tarefa agendada |
+| `picoclaw cron remove` | Remover uma tarefa agendada |
+| `picoclaw skills list` | Listar skills instaladas |
+| `picoclaw skills install` | Instalar uma skill |
+| `picoclaw migrate` | Migrar dados de versões anteriores |
+| `picoclaw auth login` | Autenticar com provedores |
+| `picoclaw model` | Ver ou trocar o modelo padrão |
+=======
## ⚙️ Configuração Detalhada
Arquivo de configuração: `~/.picoclaw/config.json`
@@ -1149,86 +775,26 @@ Para o guia de migração detalhado, consulte [docs/migration/model-list-migrati
| `picoclaw status` | Mostrar status |
| `picoclaw cron list` | Listar todas as tarefas agendadas |
| `picoclaw cron add ...` | Adicionar uma tarefa agendada |
+>>>>>>> refactor/agent
### Tarefas Agendadas / Lembretes
O PicoClaw suporta lembretes agendados e tarefas recorrentes por meio da ferramenta `cron`:
-* **Lembretes únicos**: "Remind me in 10 minutes" (Me lembre em 10 minutos) → dispara uma vez após 10min
-* **Tarefas recorrentes**: "Remind me every 2 hours" (Me lembre a cada 2 horas) → dispara a cada 2 horas
-* **Expressões Cron**: "Remind me at 9am daily" (Me lembre às 9h todos os dias) → usa expressão cron
-
-As tarefas são armazenadas em `~/.picoclaw/workspace/cron/` e processadas automaticamente.
+* **Lembretes únicos**: "Me lembre em 10 minutos" → dispara uma vez após 10min
+* **Tarefas recorrentes**: "Me lembre a cada 2 horas" → dispara a cada 2 horas
+* **Expressões Cron**: "Me lembre às 9h todos os dias" → usa expressão cron
## 🤝 Contribuir & Roadmap
PRs são bem-vindos! O código-fonte é intencionalmente pequeno e legível. 🤗
-Roadmap em breve...
+Veja nosso [Roadmap da Comunidade](https://github.com/sipeed/picoclaw/blob/main/ROADMAP.md) completo.
-Grupo de desenvolvedores em formação. Requisito de entrada: Pelo menos 1 PR com merge.
+Grupo de desenvolvedores em formação. Junte-se após seu primeiro PR com merge!
Grupos de usuários:
-Discord:
+discord:
-
-## 🐛 Solução de Problemas
-
-### Busca web mostra "API 配置问题"
-
-Isso é normal se você ainda não configurou uma API key de busca. O PicoClaw fornecerá links úteis para busca manual.
-
-Para habilitar a busca web:
-
-1. **Opção 1 (Recomendado)**: Obtenha uma API key gratuita em [https://brave.com/search/api](https://brave.com/search/api) (2000 consultas grátis/mês) para os melhores resultados.
-2. **Opção 2 (Sem Cartão de Crédito)**: Se você não tem uma key, o sistema automaticamente usa o **DuckDuckGo** como fallback (sem necessidade de key).
-
-Adicione a key em `~/.picoclaw/config.json` se usar o Brave:
-
-```json
-{
- "tools": {
- "web": {
- "brave": {
- "enabled": false,
- "api_key": "YOUR_BRAVE_API_KEY",
- "max_results": 5
- },
- "duckduckgo": {
- "enabled": true,
- "max_results": 5
- }
- }
- }
-}
-```
-
-### Erros de filtragem de conteúdo
-
-Alguns provedores (como Zhipu) possuem filtragem de conteúdo. Tente reformular sua pergunta ou use um modelo diferente.
-
-### Bot do Telegram diz "Conflict: terminated by other getUpdates"
-
-Isso acontece quando outra instância do bot está em execução. Certifique-se de que apenas um `picoclaw gateway` esteja rodando por vez.
-
----
-
-## 📝 Comparação de API Keys
-
-| Serviço | Plano Gratuito | Caso de Uso |
-| --- | --- | --- |
-| **OpenRouter** | 200K tokens/mês | Múltiplos modelos (Claude, GPT-4, etc.) |
-| **Volcengine CodingPlan** | ¥9,9/primeiro mês | Ideal para usuários chineses, múltiplos modelos SOTA (Doubao, DeepSeek, etc.) |
-| **Zhipu** | 200K tokens/mês | Adequado para usuários chineses |
-| **Brave Search** | 2000 consultas/mês | Funcionalidade de busca web |
-| **Groq** | Plano gratuito disponível | Inferência ultra-rápida (Llama, Mixtral) |
-| **Cerebras** | Plano gratuito disponível | Inferência ultra-rápida (Llama 3.3 70B) |
-| **ModelScope** | 2000 requisições/dia | Inferência gratuita (Qwen, GLM, DeepSeek, etc.) |
-
----
-
-
---
-🦐 **PicoClaw** là trợ lý AI cá nhân siêu nhẹ, lấy cảm hứng từ [nanobot](https://github.com/HKUDS/nanobot), được viết lại hoàn toàn bằng **Go** thông qua quá trình "tự khởi tạo" (self-bootstrapping) — nơi chính AI Agent đã tự dẫn dắt toàn bộ quá trình chuyển đổi kiến trúc và tối ưu hóa mã nguồn.
+> **PicoClaw** là dự án mã nguồn mở độc lập được khởi xướng bởi [Sipeed](https://sipeed.com). Được viết hoàn toàn bằng **Go** — không phải là bản fork của OpenClaw, NanoBot hay bất kỳ dự án nào khác.
-⚡️ **Cực kỳ nhẹ:** Chạy trên phần cứng chỉ **$10** với RAM **<10MB**. Tiết kiệm 99% bộ nhớ so với OpenClaw và rẻ hơn 98% so với Mac mini!
+🦐 PicoClaw là trợ lý AI cá nhân siêu nhẹ, lấy cảm hứng từ [NanoBot](https://github.com/HKUDS/nanobot), được viết lại hoàn toàn bằng Go thông qua quá trình "tự khởi tạo" (self-bootstrapping) — nơi chính AI Agent đã tự dẫn dắt toàn bộ quá trình chuyển đổi kiến trúc và tối ưu hóa mã nguồn.
+
+⚡️ Chạy trên phần cứng chỉ $10 với RAM <10MB: Tiết kiệm 99% bộ nhớ so với OpenClaw và rẻ hơn 98% so với Mac mini!
-
-
-
-
-
-
-
-
-
-
-
-
+
+
+
+
+
+
+
+
+
+
+
+
> [!CAUTION]
> **🚨 TUYÊN BỐ BẢO MẬT & KÊNH CHÍNH THỨC**
>
> * **KHÔNG CÓ CRYPTO:** PicoClaw **KHÔNG** có bất kỳ token/coin chính thức nào. Mọi thông tin trên `pump.fun` hoặc các sàn giao dịch khác đều là **LỪA ĐẢO**.
-> * **DOMAIN CHÍNH THỨC:** Website chính thức **DUY NHẤT** là **[picoclaw.io](https://picoclaw.io)**, website công ty là **[sipeed.com](https://sipeed.com)**.
-> * **Cảnh báo:** Nhiều tên miền `.ai/.org/.com/.net/...` đã bị bên thứ ba đăng ký, không phải của chúng tôi.
+>
+> * **DOMAIN CHÍNH THỨC:** Website chính thức **DUY NHẤT** là **[picoclaw.io](https://picoclaw.io)**, website công ty là **[sipeed.com](https://sipeed.com)**
+> * **Cảnh báo:** Nhiều tên miền `.ai/.org/.com/.net/...` đã bị bên thứ ba đăng ký.
> * **Cảnh báo:** PicoClaw đang trong giai đoạn phát triển sớm và có thể còn các vấn đề bảo mật mạng chưa được giải quyết. Không nên triển khai lên môi trường production trước phiên bản v1.0.
> * **Lưu ý:** PicoClaw gần đây đã merge nhiều PR, dẫn đến bộ nhớ sử dụng có thể lớn hơn (10–20MB) ở các phiên bản mới nhất. Chúng tôi sẽ ưu tiên tối ưu tài nguyên khi bộ tính năng đã ổn định.
-
## 📢 Tin tức
-2026-02-16 🎉 PicoClaw đạt 12K stars chỉ trong một tuần! Cảm ơn tất cả mọi người! PicoClaw đang phát triển nhanh hơn chúng tôi tưởng tượng. Do số lượng PR tăng cao, chúng tôi cấp thiết cần maintainer từ cộng đồng. Các vai trò tình nguyện viên và roadmap đã được công bố [tại đây](docs/ROADMAP.md) — rất mong đón nhận sự tham gia của bạn!
+2026-03-17 🚀 **v0.2.3 Phát hành!** Giao diện khay hệ thống (Windows & Linux), theo dõi trạng thái sub-agent (`spawn_status`), hot-reload gateway thử nghiệm, cổng bảo mật cron và 2 bản vá bảo mật. PicoClaw đạt **25K ⭐**!
-2026-02-13 🎉 PicoClaw đạt 5000 stars trong 4 ngày! Cảm ơn cộng đồng! Chúng tôi đang hoàn thiện **Lộ trình dự án (Roadmap)** và thiết lập **Nhóm phát triển** để đẩy nhanh tốc độ phát triển PicoClaw.
-🚀 **Kêu gọi hành động:** Vui lòng gửi yêu cầu tính năng tại GitHub Discussions. Chúng tôi sẽ xem xét và ưu tiên trong cuộc họp hàng tuần.
+2026-03-09 🎉 **v0.2.1 — Bản cập nhật lớn nhất!** Hỗ trợ giao thức MCP, 4 kênh mới (Matrix/IRC/WeCom/Discord Proxy), 3 nhà cung cấp mới (Kimi/Minimax/Avian), pipeline xử lý hình ảnh, bộ nhớ JSONL và định tuyến mô hình.
-2026-02-09 🎉 PicoClaw chính thức ra mắt! Được xây dựng trong 1 ngày để mang AI Agent đến phần cứng $10 với RAM <10MB. 🦐 PicoClaw, Lên Đường!
+2026-02-28 📦 **v0.2.0** phát hành với hỗ trợ Docker Compose và launcher Web UI.
+
+2026-02-26 🎉 PicoClaw đạt **20K stars** chỉ trong 17 ngày! Tự động điều phối kênh và giao diện năng lực đã được triển khai.
+
+
+Tin tức cũ hơn...
+
+2026-02-16 🎉 PicoClaw đạt 12K stars chỉ trong một tuần! Vai trò maintainer cộng đồng và [roadmap](ROADMAP.md) đã được công bố chính thức.
+
+2026-02-13 🎉 PicoClaw đạt 5000 stars trong 4 ngày! Lộ trình dự án và Nhóm phát triển đang được thiết lập.
+
+2026-02-09 🎉 **PicoClaw chính thức ra mắt!** Được xây dựng trong 1 ngày để mang AI Agent đến phần cứng $10 với RAM <10MB. 🦐 PicoClaw, Lên Đường!
+
+
## ✨ Tính năng nổi bật
-🪶 **Siêu nhẹ**: Bộ nhớ sử dụng <10MB — nhỏ hơn 99% so với Clawdbot (chức năng cốt lõi).
+🪶 **Siêu nhẹ**: Bộ nhớ sử dụng <10MB — nhỏ hơn 99% so với OpenClaw (chức năng cốt lõi).*
💰 **Chi phí tối thiểu**: Đủ hiệu quả để chạy trên phần cứng $10 — rẻ hơn 98% so với Mac mini.
-⚡️ **Khởi động siêu nhanh**: Nhanh gấp 400 lần, khởi động trong 1 giây ngay cả trên CPU đơn nhân 0.6GHz.
+⚡️ **Khởi động siêu nhanh**: Nhanh gấp 400 lần, khởi động trong <1 giây ngay cả trên CPU đơn nhân 0.6GHz.
🌍 **Di động thực sự**: Một file binary duy nhất chạy trên RISC-V, ARM, MIPS và x86. Một click là chạy!
🤖 **AI tự xây dựng**: Triển khai Go-native tự động — 95% mã nguồn cốt lõi được Agent tạo ra, với sự tinh chỉnh của con người.
+🔌 **Hỗ trợ MCP**: Tích hợp [Model Context Protocol](https://modelcontextprotocol.io/) gốc — kết nối bất kỳ máy chủ MCP nào để mở rộng khả năng của agent.
+
+👁️ **Pipeline Xử lý Hình ảnh**: Gửi hình ảnh và tệp trực tiếp cho agent — tự động mã hóa base64 cho các LLM đa phương thức.
+
+🧠 **Định tuyến Thông minh**: Định tuyến mô hình dựa trên quy tắc — truy vấn đơn giản chuyển đến mô hình nhẹ, tiết kiệm chi phí API.
+
+_*Các phiên bản gần đây có thể sử dụng 10–20MB do merge tính năng nhanh chóng. Tối ưu tài nguyên đang được lên kế hoạch. So sánh thời gian khởi động dựa trên benchmark đơn nhân 0.8GHz (xem bảng bên dưới)._
+
| | OpenClaw | NanoBot | **PicoClaw** |
| ----------------------------- | ------------- | ------------------------ | ----------------------------------------- |
| **Ngôn ngữ** | TypeScript | Python | **Go** |
-| **RAM** | >1GB | >100MB | **< 10MB** |
+| **RAM** | >1GB | >100MB | **< 10MB*** |
| **Thời gian khởi động**(CPU 0.8GHz) | >500s | >30s | **<1s** |
| **Chi phí** | Mac Mini $599 | Hầu hết SBC Linux ~$50 | **Mọi bo mạch Linux****Chỉ từ $10** |
+> 📋 **[Danh Sách Tương Thích Phần Cứng](docs/hardware-compatibility.md)** — Xem tất cả các board đã được kiểm tra, từ RISC-V $5 đến Raspberry Pi và điện thoại Android. Board của bạn chưa có? Gửi PR!
+
## 🦾 Demo
### 🛠️ Quy trình trợ lý tiêu chuẩn
-
-
🧩 Lập trình Full-Stack
-
🗂️ Quản lý Nhật ký & Kế hoạch
-
🔎 Tìm kiếm Web & Học hỏi
-
-
-
-
-
-
-
-
Phát triển • Triển khai • Mở rộng
-
Lên lịch • Tự động hóa • Ghi nhớ
-
Khám phá • Phân tích • Xu hướng
-
+
+
🧩 Lập trình Full-Stack
+
🗂️ Quản lý Nhật ký & Kế hoạch
+
🔎 Tìm kiếm Web & Học hỏi
+
+
+
+
+
+
+
+
Phát triển • Triển khai • Mở rộng
+
Lên lịch • Tự động hóa • Ghi nhớ
+
Khám phá • Phân tích • Xu hướng
+
+### 📱 Chạy trên điện thoại Android cũ
+
+Hãy cho chiếc điện thoại cũ một cuộc sống mới! Biến nó thành trợ lý AI thông minh với PicoClaw. Bắt đầu nhanh:
+
+1. **Cài đặt [Termux](https://github.com/termux/termux-app)** (Tải từ [GitHub Releases](https://github.com/termux/termux-app/releases), hoặc tìm trên F-Droid / Google Play).
+2. **Chạy các lệnh**
+
+```bash
+# Tải phiên bản mới nhất từ https://github.com/sipeed/picoclaw/releases
+wget https://github.com/sipeed/picoclaw/releases/latest/download/picoclaw_Linux_arm64.tar.gz
+tar xzf picoclaw_Linux_arm64.tar.gz
+pkg install proot
+termux-chroot ./picoclaw onboard # chroot cung cấp bố cục hệ thống tệp Linux tiêu chuẩn
+```
+
+Sau đó làm theo hướng dẫn trong phần "Bắt đầu nhanh" để hoàn tất cấu hình!
+
+
+
### 🐜 Triển khai sáng tạo trên phần cứng tối thiểu
PicoClaw có thể triển khai trên hầu hết mọi thiết bị Linux!
-* $9.9 [LicheeRV-Nano](https://www.aliexpress.com/item/1005006519668532.html) phiên bản E (Ethernet) hoặc W (WiFi6), dùng làm Trợ lý Gia đình tối giản.
-* $30~50 [NanoKVM](https://www.aliexpress.com/item/1005007369816019.html), hoặc $100 [NanoKVM-Pro](https://www.aliexpress.com/item/1005010048471263.html), dùng cho quản trị Server tự động.
-* $50 [MaixCAM](https://www.aliexpress.com/item/1005008053333693.html) hoặc $100 [MaixCAM2](https://www.kickstarter.com/projects/zepan/maixcam2-build-your-next-gen-4k-ai-camera), dùng cho Giám sát thông minh.
+- $9.9 [LicheeRV-Nano](https://www.aliexpress.com/item/1005006519668532.html) phiên bản E(Ethernet) hoặc W(WiFi6), dùng làm Trợ lý Gia đình tối giản
+- $30~50 [NanoKVM](https://www.aliexpress.com/item/1005007369816019.html), hoặc $100 [NanoKVM-Pro](https://www.aliexpress.com/item/1005010048471263.html) dùng cho quản trị Server tự động
+- $50 [MaixCAM](https://www.aliexpress.com/item/1005008053333693.html) hoặc $100 [MaixCAM2](https://www.kickstarter.com/projects/zepan/maixcam2-build-your-next-gen-4k-ai-camera) dùng cho Giám sát thông minh
-https://private-user-images.githubusercontent.com/83055338/547056448-e7b031ff-d6f5-4468-bcca-5726b6fecb5c.mp4
+
🌟 Nhiều hình thức triển khai hơn đang chờ bạn khám phá!
## 📦 Cài đặt
-### Cài đặt bằng binary biên dịch sẵn
+### Tải từ picoclaw.io (Khuyến nghị)
-Tải file binary cho nền tảng của bạn từ [trang Release](https://github.com/sipeed/picoclaw/releases).
+Truy cập **[picoclaw.io](https://picoclaw.io)** — trang web chính thức tự động phát hiện nền tảng của bạn và cung cấp tải xuống một cú nhấp. Không cần chọn kiến trúc thủ công.
-### Cài đặt từ mã nguồn (có tính năng mới nhất, khuyên dùng cho phát triển)
+### Tải binary đã biên dịch sẵn
+
+Hoặc tải binary cho nền tảng của bạn từ trang [GitHub Releases](https://github.com/sipeed/picoclaw/releases).
+
+### Biên dịch từ mã nguồn (cho phát triển)
```bash
git clone https://github.com/sipeed/picoclaw.git
@@ -136,444 +184,29 @@ make build
# Build cho nhiều nền tảng
make build-all
+# Build cho Raspberry Pi Zero 2 W (32-bit: make build-linux-arm; 64-bit: make build-linux-arm64)
+make build-pi-zero
+
# Build và cài đặt
make install
```
-## 🐳 Docker Compose
-
-Bạn cũng có thể chạy PicoClaw bằng Docker Compose mà không cần cài đặt gì trên máy.
-
-```bash
-# 1. Clone repo
-git clone https://github.com/sipeed/picoclaw.git
-cd picoclaw
-
-# 2. Lần chạy đầu tiên — tự tạo docker/data/config.json rồi dừng lại
-docker compose -f docker/docker-compose.yml --profile gateway up
-# Container hiển thị "First-run setup complete." rồi tự dừng.
-
-# 3. Thiết lập API Key
-vim docker/data/config.json # API key của provider, bot token, v.v.
-
-# 4. Khởi động
-docker compose -f docker/docker-compose.yml --profile gateway up -d
-```
-
-> [!TIP]
-> **Người dùng Docker**: Theo mặc định, Gateway lắng nghe trên `127.0.0.1`, không thể truy cập từ máy chủ. Nếu bạn cần truy cập các endpoint kiểm tra sức khỏe hoặc mở cổng, hãy đặt `PICOCLAW_GATEWAY_HOST=0.0.0.0` trong môi trường của bạn hoặc cập nhật `config.json`.
-
-```bash
-# 5. Xem logs
-docker compose -f docker/docker-compose.yml logs -f picoclaw-gateway
-
-# 6. Dừng
-docker compose -f docker/docker-compose.yml --profile gateway down
-```
-
-### Chế độ Agent (chạy một lần)
-
-```bash
-# Đặt câu hỏi
-docker compose -f docker/docker-compose.yml run --rm picoclaw-agent -m "2+2 bằng mấy?"
-
-# Chế độ tương tác
-docker compose -f docker/docker-compose.yml run --rm picoclaw-agent
-```
-
-### Cập nhật
-
-```bash
-docker compose -f docker/docker-compose.yml pull
-docker compose -f docker/docker-compose.yml --profile gateway up -d
-```
-
-### 🚀 Bắt đầu nhanh
-
-> [!TIP]
-> Thiết lập API key trong `~/.picoclaw/config.json`. Lấy API key: [Volcengine (CodingPlan)](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) (LLM) · [OpenRouter](https://openrouter.ai/keys) (LLM) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) (LLM). Tìm kiếm web là **tùy chọn** — lấy [Tavily API](https://tavily.com) miễn phí (1000 truy vấn/tháng) hoặc [Brave Search API](https://brave.com/search/api) (2000 truy vấn/tháng).
-
-**1. Khởi tạo**
-
-```bash
-picoclaw onboard
-```
-
-**2. Cấu hình** (`~/.picoclaw/config.json`)
-
-```json
-{
- "model_list": [
- {
- "model_name": "ark-code-latest",
- "model": "volcengine/ark-code-latest",
- "api_key": "sk-your-api-key",
- "api_base":"https://ark.cn-beijing.volces.com/api/coding/v3"
- },
- {
- "model_name": "gpt-5.4",
- "model": "openai/gpt-5.4",
- "api_key": "sk-your-openai-key",
- "request_timeout": 300,
- "api_base": "https://api.openai.com/v1"
- }
- ],
- "agents": {
- "defaults": {
- "model_name": "gpt4"
- }
- },
- "channels": {
- "telegram": {
- "enabled": true,
- "token": "YOUR_TELEGRAM_BOT_TOKEN",
- "allow_from": []
- }
- }
-}
-```
-
-> **Mới**: Định dạng cấu hình `model_list` cho phép thêm nhà cung cấp mà không cần thay đổi mã nguồn. Xem [Cấu hình Mô hình](#cấu-hình-mô-hình-model_list) để biết chi tiết.
-> `request_timeout` là tùy chọn và dùng đơn vị giây. Nếu bỏ qua hoặc đặt `<= 0`, PicoClaw sẽ dùng timeout mặc định (120s).
-
-**3. Lấy API Key**
-
-* **Nhà cung cấp LLM**: [OpenRouter](https://openrouter.ai/keys) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) · [Anthropic](https://console.anthropic.com) · [OpenAI](https://platform.openai.com) · [Gemini](https://aistudio.google.com/api-keys)
-* **Tìm kiếm Web** (tùy chọn): [Brave Search](https://brave.com/search/api) — Có gói miễn phí (2000 truy vấn/tháng)
-
-> **Lưu ý**: Xem `config.example.json` để có mẫu cấu hình đầy đủ.
-
-**4. Trò chuyện**
-
-```bash
-picoclaw agent -m "Xin chào, bạn là ai?"
-```
-
-Vậy là xong! Bạn đã có một trợ lý AI hoạt động chỉ trong 2 phút.
-
----
-
-## 💬 Tích hợp ứng dụng Chat
-
-Trò chuyện với PicoClaw qua Telegram, Discord, DingTalk, LINE hoặc WeCom.
-
-| Kênh | Mức độ thiết lập |
-| --- | --- |
-| **Telegram** | Dễ (chỉ cần token) |
-| **Discord** | Dễ (bot token + intents) |
-| **QQ** | Dễ (AppID + AppSecret) |
-| **DingTalk** | Trung bình (app credentials) |
-| **LINE** | Trung bình (credentials + webhook URL) |
-| **WeCom AI Bot** | Trung bình (Token + khóa AES) |
-
-
-Telegram (Khuyên dùng)
-
-**1. Tạo bot**
-
-* Mở Telegram, tìm `@BotFather`
-* Gửi `/newbot`, làm theo hướng dẫn
-* Sao chép token
-
-**2. Cấu hình**
-
-```json
-{
- "channels": {
- "telegram": {
- "enabled": true,
- "token": "YOUR_BOT_TOKEN",
- "allow_from": ["YOUR_USER_ID"]
- }
- }
-}
-```
-
-> Lấy User ID từ `@userinfobot` trên Telegram.
-
-**3. Chạy**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-Discord
-
-**1. Tạo bot**
-
-* Truy cập
-* Create an application → Bot → Add Bot
-* Sao chép bot token
-
-**2. Bật Intents**
-
-* Trong phần Bot settings, bật **MESSAGE CONTENT INTENT**
-* (Tùy chọn) Bật **SERVER MEMBERS INTENT** nếu muốn dùng danh sách cho phép theo thông tin thành viên
-
-**3. Lấy User ID**
-
-* Discord Settings → Advanced → bật **Developer Mode**
-* Click chuột phải vào avatar → **Copy User ID**
-
-**4. Cấu hình**
-
-```json
-{
- "channels": {
- "discord": {
- "enabled": true,
- "token": "YOUR_BOT_TOKEN",
- "allow_from": ["YOUR_USER_ID"]
- }
- }
-}
-```
-
-**5. Mời bot vào server**
-
-* OAuth2 → URL Generator
-* Scopes: `bot`
-* Bot Permissions: `Send Messages`, `Read Message History`
-* Mở URL mời được tạo và thêm bot vào server của bạn
-
-**6. Chạy**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-QQ
-
-**1. Tạo bot**
-
-* Truy cập [QQ Open Platform](https://q.qq.com/#)
-* Tạo ứng dụng → Lấy **AppID** và **AppSecret**
-
-**2. Cấu hình**
-
-```json
-{
- "channels": {
- "qq": {
- "enabled": true,
- "app_id": "YOUR_APP_ID",
- "app_secret": "YOUR_APP_SECRET",
- "allow_from": []
- }
- }
-}
-```
-
-> Để `allow_from` trống để cho phép tất cả người dùng, hoặc chỉ định số QQ để giới hạn quyền truy cập.
-
-**3. Chạy**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-DingTalk
-
-**1. Tạo bot**
-
-* Truy cập [Open Platform](https://open.dingtalk.com/)
-* Tạo ứng dụng nội bộ
-* Sao chép Client ID và Client Secret
-
-**2. Cấu hình**
-
-```json
-{
- "channels": {
- "dingtalk": {
- "enabled": true,
- "client_id": "YOUR_CLIENT_ID",
- "client_secret": "YOUR_CLIENT_SECRET",
- "allow_from": []
- }
- }
-}
-```
-
-> Để `allow_from` trống để cho phép tất cả người dùng, hoặc chỉ định ID để giới hạn quyền truy cập.
-
-**3. Chạy**
-
-```bash
-picoclaw gateway
-```
-
-
-
-
-LINE
-
-**1. Tạo tài khoản LINE Official**
-
-- Truy cập [LINE Developers Console](https://developers.line.biz/)
-- Tạo provider → Tạo Messaging API channel
-- Sao chép **Channel Secret** và **Channel Access Token**
-
-**2. Cấu hình**
-
-```json
-{
- "channels": {
- "line": {
- "enabled": true,
- "channel_secret": "YOUR_CHANNEL_SECRET",
- "channel_access_token": "YOUR_CHANNEL_ACCESS_TOKEN",
- "webhook_path": "/webhook/line",
- "allow_from": []
- }
- }
-}
-```
-
-**3. Thiết lập Webhook URL**
-
-LINE yêu cầu HTTPS cho webhook. Sử dụng reverse proxy hoặc tunnel:
-
-```bash
-# Ví dụ với ngrok
-ngrok http 18790
-```
-
-Sau đó cài đặt Webhook URL trong LINE Developers Console thành `https://your-domain/webhook/line` và bật **Use webhook**.
-
-**4. Chạy**
-
-```bash
-picoclaw gateway
-```
-
-> Trong nhóm chat, bot chỉ phản hồi khi được @mention. Các câu trả lời sẽ trích dẫn tin nhắn gốc.
-
-> **Docker Compose**: Nếu bạn cần mở port webhook cục bộ, hãy thêm một rule chuyển tiếp từ port Gateway (mặc định 18790) tới host. Lưu ý: LINE webhook được phục vụ bởi Gateway HTTP chung (mặc định 127.0.0.1:18790).
-
-
-
-
-WeCom (WeChat Work)
-
-PicoClaw hỗ trợ ba loại tích hợp WeCom:
-
-**Tùy chọn 1: WeCom Bot (Robot)** - Thiết lập dễ dàng hơn, hỗ trợ chat nhóm
-**Tùy chọn 2: WeCom App (Ứng dụng Tùy chỉnh)** - Nhiều tính năng hơn, nhắn tin chủ động, chỉ chat riêng tư
-**Tùy chọn 3: WeCom AI Bot (Bot Thông Minh)** - Bot AI chính thức, phản hồi streaming, hỗ trợ nhóm và riêng tư
-
-Xem [Hướng dẫn Cấu hình WeCom AI Bot](docs/channels/wecom/wecom_aibot/README.zh.md) để biết hướng dẫn chi tiết.
-
-**Thiết lập Nhanh - WeCom Bot:**
-
-**1. Tạo bot**
-
-* Truy cập Bảng điều khiển Quản trị WeCom → Chat Nhóm → Thêm Bot Nhóm
-* Sao chép URL webhook (định dạng: `https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=xxx`)
-
-**2. Cấu hình**
-
-```json
-{
- "channels": {
- "wecom": {
- "enabled": true,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
- "webhook_url": "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=YOUR_KEY",
- "webhook_path": "/webhook/wecom",
- "allow_from": []
- }
- }
-}
-```
-
-> **Lưu ý:** Các endpoint webhook của WeCom Bot được phục vụ bởi máy chủ Gateway HTTP dùng chung (mặc định 127.0.0.1:18790). Nếu bạn cần truy cập từ bên ngoài, hãy cấu hình reverse proxy hoặc mở cổng Gateway tương ứng.
-
-**Thiết lập Nhanh - WeCom App:**
-
-**1. Tạo ứng dụng**
-
-* Truy cập Bảng điều khiển Quản trị WeCom → Quản lý Ứng dụng → Tạo Ứng dụng
-* Sao chép **AgentId** và **Secret**
-* Truy cập trang "Công ty của tôi", sao chép **CorpID**
-
-**2. Cấu hình nhận tin nhắn**
-
-* Trong chi tiết ứng dụng, nhấp vào "Nhận Tin nhắn" → "Thiết lập API"
-* Đặt URL thành `http://your-server:18790/webhook/wecom-app`
-* Tạo **Token** và **EncodingAESKey**
-
-**3. Cấu hình**
-
-```json
-{
- "channels": {
- "wecom_app": {
- "enabled": true,
- "corp_id": "wwxxxxxxxxxxxxxxxx",
- "corp_secret": "YOUR_CORP_SECRET",
- "agent_id": 1000002,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
- "webhook_path": "/webhook/wecom-app",
- "allow_from": []
- }
- }
-}
-```
-
-**4. Chạy**
-
-```bash
-picoclaw gateway
-```
-
-> **Lưu ý**: WeCom App callback webhook được phục vụ bởi Gateway HTTP chung (mặc định 127.0.0.1:18790). Sử dụng proxy ngược để cung cấp HTTPS trong môi trường production nếu cần.
-
-**Thiết lập Nhanh - WeCom AI Bot:**
-
-**1. Tạo AI Bot**
-
-* Truy cập Bảng điều khiển Quản trị WeCom → Quản lý Ứng dụng → AI Bot
-* Cấu hình URL callback: `http://your-server:18791/webhook/wecom-aibot`
-* Sao chép **Token** và tạo **EncodingAESKey**
-
-**2. Cấu hình**
-
-```json
-{
- "channels": {
- "wecom_aibot": {
- "enabled": true,
- "token": "YOUR_TOKEN",
- "encoding_aes_key": "YOUR_43_CHAR_ENCODING_AES_KEY",
- "webhook_path": "/webhook/wecom-aibot",
- "allow_from": [],
- "welcome_message": "Xin chào! Tôi có thể giúp gì cho bạn?"
- }
- }
-}
-```
-
-**3. Chạy**
-
-```bash
-picoclaw gateway
-```
-
-> **Lưu ý**: WeCom AI Bot sử dụng giao thức pull streaming — không lo timeout phản hồi. Tác vụ dài (>5,5 phút) tự động chuyển sang gửi qua `response_url`.
-
-
+**Raspberry Pi Zero 2 W:** Sử dụng binary phù hợp với hệ điều hành: Raspberry Pi OS 32-bit → `make build-linux-arm`; 64-bit → `make build-linux-arm64`. Hoặc chạy `make build-pi-zero` để build cả hai.
+
+## 📚 Tài liệu
+
+Để xem hướng dẫn chi tiết, tham khảo tài liệu bên dưới. README này chỉ bao gồm phần bắt đầu nhanh.
+
+| Chủ đề | Mô tả |
+|--------|-------|
+| 🐳 [Docker & Bắt đầu nhanh](docs/vi/docker.md) | Thiết lập Docker Compose, chế độ Launcher/Agent, cấu hình Bắt đầu nhanh |
+| 💬 [Ứng dụng Chat](docs/vi/chat-apps.md) | Telegram, Discord, WhatsApp, Matrix, QQ, Slack, IRC, DingTalk, LINE, Feishu, WeCom và nhiều hơn |
+| ⚙️ [Cấu hình](docs/vi/configuration.md) | Biến môi trường, cấu trúc workspace, nguồn skill, sandbox bảo mật, heartbeat |
+| 🔌 [Nhà cung cấp & Mô hình](docs/vi/providers.md) | 20+ nhà cung cấp LLM, định tuyến mô hình, cấu hình model_list, kiến trúc nhà cung cấp |
+| 🔄 [Spawn & Tác vụ bất đồng bộ](docs/vi/spawn-tasks.md) | Tác vụ nhanh, tác vụ dài với spawn, điều phối sub-agent bất đồng bộ |
+| 🐛 [Xử lý sự cố](docs/vi/troubleshooting.md) | Các vấn đề thường gặp và giải pháp |
+| 🔧 [Cấu hình Công cụ](docs/vi/tools_configuration.md) | Bật/tắt từng công cụ, chính sách thực thi |
+| 📋 [Tương Thích Phần Cứng](docs/hardware-compatibility.md) | Các board đã kiểm tra, yêu cầu tối thiểu, cách thêm board |
## Tham gia Mạng xã hội Agent
@@ -581,6 +214,27 @@ Kết nối PicoClaw với Mạng xã hội Agent chỉ bằng cách gửi một
**Đọc `https://clawdchat.ai/skill.md` và làm theo hướng dẫn để tham gia [ClawdChat.ai](https://clawdchat.ai)**
+<<<<<<< HEAD
+## 🖥️ Tham chiếu CLI
+
+| Lệnh | Mô tả |
+| -------------------------- | ------------------------------ |
+| `picoclaw onboard` | Khởi tạo cấu hình & workspace |
+| `picoclaw agent -m "..."` | Trò chuyện với agent |
+| `picoclaw agent` | Chế độ chat tương tác |
+| `picoclaw gateway` | Khởi động gateway |
+| `picoclaw status` | Hiển thị trạng thái |
+| `picoclaw version` | Hiển thị thông tin phiên bản |
+| `picoclaw cron list` | Liệt kê tất cả tác vụ định kỳ |
+| `picoclaw cron add ...` | Thêm tác vụ định kỳ |
+| `picoclaw cron disable` | Tắt tác vụ định kỳ |
+| `picoclaw cron remove` | Xóa tác vụ định kỳ |
+| `picoclaw skills list` | Liệt kê các skill đã cài |
+| `picoclaw skills install` | Cài đặt một skill |
+| `picoclaw migrate` | Di chuyển dữ liệu từ phiên bản cũ |
+| `picoclaw auth login` | Xác thực với nhà cung cấp |
+| `picoclaw model` | Xem hoặc chuyển đổi model mặc định |
+=======
## ⚙️ Cấu hình chi tiết
File cấu hình: `~/.picoclaw/config.json`
@@ -1118,85 +772,26 @@ Xem hướng dẫn chuyển đổi chi tiết tại [docs/migration/model-list-m
| `picoclaw status` | Hiển thị trạng thái |
| `picoclaw cron list` | Liệt kê tất cả tác vụ định kỳ |
| `picoclaw cron add ...` | Thêm tác vụ định kỳ |
+>>>>>>> refactor/agent
### Tác vụ định kỳ / Nhắc nhở
PicoClaw hỗ trợ nhắc nhở theo lịch và tác vụ lặp lại thông qua công cụ `cron`:
-* **Nhắc nhở một lần**: "Remind me in 10 minutes" (Nhắc tôi sau 10 phút) → kích hoạt một lần sau 10 phút
-* **Tác vụ lặp lại**: "Remind me every 2 hours" (Nhắc tôi mỗi 2 giờ) → kích hoạt mỗi 2 giờ
-* **Biểu thức Cron**: "Remind me at 9am daily" (Nhắc tôi lúc 9 giờ sáng mỗi ngày) → sử dụng biểu thức cron
-
-Các tác vụ được lưu trong `~/.picoclaw/workspace/cron/` và được xử lý tự động.
+* **Nhắc nhở một lần**: "Nhắc tôi sau 10 phút" → kích hoạt một lần sau 10 phút
+* **Tác vụ lặp lại**: "Nhắc tôi mỗi 2 giờ" → kích hoạt mỗi 2 giờ
+* **Biểu thức Cron**: "Nhắc tôi lúc 9 giờ sáng mỗi ngày" → sử dụng biểu thức cron
## 🤝 Đóng góp & Lộ trình
Chào đón mọi PR! Mã nguồn được thiết kế nhỏ gọn và dễ đọc. 🤗
-Lộ trình sắp được công bố...
+Xem [Lộ trình Cộng đồng](https://github.com/sipeed/picoclaw/blob/main/ROADMAP.md) đầy đủ.
-Nhóm phát triển đang được xây dựng. Điều kiện tham gia: Ít nhất 1 PR đã được merge.
+Nhóm phát triển đang được xây dựng. Tham gia sau khi có PR đầu tiên được merge!
Nhóm người dùng:
-Discord:
+discord:
-
-## 🐛 Xử lý sự cố
-
-### Tìm kiếm web hiện "API 配置问题"
-
-Điều này là bình thường nếu bạn chưa cấu hình API key cho tìm kiếm. PicoClaw sẽ cung cấp các liên kết hữu ích để tìm kiếm thủ công.
-
-Để bật tìm kiếm web:
-
-1. **Tùy chọn 1 (Khuyên dùng)**: Lấy API key miễn phí tại [https://brave.com/search/api](https://brave.com/search/api) (2000 truy vấn miễn phí/tháng) để có kết quả tốt nhất.
-2. **Tùy chọn 2 (Không cần thẻ tín dụng)**: Nếu không có key, hệ thống tự động chuyển sang dùng **DuckDuckGo** (không cần key).
-
-Thêm key vào `~/.picoclaw/config.json` nếu dùng Brave:
-
-```json
-{
- "tools": {
- "web": {
- "brave": {
- "enabled": false,
- "api_key": "YOUR_BRAVE_API_KEY",
- "max_results": 5
- },
- "duckduckgo": {
- "enabled": true,
- "max_results": 5
- }
- }
- }
-}
-```
-
-### Gặp lỗi lọc nội dung (Content Filtering)
-
-Một số nhà cung cấp (như Zhipu) có bộ lọc nội dung nghiêm ngặt. Thử diễn đạt lại câu hỏi hoặc sử dụng model khác.
-
-### Telegram bot báo "Conflict: terminated by other getUpdates"
-
-Điều này xảy ra khi có một instance bot khác đang chạy. Đảm bảo chỉ có một tiến trình `picoclaw gateway` chạy tại một thời điểm.
-
----
-
-## 📝 So sánh API Key
-
-| Dịch vụ | Gói miễn phí | Trường hợp sử dụng |
-| --- | --- | --- |
-| **OpenRouter** | 200K tokens/tháng | Đa model (Claude, GPT-4, v.v.) |
-| **Volcengine CodingPlan** | ¥9.9/tháng đầu | Tốt nhất cho người dùng Trung Quốc, nhiều mô hình SOTA (Doubao, DeepSeek, v.v.) |
-| **Zhipu** | 200K tokens/tháng | Phù hợp cho người dùng Trung Quốc |
-| **Brave Search** | 2000 truy vấn/tháng | Chức năng tìm kiếm web |
-| **Groq** | Có gói miễn phí | Suy luận siêu nhanh (Llama, Mixtral) |
-| **ModelScope** | 2000 yêu cầu/ngày | Suy luận miễn phí (Qwen, GLM, DeepSeek, v.v.) |
-
----
-
-
diff --git a/docs/pt-br/ANTIGRAVITY_AUTH.md b/docs/pt-br/ANTIGRAVITY_AUTH.md
new file mode 100644
index 000000000..d243783cb
--- /dev/null
+++ b/docs/pt-br/ANTIGRAVITY_AUTH.md
@@ -0,0 +1,809 @@
+> Voltar ao [README](../../README.pt-br.md)
+
+# Guia de Autenticação e Integração do Antigravity
+
+## Visão Geral
+
+**Antigravity** (Google Cloud Code Assist) é um provedor de modelos de IA apoiado pelo Google que oferece acesso a modelos como Claude Opus 4.6 e Gemini através da infraestrutura de nuvem do Google. Este documento fornece um guia completo sobre como a autenticação funciona, como buscar modelos e como implementar um novo provedor no PicoClaw.
+
+---
+
+## Índice
+
+1. [Fluxo de Autenticação](#fluxo-de-autenticação)
+2. [Detalhes da Implementação OAuth](#detalhes-da-implementação-oauth)
+3. [Gerenciamento de Tokens](#gerenciamento-de-tokens)
+4. [Busca da Lista de Modelos](#busca-da-lista-de-modelos)
+5. [Rastreamento de Uso](#rastreamento-de-uso)
+6. [Estrutura do Plugin do Provedor](#estrutura-do-plugin-do-provedor)
+7. [Requisitos de Integração](#requisitos-de-integração)
+8. [Endpoints da API](#endpoints-da-api)
+9. [Configuração](#configuração)
+10. [Criando um Novo Provedor no PicoClaw](#criando-um-novo-provedor-no-picoclaw)
+
+---
+
+## Fluxo de Autenticação
+
+### 1. OAuth 2.0 com PKCE
+
+O Antigravity utiliza **OAuth 2.0 com PKCE (Proof Key for Code Exchange)** para autenticação segura:
+
+```
+┌─────────────┐ ┌─────────────────┐
+│ Client │ ───(1) Generate PKCE Pair────────> │ │
+│ │ ───(2) Open Auth URL─────────────> │ Google OAuth │
+│ │ │ Server │
+│ │ <──(3) Redirect with Code───────── │ │
+│ │ └─────────────────┘
+│ │ ───(4) Exchange Code for Tokens──> │ Token URL │
+│ │ │ │
+│ │ <──(5) Access + Refresh Tokens──── │ │
+└─────────────┘ └─────────────────┘
+```
+
+### 2. Etapas Detalhadas
+
+#### Etapa 1: Gerar Parâmetros PKCE
+```typescript
+function generatePkce(): { verifier: string; challenge: string } {
+ const verifier = randomBytes(32).toString("hex");
+ const challenge = createHash("sha256").update(verifier).digest("base64url");
+ return { verifier, challenge };
+}
+```
+
+#### Etapa 2: Construir a URL de Autorização
+```typescript
+const AUTH_URL = "https://accounts.google.com/o/oauth2/v2/auth";
+const REDIRECT_URI = "http://localhost:51121/oauth-callback";
+
+function buildAuthUrl(params: { challenge: string; state: string }): string {
+ const url = new URL(AUTH_URL);
+ url.searchParams.set("client_id", CLIENT_ID);
+ url.searchParams.set("response_type", "code");
+ url.searchParams.set("redirect_uri", REDIRECT_URI);
+ url.searchParams.set("scope", SCOPES.join(" "));
+ url.searchParams.set("code_challenge", params.challenge);
+ url.searchParams.set("code_challenge_method", "S256");
+ url.searchParams.set("state", params.state);
+ url.searchParams.set("access_type", "offline");
+ url.searchParams.set("prompt", "consent");
+ return url.toString();
+}
+```
+
+**Escopos Necessários:**
+```typescript
+const SCOPES = [
+ "https://www.googleapis.com/auth/cloud-platform",
+ "https://www.googleapis.com/auth/userinfo.email",
+ "https://www.googleapis.com/auth/userinfo.profile",
+ "https://www.googleapis.com/auth/cclog",
+ "https://www.googleapis.com/auth/experimentsandconfigs",
+];
+```
+
+#### Etapa 3: Tratar o Callback OAuth
+
+**Modo Automático (Desenvolvimento Local):**
+- Iniciar um servidor HTTP local na porta 51121
+- Aguardar o redirecionamento do Google
+- Extrair o código de autorização dos parâmetros da query
+
+**Modo Manual (Remoto/Sem Interface Gráfica):**
+- Exibir a URL de autorização para o usuário
+- O usuário completa a autenticação no navegador
+- O usuário cola a URL de redirecionamento completa no terminal
+- Analisar o código da URL colada
+
+#### Etapa 4: Trocar o Código por Tokens
+```typescript
+const TOKEN_URL = "https://oauth2.googleapis.com/token";
+
+async function exchangeCode(params: {
+ code: string;
+ verifier: string;
+}): Promise<{ access: string; refresh: string; expires: number }> {
+ const response = await fetch(TOKEN_URL, {
+ method: "POST",
+ headers: { "Content-Type": "application/x-www-form-urlencoded" },
+ body: new URLSearchParams({
+ client_id: CLIENT_ID,
+ client_secret: CLIENT_SECRET,
+ code: params.code,
+ grant_type: "authorization_code",
+ redirect_uri: REDIRECT_URI,
+ code_verifier: params.verifier,
+ }),
+ });
+
+ const data = await response.json();
+
+ return {
+ access: data.access_token,
+ refresh: data.refresh_token,
+ expires: Date.now() + data.expires_in * 1000 - 5 * 60 * 1000, // 5 min buffer
+ };
+}
+```
+
+#### Etapa 5: Buscar Dados Adicionais do Usuário
+
+**E-mail do Usuário:**
+```typescript
+async function fetchUserEmail(accessToken: string): Promise {
+ const response = await fetch(
+ "https://www.googleapis.com/oauth2/v1/userinfo?alt=json",
+ { headers: { Authorization: `Bearer ${accessToken}` } }
+ );
+ const data = await response.json();
+ return data.email;
+}
+```
+
+**ID do Projeto (Necessário para chamadas de API):**
+```typescript
+async function fetchProjectId(accessToken: string): Promise {
+ const headers = {
+ Authorization: `Bearer ${accessToken}`,
+ "Content-Type": "application/json",
+ "User-Agent": "google-api-nodejs-client/9.15.1",
+ "X-Goog-Api-Client": "google-cloud-sdk vscode_cloudshelleditor/0.1",
+ "Client-Metadata": JSON.stringify({
+ ideType: "IDE_UNSPECIFIED",
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+ }),
+ };
+
+ const response = await fetch(
+ "https://cloudcode-pa.googleapis.com/v1internal:loadCodeAssist",
+ {
+ method: "POST",
+ headers,
+ body: JSON.stringify({
+ metadata: {
+ ideType: "IDE_UNSPECIFIED",
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+ },
+ }),
+ }
+ );
+
+ const data = await response.json();
+ return data.cloudaicompanionProject || "rising-fact-p41fc"; // Valor padrão de fallback
+}
+```
+
+---
+
+## Detalhes da Implementação OAuth
+
+### Credenciais do Cliente
+
+**Importante:** Estas são codificadas em base64 no código-fonte para sincronização com pi-ai:
+
+```typescript
+const decode = (s: string) => Buffer.from(s, "base64").toString();
+
+const CLIENT_ID = decode(
+ "MTA3MTAwNjA2MDU5MS10bWhzc2luMmgyMWxjcmUyMzV2dG9sb2poNGc0MDNlcC5hcHBzLmdvb2dsZXVzZXJjb250ZW50LmNvbQ=="
+);
+const CLIENT_SECRET = decode("R09DU1BYLUs1OEZXUjQ4NkxkTEoxbUxCOHNYQzR6NnFEQWY=");
+```
+
+### Modos do Fluxo OAuth
+
+1. **Fluxo Automático** (máquinas locais com navegador):
+ - Abre o navegador automaticamente
+ - O servidor de callback local captura o redirecionamento
+ - Nenhuma interação do usuário necessária após a autenticação inicial
+
+2. **Fluxo Manual** (remoto/sem interface/WSL2):
+ - URL exibida para copiar e colar manualmente
+ - O usuário completa a autenticação em um navegador externo
+ - O usuário cola a URL de redirecionamento completa de volta
+
+```typescript
+function shouldUseManualOAuthFlow(isRemote: boolean): boolean {
+ return isRemote || isWSL2Sync();
+}
+```
+
+---
+
+## Gerenciamento de Tokens
+
+### Estrutura do Perfil de Autenticação
+
+```typescript
+type OAuthCredential = {
+ type: "oauth";
+ provider: "google-antigravity";
+ access: string; // Token de acesso
+ refresh: string; // Token de atualização
+ expires: number; // Timestamp de expiração (ms desde epoch)
+ email?: string; // E-mail do usuário
+ projectId?: string; // ID do projeto Google Cloud
+};
+```
+
+### Atualização de Tokens
+
+A credencial inclui um token de atualização que pode ser usado para obter novos tokens de acesso quando o atual expira. A expiração é definida com um buffer de 5 minutos para evitar condições de corrida.
+
+---
+
+## Busca da Lista de Modelos
+
+### Buscar Modelos Disponíveis
+
+```typescript
+const BASE_URL = "https://cloudcode-pa.googleapis.com";
+
+async function fetchAvailableModels(
+ accessToken: string,
+ projectId: string
+): Promise {
+ const headers = {
+ Authorization: `Bearer ${accessToken}`,
+ "Content-Type": "application/json",
+ "User-Agent": "antigravity",
+ "X-Goog-Api-Client": "google-cloud-sdk vscode_cloudshelleditor/0.1",
+ };
+
+ const response = await fetch(
+ `${BASE_URL}/v1internal:fetchAvailableModels`,
+ {
+ method: "POST",
+ headers,
+ body: JSON.stringify({ project: projectId }),
+ }
+ );
+
+ const data = await response.json();
+
+ // Retorna modelos com informações de cota
+ return Object.entries(data.models).map(([modelId, modelInfo]) => ({
+ id: modelId,
+ displayName: modelInfo.displayName,
+ quotaInfo: {
+ remainingFraction: modelInfo.quotaInfo?.remainingFraction,
+ resetTime: modelInfo.quotaInfo?.resetTime,
+ isExhausted: modelInfo.quotaInfo?.isExhausted,
+ },
+ }));
+}
+```
+
+### Formato da Resposta
+
+```typescript
+type FetchAvailableModelsResponse = {
+ models?: Record;
+};
+```
+
+---
+
+## Rastreamento de Uso
+
+### Buscar Dados de Uso
+
+```typescript
+export async function fetchAntigravityUsage(
+ token: string,
+ timeoutMs: number
+): Promise {
+ // 1. Buscar créditos e informações do plano
+ const loadCodeAssistRes = await fetch(
+ `${BASE_URL}/v1internal:loadCodeAssist`,
+ {
+ method: "POST",
+ headers: {
+ Authorization: `Bearer ${token}`,
+ "Content-Type": "application/json",
+ },
+ body: JSON.stringify({
+ metadata: {
+ ideType: "ANTIGRAVITY",
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+ },
+ }),
+ }
+ );
+
+ // Extrair informações de créditos
+ const { availablePromptCredits, planInfo, currentTier } = data;
+
+ // 2. Buscar cotas dos modelos
+ const modelsRes = await fetch(
+ `${BASE_URL}/v1internal:fetchAvailableModels`,
+ {
+ method: "POST",
+ headers: { Authorization: `Bearer ${token}` },
+ body: JSON.stringify({ project: projectId }),
+ }
+ );
+
+ // Construir janelas de uso
+ return {
+ provider: "google-antigravity",
+ displayName: "Google Antigravity",
+ windows: [
+ { label: "Credits", usedPercent: calculateUsedPercent(available, monthly) },
+ // Cotas individuais dos modelos...
+ ],
+ plan: currentTier?.name || planType,
+ };
+}
+```
+
+### Estrutura da Resposta de Uso
+
+```typescript
+type ProviderUsageSnapshot = {
+ provider: "google-antigravity";
+ displayName: string;
+ windows: UsageWindow[];
+ plan?: string;
+ error?: string;
+};
+
+type UsageWindow = {
+ label: string; // "Credits" ou ID do modelo
+ usedPercent: number; // 0-100
+ resetAt?: number; // Timestamp de quando a cota é redefinida
+};
+```
+
+---
+
+## Estrutura do Plugin do Provedor
+
+### Definição do Plugin
+
+```typescript
+const antigravityPlugin = {
+ id: "google-antigravity-auth",
+ name: "Google Antigravity Auth",
+ description: "OAuth flow for Google Antigravity (Cloud Code Assist)",
+ configSchema: emptyPluginConfigSchema(),
+
+ register(api: PicoClawPluginApi) {
+ api.registerProvider({
+ id: "google-antigravity",
+ label: "Google Antigravity",
+ docsPath: "/providers/models",
+ aliases: ["antigravity"],
+
+ auth: [
+ {
+ id: "oauth",
+ label: "Google OAuth",
+ hint: "PKCE + localhost callback",
+ kind: "oauth",
+ run: async (ctx: ProviderAuthContext) => {
+ // Implementação OAuth aqui
+ },
+ },
+ ],
+ });
+ },
+};
+```
+
+### ProviderAuthContext
+
+```typescript
+type ProviderAuthContext = {
+ config: PicoClawConfig;
+ agentDir?: string;
+ workspaceDir?: string;
+ prompter: WizardPrompter; // Prompts/notificações da UI
+ runtime: RuntimeEnv; // Logging, etc.
+ isRemote: boolean; // Se está executando remotamente
+ openUrl: (url: string) => Promise; // Abridor de navegador
+ oauth: {
+ createVpsAwareHandlers: Function;
+ };
+};
+```
+
+### ProviderAuthResult
+
+```typescript
+type ProviderAuthResult = {
+ profiles: Array<{
+ profileId: string;
+ credential: AuthProfileCredential;
+ }>;
+ configPatch?: Partial;
+ defaultModel?: string;
+ notes?: string[];
+};
+```
+
+---
+
+## Requisitos de Integração
+
+### 1. Ambiente/Dependências Necessários
+
+- Go ≥ 1.25
+- Base de código do PicoClaw (`pkg/providers/` e `pkg/auth/`)
+- Pacotes da biblioteca padrão `crypto` e `net/http`
+
+### 2. Cabeçalhos Necessários para Chamadas de API
+
+```typescript
+const REQUIRED_HEADERS = {
+ "Authorization": `Bearer ${accessToken}`,
+ "Content-Type": "application/json",
+ "User-Agent": "antigravity", // ou "google-api-nodejs-client/9.15.1"
+ "X-Goog-Api-Client": "google-cloud-sdk vscode_cloudshelleditor/0.1",
+};
+
+// Para chamadas loadCodeAssist, incluir também:
+const CLIENT_METADATA = {
+ ideType: "ANTIGRAVITY", // ou "IDE_UNSPECIFIED"
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+};
+```
+
+### 3. Sanitização de Schemas de Modelos
+
+O Antigravity usa modelos compatíveis com Gemini, então os schemas de ferramentas devem ser sanitizados:
+
+```typescript
+const GOOGLE_SCHEMA_UNSUPPORTED_KEYWORDS = new Set([
+ "patternProperties",
+ "additionalProperties",
+ "$schema",
+ "$id",
+ "$ref",
+ "$defs",
+ "definitions",
+ "examples",
+ "minLength",
+ "maxLength",
+ "minimum",
+ "maximum",
+ "multipleOf",
+ "pattern",
+ "format",
+ "minItems",
+ "maxItems",
+ "uniqueItems",
+ "minProperties",
+ "maxProperties",
+]);
+
+// Limpar schema antes de enviar
+function cleanToolSchemaForGemini(schema: Record): unknown {
+ // Remover palavras-chave não suportadas
+ // Garantir que o nível superior tenha type: "object"
+ // Achatar uniões anyOf/oneOf
+}
+```
+
+### 4. Tratamento de Blocos de Pensamento (Modelos Claude)
+
+Para modelos Claude via Antigravity, os blocos de pensamento requerem tratamento especial:
+
+```typescript
+const ANTIGRAVITY_SIGNATURE_RE = /^[A-Za-z0-9+/]+={0,2}$/;
+
+export function sanitizeAntigravityThinkingBlocks(
+ messages: AgentMessage[]
+): AgentMessage[] {
+ // Validar assinaturas de pensamento
+ // Normalizar campos de assinatura
+ // Descartar blocos de pensamento não assinados
+}
+```
+
+---
+
+## Endpoints da API
+
+### Endpoints de Autenticação
+
+| Endpoint | Método | Finalidade |
+|----------|--------|-----------|
+| `https://accounts.google.com/o/oauth2/v2/auth` | GET | Autorização OAuth |
+| `https://oauth2.googleapis.com/token` | POST | Troca de tokens |
+| `https://www.googleapis.com/oauth2/v1/userinfo` | GET | Informações do usuário (e-mail) |
+
+### Endpoints do Cloud Code Assist
+
+| Endpoint | Método | Finalidade |
+|----------|--------|-----------|
+| `https://cloudcode-pa.googleapis.com/v1internal:loadCodeAssist` | POST | Carregar informações do projeto, créditos, plano |
+| `https://cloudcode-pa.googleapis.com/v1internal:fetchAvailableModels` | POST | Listar modelos disponíveis com cotas |
+| `https://cloudcode-pa.googleapis.com/v1internal:streamGenerateContent?alt=sse` | POST | Endpoint de streaming de chat |
+
+**Formato de Requisição da API (Chat):**
+O endpoint `v1internal:streamGenerateContent` espera um envelope encapsulando a requisição Gemini padrão:
+
+```json
+{
+ "project": "your-project-id",
+ "model": "model-id",
+ "request": {
+ "contents": [...],
+ "systemInstruction": {...},
+ "generationConfig": {...},
+ "tools": [...]
+ },
+ "requestType": "agent",
+ "userAgent": "antigravity",
+ "requestId": "agent-timestamp-random"
+}
+```
+
+**Formato de Resposta da API (SSE):**
+Cada mensagem SSE (`data: {...}`) é encapsulada em um campo `response`:
+
+```json
+{
+ "response": {
+ "candidates": [...],
+ "usageMetadata": {...},
+ "modelVersion": "...",
+ "responseId": "..."
+ },
+ "traceId": "...",
+ "metadata": {}
+}
+```
+
+---
+
+## Configuração
+
+### Configuração do config.json
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "gemini-flash",
+ "model": "antigravity/gemini-3-flash",
+ "auth_method": "oauth"
+ }
+ ],
+ "agents": {
+ "defaults": {
+ "model_name": "gemini-flash"
+ }
+ }
+}
+```
+
+### Armazenamento do Perfil de Autenticação
+
+Os perfis de autenticação são armazenados em `~/.picoclaw/auth.json`:
+
+```json
+{
+ "credentials": {
+ "google-antigravity": {
+ "access_token": "ya29...",
+ "refresh_token": "1//...",
+ "expires_at": "2026-01-01T00:00:00Z",
+ "provider": "google-antigravity",
+ "auth_method": "oauth",
+ "email": "user@example.com",
+ "project_id": "my-project-id"
+ }
+ }
+}
+```
+
+---
+
+## Criando um Novo Provedor no PicoClaw
+
+Os provedores do PicoClaw são implementados como pacotes Go em `pkg/providers/`. Para adicionar um novo provedor:
+
+### Implementação Passo a Passo
+
+#### 1. Criar o Arquivo do Provedor
+
+Crie um novo arquivo Go em `pkg/providers/`:
+
+```
+pkg/providers/
+└── your_provider.go
+```
+
+#### 2. Implementar a Interface Provider
+
+Seu provedor deve implementar a interface `Provider` definida em `pkg/providers/types.go`:
+
+```go
+package providers
+
+type YourProvider struct {
+ apiKey string
+ apiBase string
+}
+
+func NewYourProvider(apiKey, apiBase, proxy string) *YourProvider {
+ if apiBase == "" {
+ apiBase = "https://api.your-provider.com/v1"
+ }
+ return &YourProvider{apiKey: apiKey, apiBase: apiBase}
+}
+
+func (p *YourProvider) Chat(ctx context.Context, messages []Message, tools []Tool, cb StreamCallback) error {
+ // Implementar conclusão de chat com streaming
+}
+```
+
+#### 3. Registrar na Factory
+
+Adicione seu provedor ao switch de protocolo em `pkg/providers/factory.go`:
+
+```go
+case "your-provider":
+ return NewYourProvider(sel.apiKey, sel.apiBase, sel.proxy), nil
+```
+
+#### 4. Adicionar Configuração Padrão (Opcional)
+
+Adicione uma entrada padrão em `pkg/config/defaults.go`:
+
+```go
+{
+ ModelName: "your-model",
+ Model: "your-provider/model-name",
+ APIKey: "",
+},
+```
+
+#### 5. Adicionar Suporte de Autenticação (Opcional)
+
+Se seu provedor requer OAuth ou autenticação especial, adicione um caso em `cmd/picoclaw/internal/auth/helpers.go`:
+
+```go
+case "your-provider":
+ authLoginYourProvider()
+```
+
+#### 6. Configurar via `config.json`
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "your-model",
+ "model": "your-provider/model-name",
+ "api_key": "your-api-key",
+ "api_base": "https://api.your-provider.com/v1"
+ }
+ ]
+}
+```
+
+---
+
+## Testando Sua Implementação
+
+### Comandos CLI
+
+```bash
+# Autenticar com um provedor
+picoclaw auth login --provider your-provider
+
+# Listar modelos (para Antigravity)
+picoclaw auth models
+
+# Iniciar o gateway
+picoclaw gateway
+
+# Executar um agente com um modelo específico
+picoclaw agent -m "Hello" --model your-model
+```
+
+### Variáveis de Ambiente para Testes
+
+```bash
+# Substituir o modelo padrão
+export PICOCLAW_AGENTS_DEFAULTS_MODEL=your-model
+
+# Substituir configurações do provedor
+export PICOCLAW_MODEL_LIST='[{"model_name":"your-model","model":"your-provider/model-name","api_key":"..."}]'
+```
+
+---
+
+## Referências
+
+- **Arquivos Fonte:**
+ - `pkg/providers/antigravity_provider.go` - Implementação do provedor Antigravity
+ - `pkg/auth/oauth.go` - Implementação do fluxo OAuth
+ - `pkg/auth/store.go` - Armazenamento de credenciais de autenticação (`~/.picoclaw/auth.json`)
+ - `pkg/providers/factory.go` - Factory de provedores e roteamento de protocolo
+ - `pkg/providers/types.go` - Definições da interface do provedor
+ - `cmd/picoclaw/internal/auth/helpers.go` - Comandos CLI de autenticação
+
+- **Documentação:**
+ - `docs/ANTIGRAVITY_USAGE.md` - Guia de uso do Antigravity
+ - `docs/migration/model-list-migration.md` - Guia de migração
+
+---
+
+## Observações
+
+1. **Projeto Google Cloud:** O Antigravity requer que o Gemini for Google Cloud esteja habilitado no seu projeto Google Cloud
+2. **Cotas:** Usa cotas do projeto Google Cloud (sem cobrança separada)
+3. **Acesso a Modelos:** Os modelos disponíveis dependem da configuração do seu projeto Google Cloud
+4. **Blocos de Pensamento:** Modelos Claude via Antigravity requerem tratamento especial de blocos de pensamento com assinaturas
+5. **Sanitização de Schemas:** Os schemas de ferramentas devem ser sanitizados para remover palavras-chave JSON Schema não suportadas
+
+---
+
+---
+
+## Tratamento de Erros Comuns
+
+### 1. Limitação de Taxa (HTTP 429)
+
+O Antigravity retorna um erro 429 quando as cotas do projeto/modelo estão esgotadas. A resposta de erro frequentemente contém um `quotaResetDelay` no campo `details`.
+
+**Exemplo de Erro 429:**
+```json
+{
+ "error": {
+ "code": 429,
+ "message": "You have exhausted your capacity on this model. Your quota will reset after 4h30m28s.",
+ "status": "RESOURCE_EXHAUSTED",
+ "details": [
+ {
+ "@type": "type.googleapis.com/google.rpc.ErrorInfo",
+ "metadata": {
+ "quotaResetDelay": "4h30m28.060903746s"
+ }
+ }
+ ]
+ }
+}
+```
+
+### 2. Respostas Vazias (Modelos Restritos)
+
+Alguns modelos podem aparecer na lista de modelos disponíveis, mas retornar uma resposta vazia (200 OK mas stream SSE vazio). Isso geralmente acontece com modelos em preview ou restritos que o projeto atual não tem permissão para usar.
+
+**Tratamento:** Tratar respostas vazias como erros informando ao usuário que o modelo pode estar restrito ou inválido para seu projeto.
+
+---
+
+## Solução de Problemas
+
+### "Token expired" (token expirado)
+- Atualizar tokens OAuth: `picoclaw auth login --provider antigravity`
+
+### "Gemini for Google Cloud is not enabled" (Gemini for Google Cloud não está habilitado)
+- Habilitar a API no seu Google Cloud Console
+
+### "Project not found" (projeto não encontrado)
+- Verificar se seu projeto Google Cloud tem as APIs necessárias habilitadas
+- Verificar se o ID do projeto foi obtido corretamente durante a autenticação
+
+### Modelos não aparecem na lista
+- Verificar se a autenticação OAuth foi concluída com sucesso
+- Verificar o armazenamento do perfil de autenticação: `~/.picoclaw/auth.json`
+- Executar novamente `picoclaw auth login --provider antigravity`
diff --git a/docs/pt-br/ANTIGRAVITY_USAGE.md b/docs/pt-br/ANTIGRAVITY_USAGE.md
new file mode 100644
index 000000000..d4b681ad0
--- /dev/null
+++ b/docs/pt-br/ANTIGRAVITY_USAGE.md
@@ -0,0 +1,72 @@
+> Voltar ao [README](../../README.pt-br.md)
+
+# Usando o provedor Antigravity no PicoClaw
+
+Este guia explica como configurar e usar o provedor **Antigravity** (Google Cloud Code Assist) no PicoClaw.
+
+## Pré-requisitos
+
+1. Uma conta Google.
+2. Google Cloud Code Assist habilitado (geralmente disponível através da integração "Gemini for Google Cloud").
+
+## 1. Autenticação
+
+Para se autenticar com o Antigravity, execute o seguinte comando:
+
+```bash
+picoclaw auth login --provider antigravity
+```
+
+### Autenticação manual (Headless/VPS)
+Se você está executando em um servidor (Coolify/Docker) e não consegue acessar `localhost`, siga estas etapas:
+1. Execute o comando acima.
+2. Copie a URL fornecida e abra-a no seu navegador local.
+3. Complete o login.
+4. Seu navegador será redirecionado para uma URL `localhost:51121` (que não carregará).
+5. **Copie essa URL final** da barra de endereços do seu navegador.
+6. **Cole-a de volta no terminal** onde o PicoClaw está aguardando.
+
+O PicoClaw extrairá automaticamente o código de autorização e completará o processo.
+
+## 2. Gerenciando modelos
+
+### Listar modelos disponíveis
+Para ver quais modelos seu projeto tem acesso e verificar suas cotas:
+
+```bash
+picoclaw auth models
+```
+
+### Trocar de modelo
+Você pode alterar o modelo padrão em `~/.picoclaw/config.json` ou substituí-lo via CLI:
+
+```bash
+# Substituir para um único comando
+picoclaw agent -m "Hello" --model claude-opus-4-6-thinking
+```
+
+## 3. Uso em produção (Coolify/Docker)
+
+Se você está implantando via Coolify ou Docker, siga estas etapas para testar:
+
+1. **Variáveis de ambiente**:
+ * `PICOCLAW_AGENTS_DEFAULTS_MODEL=gemini-flash`
+2. **Persistência da autenticação**:
+ Se você já fez login localmente, pode copiar suas credenciais para o servidor:
+ ```bash
+ scp ~/.picoclaw/auth.json user@your-server:~/.picoclaw/
+ ```
+ *Alternativamente*, execute o comando `auth login` uma vez no servidor se você tiver acesso ao terminal.
+
+## 4. Solução de problemas
+
+* **Resposta vazia**: Se um modelo retorna uma resposta vazia, ele pode estar restrito para o seu projeto. Tente `gemini-3-flash` ou `claude-opus-4-6-thinking`.
+* **429 Limite de taxa**: O Antigravity possui cotas rigorosas. O PicoClaw exibirá o "tempo de redefinição" na mensagem de erro se você atingir um limite.
+* **404 Não encontrado**: Certifique-se de que está usando um ID de modelo da lista `picoclaw auth models`. Use o ID curto (ex.: `gemini-3-flash`) e não o caminho completo.
+
+## 5. Resumo dos modelos funcionais
+
+Com base nos testes, os seguintes modelos são os mais confiáveis:
+* `gemini-3-flash` (Rápido, alta disponibilidade)
+* `gemini-2.5-flash-lite` (Leve)
+* `claude-opus-4-6-thinking` (Poderoso, inclui raciocínio)
diff --git a/docs/pt-br/chat-apps.md b/docs/pt-br/chat-apps.md
new file mode 100644
index 000000000..08ef292fa
--- /dev/null
+++ b/docs/pt-br/chat-apps.md
@@ -0,0 +1,624 @@
+# 💬 Configuração de Aplicativos de Chat
+
+> Voltar ao [README](../../README.pt-br.md)
+
+## 💬 Aplicativos de Chat
+
+Converse com seu picoclaw através do Telegram, Discord, WhatsApp, Matrix, QQ, DingTalk, LINE, WeCom, Feishu, Slack, IRC, OneBot ou MaixCam
+
+> **Nota**: Todos os canais baseados em webhook (LINE, WeCom, etc.) são servidos em um único servidor HTTP Gateway compartilhado (`gateway.host`:`gateway.port`, padrão `127.0.0.1:18790`). Não há portas por canal para configurar. Nota: Feishu usa o modo WebSocket/SDK e não utiliza o servidor HTTP webhook compartilhado.
+
+| Canal | Dificuldade | Descrição | Documentação |
+| -------------------- | ------------------ | ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- |
+| **Telegram** | ⭐ Fácil | Recomendado, voz para texto, long polling (sem IP público) | [Documentação](../channels/telegram/README.pt-br.md) |
+| **Discord** | ⭐ Fácil | Socket Mode, suporte a grupos/DM, ecossistema bot rico | [Documentação](../channels/discord/README.pt-br.md) |
+| **WhatsApp** | ⭐ Fácil | Nativo (scan QR) ou Bridge URL | [Documentação](#whatsapp) |
+| **Slack** | ⭐ Fácil | **Socket Mode** (sem IP público), empresarial | [Documentação](../channels/slack/README.pt-br.md) |
+| **Matrix** | ⭐⭐ Médio | Protocolo federado, suporte a auto-hospedagem | [Documentação](../channels/matrix/README.pt-br.md) |
+| **QQ** | ⭐⭐ Médio | API bot oficial, comunidade chinesa | [Documentação](../channels/qq/README.pt-br.md) |
+| **DingTalk** | ⭐⭐ Médio | Modo Stream (sem IP público), empresarial | [Documentação](../channels/dingtalk/README.pt-br.md) |
+| **LINE** | ⭐⭐⭐ Avançado | HTTPS Webhook obrigatório | [Documentação](../channels/line/README.pt-br.md) |
+| **WeCom (企业微信)** | ⭐⭐⭐ Avançado | Bot de grupo (Webhook), app personalizado (API), AI Bot | [Bot](../channels/wecom/wecom_bot/README.pt-br.md) / [App](../channels/wecom/wecom_app/README.pt-br.md) / [AI Bot](../channels/wecom/wecom_aibot/README.pt-br.md) |
+| **Feishu (飞书)** | ⭐⭐⭐ Avançado | Colaboração empresarial, rico em recursos | [Documentação](../channels/feishu/README.pt-br.md) |
+| **IRC** | ⭐⭐ Médio | Servidor + configuração TLS | - |
+| **OneBot** | ⭐⭐ Médio | Compatível com NapCat/Go-CQHTTP, ecossistema comunitário | [Documentação](../channels/onebot/README.pt-br.md) |
+| **MaixCam** | ⭐ Fácil | Canal de integração de hardware para câmeras AI Sipeed | [Documentação](../channels/maixcam/README.pt-br.md) |
+| **Pico** | ⭐ Fácil | Canal de protocolo nativo PicoClaw | |
+
+
+Telegram (Recomendado)
+
+**1. Criar um bot**
+
+* Abra o Telegram, pesquise `@BotFather`
+* Envie `/newbot`, siga as instruções
+* Copie o token
+
+**2. Configurar**
+
+```json
+{
+ "channels": {
+ "telegram": {
+ "enabled": true,
+ "token": "YOUR_BOT_TOKEN",
+ "allow_from": ["YOUR_USER_ID"]
+ }
+ }
+}
+```
+
+> Obtenha seu ID de usuário com `@userinfobot` no Telegram.
+
+**3. Executar**
+
+```bash
+picoclaw gateway
+```
+
+**4. Menu de comandos do Telegram (registrado automaticamente na inicialização)**
+
+O PicoClaw agora mantém definições de comandos em um registro compartilhado. Na inicialização, o Telegram registrará automaticamente os comandos de bot suportados (por exemplo `/start`, `/help`, `/show`, `/list`) para que o menu de comandos e o comportamento em tempo de execução permaneçam sincronizados.
+O registro do menu de comandos do Telegram permanece como descoberta UX local do canal; a execução genérica de comandos é tratada centralmente no loop do agente via commands executor.
+
+Se o registro de comandos falhar (erros transitórios de rede/API), o canal ainda inicia e o PicoClaw tenta novamente o registro em segundo plano.
+
+
+
+
+Discord
+
+**1. Criar um bot**
+
+* Acesse
+* Crie um aplicativo → Bot → Add Bot
+* Copie o token do bot
+
+**2. Habilitar intents**
+
+* Nas configurações do Bot, habilite **MESSAGE CONTENT INTENT**
+* (Opcional) Habilite **SERVER MEMBERS INTENT** se planeja usar listas de permissão baseadas em dados de membros
+
+**3. Obter seu User ID**
+* Configurações do Discord → Avançado → habilite **Developer Mode**
+* Clique com o botão direito no seu avatar → **Copy User ID**
+
+**4. Configurar**
+
+```json
+{
+ "channels": {
+ "discord": {
+ "enabled": true,
+ "token": "YOUR_BOT_TOKEN",
+ "allow_from": ["YOUR_USER_ID"]
+ }
+ }
+}
+```
+
+**5. Convidar o bot**
+
+* OAuth2 → URL Generator
+* Scopes: `bot`
+* Bot Permissions: `Send Messages`, `Read Message History`
+* Abra a URL de convite gerada e adicione o bot ao seu servidor
+
+**Opcional: Modo de ativação em grupo**
+
+Por padrão, o bot responde a todas as mensagens em um canal do servidor. Para restringir respostas apenas a @menções, adicione:
+
+```json
+{
+ "channels": {
+ "discord": {
+ "group_trigger": { "mention_only": true }
+ }
+ }
+}
+```
+
+Você também pode ativar por prefixos de palavras-chave (ex.: `!bot`):
+
+```json
+{
+ "channels": {
+ "discord": {
+ "group_trigger": { "prefixes": ["!bot"] }
+ }
+ }
+}
+```
+
+**6. Executar**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+WhatsApp (nativo via whatsmeow)
+
+O PicoClaw pode se conectar ao WhatsApp de duas formas:
+
+- **Nativo (recomendado):** In-process usando [whatsmeow](https://github.com/tulir/whatsmeow). Sem bridge separado. Defina `"use_native": true` e deixe `bridge_url` vazio. Na primeira execução, escaneie o QR code com o WhatsApp (Dispositivos Vinculados). A sessão é armazenada no seu workspace (ex.: `workspace/whatsapp/`). O canal nativo é **opcional** para manter o binário padrão pequeno; compile com `-tags whatsapp_native` (ex.: `make build-whatsapp-native` ou `go build -tags whatsapp_native ./cmd/...`).
+- **Bridge:** Conecte-se a um bridge WebSocket externo. Defina `bridge_url` (ex.: `ws://localhost:3001`) e mantenha `use_native` como false.
+
+**Configurar (nativo)**
+
+```json
+{
+ "channels": {
+ "whatsapp": {
+ "enabled": true,
+ "use_native": true,
+ "session_store_path": "",
+ "allow_from": []
+ }
+ }
+}
+```
+
+Se `session_store_path` estiver vazio, a sessão é armazenada em `/whatsapp/`. Execute `picoclaw gateway`; na primeira execução, escaneie o QR code impresso no terminal com WhatsApp → Dispositivos Vinculados.
+
+
+
+
+QQ
+
+**Configuração rápida (recomendada)**
+
+A QQ Open Platform oferece uma página de configuração com um clique para bots compatíveis com OpenClaw:
+
+1. Abra o [QQ Bot Quick Start](https://q.qq.com/qqbot/openclaw/index.html) e escaneie o QR code para fazer login
+2. Um bot é criado automaticamente — copie o **App ID** e o **App Secret**
+3. Configure o PicoClaw:
+
+```json
+{
+ "channels": {
+ "qq": {
+ "enabled": true,
+ "app_id": "YOUR_APP_ID",
+ "app_secret": "YOUR_APP_SECRET",
+ "allow_from": []
+ }
+ }
+}
+```
+
+4. Execute `picoclaw gateway` e abra o QQ para conversar com seu bot
+
+> O App Secret é exibido apenas uma vez. Salve-o imediatamente — visualizá-lo novamente forçará uma redefinição.
+>
+> Bots criados pela página de configuração rápida são inicialmente apenas para o criador e não suportam chats de grupo. Para habilitar o acesso em grupo, configure o modo sandbox na [QQ Open Platform](https://q.qq.com/).
+
+**Configuração manual**
+
+Se preferir criar o bot manualmente:
+
+* Faça login na [QQ Open Platform](https://q.qq.com/) para se registrar como desenvolvedor
+* Crie um bot QQ — personalize seu avatar e nome
+* Copie o **App ID** e o **App Secret** nas configurações do bot
+* Configure conforme mostrado acima e execute `picoclaw gateway`
+
+
+
+
+DingTalk
+
+**1. Criar um bot**
+
+* Acesse a [Open Platform](https://open.dingtalk.com/)
+* Crie um aplicativo interno
+* Copie o Client ID e o Client Secret
+
+**2. Configurar**
+
+```json
+{
+ "channels": {
+ "dingtalk": {
+ "enabled": true,
+ "client_id": "YOUR_CLIENT_ID",
+ "client_secret": "YOUR_CLIENT_SECRET",
+ "allow_from": []
+ }
+ }
+}
+```
+
+> Defina `allow_from` como vazio para permitir todos os usuários, ou especifique IDs de usuário DingTalk para restringir o acesso.
+
+**3. Executar**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+MaixCam
+
+Canal de integração projetado especificamente para hardware de câmera AI Sipeed.
+
+```json
+{
+ "channels": {
+ "maixcam": {
+ "enabled": true
+ }
+ }
+}
+```
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+
+Matrix
+
+**1. Preparar conta do bot**
+
+* Use seu homeserver preferido (ex.: `https://matrix.org` ou auto-hospedado)
+* Crie um usuário bot e obtenha seu access token
+
+**2. Configurar**
+
+```json
+{
+ "channels": {
+ "matrix": {
+ "enabled": true,
+ "homeserver": "https://matrix.org",
+ "user_id": "@your-bot:matrix.org",
+ "access_token": "YOUR_MATRIX_ACCESS_TOKEN",
+ "allow_from": []
+ }
+ }
+}
+```
+
+**3. Executar**
+
+```bash
+picoclaw gateway
+```
+
+Para opções completas (`device_id`, `join_on_invite`, `group_trigger`, `placeholder`, `reasoning_channel_id`), veja o [Guia de Configuração do Canal Matrix](../channels/matrix/README.md).
+
+
+
+
+LINE
+
+**1. Criar uma Conta Oficial LINE**
+
+- Acesse o [LINE Developers Console](https://developers.line.biz/)
+- Crie um provider → Crie um canal Messaging API
+- Copie o **Channel Secret** e o **Channel Access Token**
+
+**2. Configurar**
+
+```json
+{
+ "channels": {
+ "line": {
+ "enabled": true,
+ "channel_secret": "YOUR_CHANNEL_SECRET",
+ "channel_access_token": "YOUR_CHANNEL_ACCESS_TOKEN",
+ "webhook_path": "/webhook/line",
+ "allow_from": []
+ }
+ }
+}
+```
+
+> O webhook do LINE é servido no servidor Gateway compartilhado (`gateway.host`:`gateway.port`, padrão `127.0.0.1:18790`).
+
+**3. Configurar URL do Webhook**
+
+O LINE requer HTTPS para webhooks. Use um proxy reverso ou túnel:
+
+```bash
+# Exemplo com ngrok (porta padrão do gateway é 18790)
+ngrok http 18790
+```
+
+Em seguida, defina a URL do Webhook no LINE Developers Console como `https://your-domain/webhook/line` e habilite **Use webhook**.
+
+**4. Executar**
+
+```bash
+picoclaw gateway
+```
+
+> Em chats de grupo, o bot responde apenas quando @mencionado. As respostas citam a mensagem original.
+
+
+
+
+WeCom (企业微信)
+
+O PicoClaw suporta três tipos de integração WeCom:
+
+**Opção 1: WeCom Bot (Bot)** - Configuração mais fácil, suporta chats de grupo
+**Opção 2: WeCom App (App Personalizado)** - Mais recursos, mensagens proativas, apenas chat privado
+**Opção 3: WeCom AI Bot (AI Bot)** - AI Bot oficial, respostas em streaming, suporta chat de grupo e privado
+
+Veja o [Guia de Configuração do WeCom AI Bot](../channels/wecom/wecom_aibot/README.pt-br.md) para instruções detalhadas de configuração.
+
+**Configuração Rápida - WeCom Bot:**
+
+**1. Criar um bot**
+
+* Acesse o Console de Administração WeCom → Chat de Grupo → Adicionar Bot de Grupo
+* Copie a URL do webhook (formato: `https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=xxx`)
+
+**2. Configurar**
+
+```json
+{
+ "channels": {
+ "wecom": {
+ "enabled": true,
+ "token": "YOUR_TOKEN",
+ "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
+ "webhook_url": "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=YOUR_KEY",
+ "webhook_path": "/webhook/wecom",
+ "allow_from": []
+ }
+ }
+}
+```
+
+> O webhook do WeCom é servido no servidor Gateway compartilhado (`gateway.host`:`gateway.port`, padrão `127.0.0.1:18790`).
+
+**Configuração Rápida - WeCom App:**
+
+**1. Criar um aplicativo**
+
+* Acesse o Console de Administração WeCom → Gerenciamento de Apps → Criar App
+* Copie o **AgentId** e o **Secret**
+* Acesse a página "Minha Empresa", copie o **CorpID**
+
+**2. Configurar recebimento de mensagens**
+
+* Nos detalhes do App, clique em "Receber Mensagem" → "Configurar API"
+* Defina a URL como `http://your-server:18790/webhook/wecom-app`
+* Gere o **Token** e o **EncodingAESKey**
+
+**3. Configurar**
+
+```json
+{
+ "channels": {
+ "wecom_app": {
+ "enabled": true,
+ "corp_id": "wwxxxxxxxxxxxxxxxx",
+ "corp_secret": "YOUR_CORP_SECRET",
+ "agent_id": 1000002,
+ "token": "YOUR_TOKEN",
+ "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
+ "webhook_path": "/webhook/wecom-app",
+ "allow_from": []
+ }
+ }
+}
+```
+
+**4. Executar**
+
+```bash
+picoclaw gateway
+```
+
+> **Nota**: Os callbacks de webhook do WeCom são servidos na porta do Gateway (padrão 18790). Use um proxy reverso para HTTPS.
+
+**Configuração Rápida - WeCom AI Bot:**
+
+**1. Criar um AI Bot**
+
+* Acesse o Console de Administração WeCom → Gerenciamento de Apps → AI Bot
+* Nas configurações do AI Bot, configure a URL de callback: `http://your-server:18790/webhook/wecom-aibot`
+* Copie o **Token** e clique em "Gerar Aleatoriamente" para o **EncodingAESKey**
+
+**2. Configurar**
+
+```json
+{
+ "channels": {
+ "wecom_aibot": {
+ "enabled": true,
+ "token": "YOUR_TOKEN",
+ "encoding_aes_key": "YOUR_43_CHAR_ENCODING_AES_KEY",
+ "webhook_path": "/webhook/wecom-aibot",
+ "allow_from": [],
+ "welcome_message": "Hello! How can I help you?"
+ }
+ }
+}
+```
+
+**3. Executar**
+
+```bash
+picoclaw gateway
+```
+
+> **Nota**: O WeCom AI Bot usa protocolo de streaming pull — sem preocupações com timeout de resposta. Tarefas longas (>30 segundos) mudam automaticamente para entrega via `response_url` push.
+
+
+
+
+Feishu (Lark)
+
+O PicoClaw se conecta ao Feishu via modo WebSocket/SDK — não é necessário URL de webhook público nem servidor de callback.
+
+**1. Criar um aplicativo**
+
+* Acesse a [Feishu Open Platform](https://open.feishu.cn/) e crie um aplicativo
+* Nas configurações do aplicativo, habilite a capacidade **Bot**
+* Crie uma versão e publique o aplicativo (o aplicativo deve ser publicado para funcionar)
+* Copie o **App ID** (começa com `cli_`) e o **App Secret**
+
+**2. Configurar**
+
+```json
+{
+ "channels": {
+ "feishu": {
+ "enabled": true,
+ "app_id": "cli_xxx",
+ "app_secret": "YOUR_APP_SECRET",
+ "allow_from": []
+ }
+ }
+}
+```
+
+Opcional: `encrypt_key` e `verification_token` para criptografia de eventos (recomendado para produção).
+
+**3. Executar e conversar**
+
+```bash
+picoclaw gateway
+```
+
+Abra o Feishu, pesquise o nome do seu bot e comece a conversar. Você também pode adicionar o bot a um grupo — use `group_trigger.mention_only: true` para responder apenas quando @mencionado.
+
+Para opções completas, veja o [Guia de Configuração do Canal Feishu](../channels/feishu/README.pt-br.md).
+
+
+
+
+Slack
+
+**1. Criar um aplicativo Slack**
+
+* Acesse a [Slack API](https://api.slack.com/apps) e crie um novo aplicativo
+* Em **OAuth & Permissions**, adicione os escopos do bot: `chat:write`, `app_mentions:read`, `im:history`, `im:read`, `im:write`
+* Instale o aplicativo no seu workspace
+* Copie o **Bot Token** (`xoxb-...`) e o **App-Level Token** (`xapp-...`, habilite Socket Mode para obtê-lo)
+
+**2. Configurar**
+
+```json
+{
+ "channels": {
+ "slack": {
+ "enabled": true,
+ "bot_token": "xoxb-YOUR-BOT-TOKEN",
+ "app_token": "xapp-YOUR-APP-TOKEN",
+ "allow_from": []
+ }
+ }
+}
+```
+
+**3. Executar**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+IRC
+
+**1. Configurar**
+
+```json
+{
+ "channels": {
+ "irc": {
+ "enabled": true,
+ "server": "irc.libera.chat:6697",
+ "tls": true,
+ "nick": "picoclaw-bot",
+ "channels": ["#your-channel"],
+ "password": "",
+ "allow_from": []
+ }
+ }
+}
+```
+
+Opcional: `nickserv_password` para autenticação NickServ, `sasl_user`/`sasl_password` para autenticação SASL.
+
+**2. Executar**
+
+```bash
+picoclaw gateway
+```
+
+O bot se conectará ao servidor IRC e entrará nos canais especificados.
+
+
+
+
+OneBot (QQ via protocolo OneBot)
+
+OneBot é um protocolo aberto para bots QQ. O PicoClaw se conecta a qualquer implementação compatível com OneBot v11 (ex.: [Lagrange](https://github.com/LagrangeDev/Lagrange.Core), [NapCat](https://github.com/NapNeko/NapCatQQ)) via WebSocket.
+
+**1. Configurar uma implementação OneBot**
+
+Instale e execute um framework de bot QQ compatível com OneBot v11. Habilite seu servidor WebSocket.
+
+**2. Configurar**
+
+```json
+{
+ "channels": {
+ "onebot": {
+ "enabled": true,
+ "ws_url": "ws://127.0.0.1:8080",
+ "access_token": "",
+ "allow_from": []
+ }
+ }
+}
+```
+
+| Campo | Descrição |
+|-------|-----------|
+| `ws_url` | URL WebSocket da implementação OneBot |
+| `access_token` | Token de acesso para autenticação (se configurado no OneBot) |
+| `reconnect_interval` | Intervalo de reconexão em segundos (padrão: 5) |
+
+**3. Executar**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+MaixCam
+
+Canal de integração projetado especificamente para hardware de câmera AI Sipeed.
+
+```json
+{
+ "channels": {
+ "maixcam": {
+ "enabled": true
+ }
+ }
+}
+```
+
+```bash
+picoclaw gateway
+```
+
+
diff --git a/docs/pt-br/configuration.md b/docs/pt-br/configuration.md
new file mode 100644
index 000000000..ee14ca724
--- /dev/null
+++ b/docs/pt-br/configuration.md
@@ -0,0 +1,219 @@
+# ⚙️ Guia de Configuração
+
+> Voltar ao [README](../../README.pt-br.md)
+
+## ⚙️ Configuração
+
+Arquivo de configuração: `~/.picoclaw/config.json`
+
+### Variáveis de Ambiente
+
+Você pode substituir os caminhos padrão usando variáveis de ambiente. Isso é útil para instalações portáteis, implantações em contêineres ou execução do picoclaw como serviço do sistema. Essas variáveis são independentes e controlam caminhos diferentes.
+
+| Variável | Descrição | Caminho Padrão |
+|-------------------|-----------------------------------------------------------------------------------------------------------------------------------------|---------------------------|
+| `PICOCLAW_CONFIG` | Substitui o caminho para o arquivo de configuração. Isso indica diretamente ao picoclaw qual `config.json` carregar, ignorando todos os outros locais. | `~/.picoclaw/config.json` |
+| `PICOCLAW_HOME` | Substitui o diretório raiz para dados do picoclaw. Isso altera o local padrão do `workspace` e outros diretórios de dados. | `~/.picoclaw` |
+
+**Exemplos:**
+
+```bash
+# Executar picoclaw usando um arquivo de configuração específico
+# O caminho do workspace será lido de dentro desse arquivo de configuração
+PICOCLAW_CONFIG=/etc/picoclaw/production.json picoclaw gateway
+
+# Executar picoclaw com todos os dados armazenados em /opt/picoclaw
+# A configuração será carregada do padrão ~/.picoclaw/config.json
+# O workspace será criado em /opt/picoclaw/workspace
+PICOCLAW_HOME=/opt/picoclaw picoclaw agent
+
+# Usar ambos para uma configuração totalmente personalizada
+PICOCLAW_HOME=/srv/picoclaw PICOCLAW_CONFIG=/srv/picoclaw/main.json picoclaw gateway
+```
+
+### Layout do Workspace
+
+O PicoClaw armazena dados no seu workspace configurado (padrão: `~/.picoclaw/workspace`):
+
+```
+~/.picoclaw/workspace/
+├── sessions/ # Sessões de conversa e histórico
+├── memory/ # Memória de longo prazo (MEMORY.md)
+├── state/ # Estado persistente (último canal, etc.)
+├── cron/ # Banco de dados de tarefas agendadas
+├── skills/ # Skills personalizadas
+├── AGENT.md # Guia de comportamento do agente
+├── HEARTBEAT.md # Prompts de tarefas periódicas (verificados a cada 30 min)
+├── IDENTITY.md # Identidade do agente
+├── SOUL.md # Alma do agente
+└── USER.md # Preferências do usuário
+```
+
+> **Nota:** Alterações em `AGENT.md`, `SOUL.md`, `USER.md` e `memory/MEMORY.md` são detectadas automaticamente em tempo de execução via rastreamento de data de modificação (mtime). **Não é necessário reiniciar o gateway** após editar esses arquivos — o agente carrega o novo conteúdo na próxima requisição.
+
+### Fontes de Skills
+
+Por padrão, as skills são carregadas de:
+
+1. `~/.picoclaw/workspace/skills` (workspace)
+2. `~/.picoclaw/skills` (global)
+3. `/skills` (embutido)
+
+Para configurações avançadas/de teste, você pode substituir o diretório raiz de skills builtin com:
+
+```bash
+export PICOCLAW_BUILTIN_SKILLS=/path/to/skills
+```
+
+### Política Unificada de Execução de Comandos
+
+- Comandos slash genéricos são executados através de um único caminho em `pkg/agent/loop.go` via `commands.Executor`.
+- Os adaptadores de canal não consomem mais comandos genéricos localmente; eles encaminham o texto de entrada para o caminho bus/agent. O Telegram ainda registra automaticamente os comandos suportados na inicialização.
+- Comando slash desconhecido (por exemplo `/foo`) passa para o processamento normal do LLM.
+- Comando registrado mas não suportado no canal atual (por exemplo `/show` no WhatsApp) retorna um erro explícito ao usuário e interrompe o processamento.
+
+### 🔒 Sandbox de Segurança
+
+O PicoClaw é executado em um ambiente sandbox por padrão. O agente só pode acessar arquivos e executar comandos dentro do workspace configurado.
+
+#### Configuração Padrão
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "workspace": "~/.picoclaw/workspace",
+ "restrict_to_workspace": true
+ }
+ }
+}
+```
+
+| Opção | Padrão | Descrição |
+| ----------------------- | ----------------------- | ----------------------------------------- |
+| `workspace` | `~/.picoclaw/workspace` | Diretório de trabalho do agente |
+| `restrict_to_workspace` | `true` | Restringir acesso a arquivos/comandos ao workspace |
+
+#### Ferramentas Protegidas
+
+Quando `restrict_to_workspace: true`, as seguintes ferramentas são isoladas:
+
+| Ferramenta | Função | Restrição |
+| ------------- | ---------------- | -------------------------------------- |
+| `read_file` | Ler arquivos | Apenas arquivos dentro do workspace |
+| `write_file` | Escrever arquivos| Apenas arquivos dentro do workspace |
+| `list_dir` | Listar diretórios| Apenas diretórios dentro do workspace |
+| `edit_file` | Editar arquivos | Apenas arquivos dentro do workspace |
+| `append_file` | Anexar a arquivos| Apenas arquivos dentro do workspace |
+| `exec` | Executar comandos| Caminhos de comando devem estar dentro do workspace |
+
+#### Proteção Adicional do Exec
+
+Mesmo com `restrict_to_workspace: false`, a ferramenta `exec` bloqueia estes comandos perigosos:
+
+* `rm -rf`, `del /f`, `rmdir /s` — Exclusão em massa
+* `format`, `mkfs`, `diskpart` — Formatação de disco
+* `dd if=` — Imagem de disco
+* Escrita em `/dev/sd[a-z]` — Escritas diretas em disco
+* `shutdown`, `reboot`, `poweroff` — Desligamento do sistema
+* Fork bomb `:(){ :|:& };:`
+
+### Controle de Acesso a Arquivos
+
+| Config Key | Type | Default | Description |
+|------------|------|---------|-------------|
+| `tools.allow_read_paths` | string[] | `[]` | Additional paths allowed for reading outside workspace |
+| `tools.allow_write_paths` | string[] | `[]` | Additional paths allowed for writing outside workspace |
+
+### Segurança do Exec
+
+| Config Key | Type | Default | Description |
+|------------|------|---------|-------------|
+| `tools.exec.allow_remote` | bool | `false` | Allow exec tool from remote channels (Telegram/Discord etc.) |
+| `tools.exec.enable_deny_patterns` | bool | `true` | Enable dangerous command interception |
+| `tools.exec.custom_deny_patterns` | string[] | `[]` | Custom regex patterns to block |
+| `tools.exec.custom_allow_patterns` | string[] | `[]` | Custom regex patterns to allow |
+
+> **Nota de Segurança:** A proteção contra symlinks é habilitada por padrão — todos os caminhos de arquivo são resolvidos através de `filepath.EvalSymlinks` antes da correspondência com a whitelist, prevenindo ataques de escape via symlink.
+
+#### Limitação Conhecida: Processos Filhos de Ferramentas de Build
+
+O guard de segurança do exec inspeciona apenas a linha de comando que o PicoClaw executa diretamente. Ele não inspeciona recursivamente processos filhos gerados por ferramentas de desenvolvimento permitidas como `make`, `go run`, `cargo`, `npm run` ou scripts de build personalizados.
+
+Isso significa que um comando de nível superior ainda pode compilar ou executar outros binários após passar pela verificação inicial do guard. Na prática, trate scripts de build, Makefiles, scripts de pacotes e binários gerados como código executável que precisa do mesmo nível de revisão que um comando shell direto.
+
+Para ambientes de maior risco:
+
+* Revise scripts de build antes da execução.
+* Prefira aprovação/revisão manual para fluxos de trabalho de compilação e execução.
+* Execute o PicoClaw dentro de um contêiner ou VM se precisar de isolamento mais forte do que o guard integrado oferece.
+
+#### Exemplos de Erro
+
+```
+[ERROR] tool: Tool execution failed
+{tool=exec, error=Command blocked by safety guard (path outside working dir)}
+```
+
+```
+[ERROR] tool: Tool execution failed
+{tool=exec, error=Command blocked by safety guard (dangerous pattern detected)}
+```
+
+#### Desabilitando Restrições (Risco de Segurança)
+
+Se você precisar que o agente acesse caminhos fora do workspace:
+
+**Método 1: Arquivo de configuração**
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "restrict_to_workspace": false
+ }
+ }
+}
+```
+
+**Método 2: Variável de ambiente**
+
+```bash
+export PICOCLAW_AGENTS_DEFAULTS_RESTRICT_TO_WORKSPACE=false
+```
+
+> ⚠️ **Aviso**: Desabilitar esta restrição permite que o agente acesse qualquer caminho no seu sistema. Use com cautela apenas em ambientes controlados.
+
+#### Consistência do Limite de Segurança
+
+A configuração `restrict_to_workspace` se aplica consistentemente em todos os caminhos de execução:
+
+| Caminho de Execução | Limite de Segurança |
+| -------------------- | ---------------------------- |
+| Main Agent | `restrict_to_workspace` ✅ |
+| Subagent / Spawn | Herda a mesma restrição ✅ |
+| Heartbeat tasks | Herda a mesma restrição ✅ |
+
+Todos os caminhos compartilham a mesma restrição de workspace — não há como contornar o limite de segurança através de subagentes ou tarefas agendadas.
+
+### Heartbeat (Tarefas Periódicas)
+
+O PicoClaw pode executar tarefas periódicas automaticamente. Crie um arquivo `HEARTBEAT.md` no seu workspace:
+
+```markdown
+# Tarefas Periódicas
+
+- Verificar meu e-mail para mensagens importantes
+- Revisar meu calendário para eventos próximos
+- Verificar a previsão do tempo
+```
+
+O agente lerá este arquivo a cada 30 minutos (configurável) e executará quaisquer tarefas usando as ferramentas disponíveis.
+
+#### Tarefas Assíncronas com Spawn
+
+Para tarefas de longa duração (busca na web, chamadas de API), use a ferramenta `spawn` para criar um **subagente**:
+
+```markdown
+# Tarefas Periódicas
+```
diff --git a/docs/pt-br/credential_encryption.md b/docs/pt-br/credential_encryption.md
new file mode 100644
index 000000000..59a31e438
--- /dev/null
+++ b/docs/pt-br/credential_encryption.md
@@ -0,0 +1,159 @@
+> Voltar ao [README](../../README.pt-br.md)
+
+# Criptografia de Credenciais
+
+O PicoClaw suporta a criptografia de valores `api_key` nas entradas de configuração `model_list`.
+As chaves criptografadas são armazenadas como strings `enc://` e descriptografadas automaticamente na inicialização.
+
+---
+
+## Início Rápido
+
+**1. Defina sua frase secreta**
+
+```bash
+export PICOCLAW_KEY_PASSPHRASE="your-passphrase"
+```
+
+**2. Criptografe uma chave de API**
+
+Execute `picoclaw onboard` — ele solicita sua frase secreta e gera a chave SSH,
+depois recriptografa automaticamente quaisquer entradas `api_key` em texto simples na sua configuração
+na próxima chamada `SaveConfig`. O valor `enc://` resultante será semelhante a:
+
+```
+enc://AAAA...base64...
+```
+
+**3. Cole a saída na sua configuração**
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "gpt-4o",
+ "model": "openai/gpt-4o",
+ "api_key": "enc://AAAA...base64...",
+ "api_base": "https://api.openai.com/v1"
+ }
+ ]
+}
+```
+
+---
+
+## Formatos de `api_key` Suportados
+
+| Formato | Exemplo | Comportamento |
+|---------|---------|---------------|
+| Texto simples | `sk-abc123` | Usado como está |
+| Referência de arquivo | `file://openai.key` | Conteúdo lido do mesmo diretório do arquivo de configuração |
+| Criptografado | `enc://` | Descriptografado na inicialização usando `PICOCLAW_KEY_PASSPHRASE` |
+| Vazio | `""` | Passado sem alteração (usado com `auth_method: oauth`) |
+
+---
+
+## Design Criptográfico
+
+### Derivação de Chave
+
+A criptografia utiliza **HKDF-SHA256** com uma chave privada SSH como segundo fator.
+
+```
+sshHash = SHA256(ssh_private_key_file_bytes)
+ikm = HMAC-SHA256(key=sshHash, message=passphrase)
+aes_key = HKDF-SHA256(ikm, salt, info="picoclaw-credential-v1", 32 bytes)
+```
+
+### Criptografia
+
+```
+AES-256-GCM(key=aes_key, nonce=random[12], plaintext=api_key)
+```
+
+### Formato de Transmissão
+
+```
+enc://
+```
+
+| Campo | Tamanho | Descrição |
+|-------|---------|-----------|
+| `salt` | 16 bytes | Aleatório por criptografia; alimentado no HKDF |
+| `nonce` | 12 bytes | Aleatório por criptografia; IV do AES-GCM |
+| `ciphertext` | variável | Texto cifrado AES-256-GCM + tag de autenticação de 16 bytes |
+
+O tag de autenticação GCM é anexado automaticamente ao texto cifrado. Qualquer adulteração faz com que a descriptografia falhe com um erro em vez de retornar texto simples corrompido.
+
+### Desempenho
+
+| Operação | Tempo (ARM Cortex-A) |
+|----------|----------------------|
+| Derivação de chave (HKDF) | < 1 ms |
+| Descriptografia AES-256-GCM | < 1 ms |
+| **Sobrecarga total na inicialização** | **< 2 ms por chave** |
+
+---
+
+## Segurança de Dois Fatores com Chave SSH
+
+Quando uma chave privada SSH é fornecida, quebrar a criptografia requer **ambos**:
+
+1. A **frase secreta** (`PICOCLAW_KEY_PASSPHRASE`)
+2. O **arquivo de chave privada SSH**
+
+Isso significa que um arquivo de configuração vazado sozinho não é suficiente para recuperar a chave de API, mesmo que a frase secreta seja fraca. A chave SSH contribui com 256 bits de entropia (Ed25519) independentemente da força da frase secreta.
+
+### Modelo de Ameaça
+
+| O que o atacante possui | Pode descriptografar? |
+|------------------------|----------------------|
+| Apenas o arquivo de configuração | Não — necessita da frase secreta + chave SSH |
+| Apenas a chave SSH | Não — necessita da frase secreta |
+| Apenas a frase secreta | Não — necessita da chave SSH |
+| Arquivo de configuração + chave SSH + frase secreta | Sim — comprometimento total |
+
+---
+
+## Variáveis de Ambiente
+
+| Variável | Obrigatório | Descrição |
+|----------|-------------|-----------|
+| `PICOCLAW_KEY_PASSPHRASE` | Sim (para `enc://`) | Frase secreta usada para derivação de chave |
+| `PICOCLAW_SSH_KEY_PATH` | Não | Caminho para a chave privada SSH. Se não definido, detecta automaticamente em `~/.ssh/picoclaw_ed25519.key` |
+
+### Detecção Automática da Chave SSH
+
+Se `PICOCLAW_SSH_KEY_PATH` não estiver definido, o PicoClaw procura a chave dedicada:
+
+```
+~/.ssh/picoclaw_ed25519.key
+```
+
+Este arquivo dedicado evita conflitos com as chaves SSH existentes do usuário.
+Execute `picoclaw onboard` para gerá-lo automaticamente.
+
+`os.UserHomeDir()` é usado para resolução multiplataforma do diretório home (lê `USERPROFILE` no Windows, `HOME` no Unix/macOS).
+
+> **Nota:** Um arquivo de chave SSH é obrigatório para a criptografia de credenciais. Se nenhuma chave for encontrada e `PICOCLAW_SSH_KEY_PATH` não estiver definido, a criptografia/descriptografia falhará. Execute `picoclaw onboard` para gerar a chave automaticamente.
+
+---
+
+## Migração
+
+Como os únicos materiais secretos são `PICOCLAW_KEY_PASSPHRASE` e o arquivo de chave privada SSH, a migração é simples:
+
+1. Copie o arquivo de configuração para a nova máquina.
+2. Defina `PICOCLAW_KEY_PASSPHRASE` com o mesmo valor.
+3. Copie o arquivo de chave privada SSH para o mesmo caminho (ou defina `PICOCLAW_SSH_KEY_PATH` para sua nova localização).
+
+Nenhuma recriptografia é necessária.
+
+---
+
+## Considerações de Segurança
+
+- **Tanto a frase secreta quanto a chave SSH são obrigatórias.** A chave SSH atua como um segundo fator — sem ela, a criptografia/descriptografia falhará. Execute `picoclaw onboard` para gerar a chave se ela não existir.
+- **A chave SSH é somente leitura em tempo de execução.** O PicoClaw nunca escreve ou modifica o arquivo de chave SSH.
+- **Chaves em texto simples continuam sendo suportadas.** Configurações existentes sem `enc://` não são afetadas.
+- **O formato `enc://` é versionado** através do campo `info` do HKDF (`picoclaw-credential-v1`), permitindo futuras atualizações de algoritmo sem quebrar valores criptografados existentes.
diff --git a/docs/pt-br/debug.md b/docs/pt-br/debug.md
new file mode 100644
index 000000000..8614cd5ed
--- /dev/null
+++ b/docs/pt-br/debug.md
@@ -0,0 +1,36 @@
+# Depuração do PicoClaw
+
+> Voltar ao [README](../../README.pt-br.md)
+
+O PicoClaw realiza múltiplas interações complexas nos bastidores para cada requisição que recebe — desde o roteamento de mensagens e avaliação de complexidade, até a execução de ferramentas e adaptação a falhas de modelo. Poder ver exatamente o que está acontecendo é crucial, não apenas para solucionar problemas potenciais, mas também para realmente entender como o agente opera.
+
+## Iniciando o PicoClaw em modo de depuração
+
+Para obter informações detalhadas sobre o que o agente está fazendo (requisições LLM, chamadas de ferramentas, roteamento de mensagens), você pode iniciar o gateway do PicoClaw com a flag de depuração:
+
+```bash
+picoclaw gateway --debug
+# or
+picoclaw gateway -d
+```
+
+Neste modo, o sistema formata os logs de forma detalhada e exibe prévias dos prompts do sistema e dos resultados de execução das ferramentas.
+
+## Desabilitando a truncagem de logs (logs completos)
+
+Por padrão, o PicoClaw trunca strings muito longas (como o *Prompt do Sistema* ou resultados JSON grandes) nos logs de depuração para manter o console legível.
+
+Se você precisar inspecionar a saída completa de um comando ou o payload exato enviado ao modelo LLM, pode usar a flag `--no-truncate`.
+
+**Nota:** Esta flag *só* funciona quando combinada com o modo `--debug`.
+
+```bash
+picoclaw gateway --debug --no-truncate
+
+```
+
+Quando esta flag está ativa, a função de truncagem global é desabilitada. Isso é extremamente útil para:
+
+* Verificar a sintaxe exata das mensagens enviadas ao provedor.
+* Ler a saída completa de ferramentas como `exec`, `web_fetch` ou `read_file`.
+* Depurar o histórico de sessão salvo na memória.
diff --git a/docs/pt-br/docker.md b/docs/pt-br/docker.md
new file mode 100644
index 000000000..bac48954b
--- /dev/null
+++ b/docs/pt-br/docker.md
@@ -0,0 +1,167 @@
+# 🐳 Docker e Início Rápido
+
+> Voltar ao [README](../../README.pt-br.md)
+
+## 🐳 Docker Compose
+
+Você também pode executar o PicoClaw usando Docker Compose sem instalar nada localmente.
+
+```bash
+# 1. Clone este repositório
+git clone https://github.com/sipeed/picoclaw.git
+cd picoclaw
+
+# 2. Primeira execução — gera automaticamente docker/data/config.json e encerra
+# (só é acionado quando config.json e workspace/ estão ambos ausentes)
+docker compose -f docker/docker-compose.yml --profile gateway up
+# O contêiner exibe "First-run setup complete." e para.
+
+# 3. Configure suas chaves de API
+vim docker/data/config.json # Set provider API keys, bot tokens, etc.
+
+# 4. Iniciar
+docker compose -f docker/docker-compose.yml --profile gateway up -d
+```
+
+> [!TIP]
+> **Usuários Docker**: Por padrão, o Gateway escuta em `127.0.0.1`, que não é acessível a partir do host. Se você precisar acessar os endpoints de saúde ou expor portas, defina `PICOCLAW_GATEWAY_HOST=0.0.0.0` no seu ambiente ou atualize o `config.json`.
+
+```bash
+# 5. Verificar logs
+docker compose -f docker/docker-compose.yml logs -f picoclaw-gateway
+
+# 6. Parar
+docker compose -f docker/docker-compose.yml --profile gateway down
+```
+
+### Modo Launcher (Console Web)
+
+A imagem `launcher` inclui os três binários (`picoclaw`, `picoclaw-launcher`, `picoclaw-launcher-tui`) e inicia o console web por padrão, que fornece uma interface baseada em navegador para configuração e chat.
+
+```bash
+docker compose -f docker/docker-compose.yml --profile launcher up -d
+```
+
+Abra http://localhost:18800 no seu navegador. O launcher gerencia o processo do gateway automaticamente.
+
+> [!WARNING]
+> O console web ainda não suporta autenticação. Evite expô-lo na internet pública.
+
+### Modo Agent (One-shot)
+
+```bash
+# Fazer uma pergunta
+docker compose -f docker/docker-compose.yml run --rm picoclaw-agent -m "What is 2+2?"
+
+# Modo interativo
+docker compose -f docker/docker-compose.yml run --rm picoclaw-agent
+```
+
+### Atualização
+
+```bash
+docker compose -f docker/docker-compose.yml pull
+docker compose -f docker/docker-compose.yml --profile gateway up -d
+```
+
+### 🚀 Início Rápido
+
+> [!TIP]
+> Configure sua chave de API em `~/.picoclaw/config.json`. Obtenha chaves de API: [Volcengine (CodingPlan)](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) (LLM) · [OpenRouter](https://openrouter.ai/keys) (LLM) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) (LLM). A busca na web é opcional — obtenha gratuitamente uma [API Tavily](https://tavily.com) (1000 consultas gratuitas/mês) ou [API Brave Search](https://brave.com/search/api) (2000 consultas gratuitas/mês).
+
+**1. Inicializar**
+
+```bash
+picoclaw onboard
+```
+
+**2. Configurar** (`~/.picoclaw/config.json`)
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "workspace": "~/.picoclaw/workspace",
+ "model_name": "gpt-5.4",
+ "max_tokens": 8192,
+ "temperature": 0.7,
+ "max_tool_iterations": 20
+ }
+ },
+ "model_list": [
+ {
+ "model_name": "ark-code-latest",
+ "model": "volcengine/ark-code-latest",
+ "api_key": "sk-your-api-key",
+ "api_base":"https://ark.cn-beijing.volces.com/api/coding/v3"
+ },
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_key": "your-api-key",
+ "request_timeout": 300
+ },
+ {
+ "model_name": "claude-sonnet-4.6",
+ "model": "anthropic/claude-sonnet-4.6",
+ "api_key": "your-anthropic-key"
+ }
+ ],
+ "tools": {
+ "web": {
+ "enabled": true,
+ "fetch_limit_bytes": 10485760,
+ "format": "plaintext",
+ "brave": {
+ "enabled": false,
+ "api_key": "YOUR_BRAVE_API_KEY",
+ "max_results": 5
+ },
+ "tavily": {
+ "enabled": false,
+ "api_key": "YOUR_TAVILY_API_KEY",
+ "max_results": 5
+ },
+ "duckduckgo": {
+ "enabled": true,
+ "max_results": 5
+ },
+ "perplexity": {
+ "enabled": false,
+ "api_key": "YOUR_PERPLEXITY_API_KEY",
+ "max_results": 5
+ },
+ "searxng": {
+ "enabled": false,
+ "base_url": "http://your-searxng-instance:8888",
+ "max_results": 5
+ }
+ }
+ }
+}
+```
+
+> **Novo**: O formato de configuração `model_list` permite adicionar provedores sem alteração de código. Veja [Configuração de Modelos](#configuração-de-modelos-model_list) para detalhes.
+> `request_timeout` é opcional e usa segundos. Se omitido ou definido como `<= 0`, o PicoClaw usa o timeout padrão (120s).
+
+**3. Obter chaves de API**
+
+* **Provedor LLM**: [OpenRouter](https://openrouter.ai/keys) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) · [Anthropic](https://console.anthropic.com) · [OpenAI](https://platform.openai.com) · [Gemini](https://aistudio.google.com/api-keys)
+* **Busca na Web** (opcional):
+ * [Brave Search](https://brave.com/search/api) - Pago ($5/1000 consultas, ~$5-6/mês)
+ * [Perplexity](https://www.perplexity.ai) - Busca com IA e interface de chat
+ * [SearXNG](https://github.com/searxng/searxng) - Metabuscador auto-hospedado (gratuito, sem necessidade de chave de API)
+ * [Tavily](https://tavily.com) - Otimizado para agentes de IA (1000 requisições/mês)
+ * DuckDuckGo - Fallback integrado (sem necessidade de chave de API)
+
+> **Nota**: Veja `config.example.json` para um modelo de configuração completo.
+
+**4. Conversar**
+
+```bash
+picoclaw agent -m "What is 2+2?"
+```
+
+Pronto! Você tem um assistente de IA funcionando em 2 minutos.
+
+---
diff --git a/docs/pt-br/hardware-compatibility.md b/docs/pt-br/hardware-compatibility.md
new file mode 100644
index 000000000..771621014
--- /dev/null
+++ b/docs/pt-br/hardware-compatibility.md
@@ -0,0 +1,152 @@
+> Voltar ao [README](../../README.pt-br.md)
+
+# 🖥️ PicoClaw Lista de compatibilidade de hardware
+
+O PicoClaw roda em praticamente qualquer dispositivo Linux. Esta página registra chips, produtos e placas de desenvolvimento verificados.
+
+**Seu hardware não está na lista?** Envie um PR para adicioná-lo! Fabricantes de hardware são bem-vindos para contribuir e co-promover.
+
+---
+
+## 1. Suporte a chips verificado
+
+### x86
+
+| Fabricante | Chip | Notas |
+|------------|------|-------|
+| Intel | Any x86 CPU (i386+) | Todos os processadores desktop/servidor/notebook |
+| AMD | Any x86 CPU | Todos os processadores desktop/servidor/notebook |
+
+### ARM
+
+| Sub-arq | Chips típicos | Notas |
+|---------|---------------|-------|
+| ARMv6 | [BCM2835](https://www.raspberrypi.com/documentation/computers/processors.html#bcm2835) (Raspberry Pi 1/Zero) | Single-core ARM1176JZF-S |
+| ARMv7 | [Allwinner V3s](https://linux-sunxi.org/V3s) | Single-core Cortex-A7, usado no LicheePi Zero |
+| ARM64 | [Allwinner H618](https://linux-sunxi.org/H618) | Quad-core Cortex-A53, usado no Orange Pi Zero 3 |
+| ARM64 | [BCM2711](https://www.raspberrypi.com/documentation/computers/processors.html#bcm2711) (Raspberry Pi 4) | Quad-core Cortex-A72 |
+| ARM64 | [BCM2712](https://www.raspberrypi.com/documentation/computers/processors.html#bcm2712) (Raspberry Pi 5) | Quad-core Cortex-A76 |
+| ARM64 | [AX630C](https://www.axera-tech.com/) (爱芯元智) | Dual-core Cortex-A53 + NPU, usado no NanoKVM-Pro / MaixCAM2 |
+
+### RISC-V (riscv64)
+
+| Fabricante | Chip | Núcleo | Notas |
+|------------|------|--------|-------|
+| [SOPHGO (算能)](https://www.sophgo.com/) | SG2002 | C906 @ 1GHz | 256MB DDR3 integrado, usado no LicheeRV-Nano / NanoKVM / MaixCAM |
+| [Allwinner (全志)](https://www.allwinnertech.com/) | V861 | Dual C907 | 128MB DDR3L integrado, 1 TOPS NPU, câmera AI 4K SiP |
+| [Allwinner (全志)](https://www.allwinnertech.com/) | V881 | C907 | Série de câmeras AI RISC-V |
+| [Arterytek (匠芯创)](https://www.arterytek.com/) | D213 | RISC-V | Usado no HaaS506-LD1 RTU industrial |
+| [SpacemiT (进迭)](https://www.spacemit.com/) | K1 | 8x X60 @ 1.8GHz | Usado no Milk-V Jupiter, BananaPi BPI-F3 |
+| [SpacemiT (进迭)](https://www.spacemit.com/) | K3 | 8x X100 @ 2.5GHz | Compatível com RVA23, RVV de 1024 bits, inferência AI FP8 |
+| [Zhihe (知合)](https://www.zhihe-tech.com/) | A210 | High-perf RISC-V | 8 núcleos, 16MB cache L3, classe desktop |
+| [Canaan (嘉楠)](https://www.canaan-creative.com/) | K230 | Dual C908 @ 1.6GHz | 6 TOPS KPU, usado no CanMV-K230 |
+
+### MIPS
+
+| Fabricante | Chip | Notas |
+|------------|------|-------|
+| MediaTek | [MT7620](https://www.mediatek.com/products/home-networking/mt7620) | MIPS24KEc @ 580MHz, usado em muitos roteadores OpenWrt (ex. Xiaomi Router 3G) |
+
+### LoongArch (loong64)
+
+| Fabricante | Chip | Notas |
+|------------|------|-------|
+| [Loongson (龙芯)](https://www.loongson.cn/) | 3A5000 | Quad-core LA464 @ 2.5GHz, desktop/estação de trabalho |
+| [Loongson (龙芯)](https://www.loongson.cn/) | 3A6000 | Quad-core 4C/8T @ 2.5GHz, IPC comparável ao Intel 10ª geração |
+| [Loongson (龙芯)](https://www.loongson.cn/) | 2K1000LA | Dual-core @ 1GHz, aplicações industriais/IoT |
+
+---
+
+## 2. Produtos verificados (por data de lançamento)
+
+Produtos de consumo, roteadores e dispositivos industriais testados com o PicoClaw.
+
+| Ano | Produto | Arq | SoC | RAM | Categoria |
+|-----|---------|-----|-----|-----|-----------|
+| 2009 | Nokia N900 | ARM (A8) | OMAP3430 | 256MB | Smartphone |
+| 2012 | Samsung Galaxy Note 10.1 (N8000) | ARM (A9) | Exynos 4412 | 2GB | Tablet |
+| 2016 | Xiaomi Router 3G (小米路由器3G) | MIPS | MT7620 | 256MB | Roteador (OpenWrt) |
+| 2018 | Phicomm N1 (斐讯N1) | ARM64 (A53) | S905D | 2GB | TV Box / Servidor doméstico |
+| 2019 | Xiaomi AI Speaker (小爱音箱) | ARM64 (A53) | — | 256MB | Alto-falante inteligente |
+| 2024 | [NanoKVM](https://wiki.sipeed.com/hardware/en/kvm/NanoKVM/introduction.html) | RISC-V | SG2002 | 256MB | IP-KVM |
+| 2025 | HaaS506-LD1 | RISC-V | D213 | 128MB | RTU industrial |
+| 2025 | [NanoKVM-Pro](https://wiki.sipeed.com/hardware/en/kvm/NanoKVM_Pro/introduction.html) | ARM64 (A53) | AX630C | 1GB | IP-KVM Pro |
+| 2026 | [MaixCAM2](https://wiki.sipeed.com/hardware/en/maixcam/index.html) | ARM64 (A53) | AX630C | 1/4GB | Câmera AI 4K |
+
+---
+
+## 3. Placas de desenvolvimento verificadas (por data de lançamento)
+
+| Ano | Placa | Arq | SoC | RAM | Link de compra |
+|-----|-------|-----|-----|-----|----------------|
+| 2012 | [Raspberry Pi 1 Model B](https://www.raspberrypi.com/products/) | ARMv6 | BCM2835 | 512MB | — |
+| 2015 | [Raspberry Pi 2 Model B](https://www.raspberrypi.com/products/raspberry-pi-2-model-b/) | ARMv7 (A7) | BCM2836 | 1GB | — |
+| 2015 | [Raspberry Pi Zero](https://www.raspberrypi.com/products/raspberry-pi-zero/) | ARMv6 | BCM2835 | 512MB | — |
+| 2016 | [Raspberry Pi 3 Model B](https://www.raspberrypi.com/products/raspberry-pi-3-model-b/) | ARM64 (A53) | BCM2837 | 1GB | — |
+| 2017 | [LicheePi Zero](https://wiki.sipeed.com/hardware/en/lichee/Zero/Zero.html) | ARMv7 (A7) | Allwinner V3s | 64MB | [Sipeed](https://sipeed.com/) |
+| 2019 | [Raspberry Pi 4 Model B](https://www.raspberrypi.com/products/raspberry-pi-4-model-b/) | ARM64 (A72) | BCM2711 | 1~8GB | [RPi](https://www.raspberrypi.com/) |
+| 2023 | [Raspberry Pi 5](https://www.raspberrypi.com/products/raspberry-pi-5/) | ARM64 (A76) | BCM2712 | 2~8GB | [RPi](https://www.raspberrypi.com/) |
+| 2024 | [LicheeRV-Nano](https://wiki.sipeed.com/hardware/en/lichee/RV_Nano/1_intro.html) | RISC-V | SG2002 | 256MB | [AliExpress](https://www.aliexpress.com/item/1005006519668532.html) |
+| 2024 | [MaixCAM-Pro](https://wiki.sipeed.com/hardware/en/maixcam/index.html) | RISC-V | SG2002 | 256MB | [Sipeed](https://sipeed.com/) |
+| 2024 | [Milk-V Duo 64M](https://milkv.io/docs/duo/getting-started/duo) | RISC-V | CV1800B | 64MB | [Milk-V](https://milkv.io/) |
+| 2024 | [CanMV-K230](https://developer.canaan-creative.com/k230_canmv/en/main/) | RISC-V | K230 | 512MB | [Canaan](https://www.canaan-creative.com/) |
+
+---
+
+## 4. Também funciona em
+
+### Celulares Android (via Termux)
+
+Qualquer celular Android ARM64 (2015+) com 1GB+ de RAM. Instale o [Termux](https://github.com/termux/termux-app), use `proot` para rodar o PicoClaw.
+
+> Veja [README: Rodar em celulares Android antigos](../../README.pt-br.md#-run-on-old-android-phones) para instruções de configuração.
+
+### Desktop / Servidor / Nuvem
+
+| Plataforma | Notas |
+|------------|-------|
+| x86_64 Linux | Binário nativo, sem dependências |
+| x86_64 Windows | Binário nativo |
+| macOS (Intel / Apple Silicon) | Binário nativo |
+| Docker (any platform) | `docker compose` em uma linha, veja [Guia Docker](docker.md) |
+| OpenWrt routers | Builds MIPS/ARM, requer >32MB de RAM livre |
+| FreeBSD / NetBSD | Builds x86_64 e arm64 disponíveis |
+
+---
+
+## 5. Requisitos mínimos
+
+| Recurso | Mínimo | Recomendado |
+|---------|--------|-------------|
+| RAM | 10MB livres | 32MB+ livres |
+| Armazenamento | 20MB (binário) | 50MB+ (com workspace) |
+| CPU | Qualquer (single-core 0,6GHz+) | — |
+| OS | Linux (kernel 3.x+) | Linux 5.x+ |
+| Rede | Necessária (para chamadas de API LLM) | Ethernet ou WiFi |
+
+---
+
+## 6. Como testar e contribuir
+
+```bash
+# 1. Baixar para sua arquitetura
+wget https://github.com/sipeed/picoclaw/releases/latest/download/picoclaw_Linux_arm64.tar.gz
+tar xzf picoclaw_Linux_arm64.tar.gz
+
+# 2. Inicializar
+./picoclaw onboard
+
+# 3. Testar
+./picoclaw agent -m "Hello, what board am I running on?"
+```
+
+Builds disponíveis: `linux-amd64`, `linux-arm64`, `linux-arm`, `linux-riscv64`, `linux-loong64`, `linux-mipsle`
+
+### Adicionar seu hardware
+
+1. Faça fork deste repositório
+2. Adicione seu chip / produto / placa na tabela apropriada
+3. Inclua: nome, arquitetura, SoC, RAM, ano e um link se disponível
+4. Envie um PR
+
+Fabricantes de hardware: deseja adicionar suporte oficial ou co-promover? Abra uma issue ou entre em contato via [Discord](https://discord.gg/V4sAZ9XWpN).
diff --git a/docs/pt-br/providers.md b/docs/pt-br/providers.md
new file mode 100644
index 000000000..0f7a4b5a1
--- /dev/null
+++ b/docs/pt-br/providers.md
@@ -0,0 +1,433 @@
+# 🔌 Provedores e Configuração de Modelos
+
+> Voltar ao [README](../../README.pt-br.md)
+
+### Provedores
+
+> [!NOTE]
+> O Groq fornece transcrição de voz gratuita via Whisper. Se configurado, mensagens de áudio de qualquer canal serão automaticamente transcritas no nível do agente.
+
+| Provider | Purpose | Get API Key |
+| ------------ | --------------------------------------- | ------------------------------------------------------------ |
+| `gemini` | LLM (Gemini direct) | [aistudio.google.com](https://aistudio.google.com) |
+| `zhipu` | LLM (Zhipu direct) | [bigmodel.cn](https://bigmodel.cn) |
+| `volcengine` | LLM(Volcengine direct) | [volcengine.com](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) |
+| `openrouter` | LLM (recommended, access to all models) | [openrouter.ai](https://openrouter.ai) |
+| `anthropic` | LLM (Claude direct) | [console.anthropic.com](https://console.anthropic.com) |
+| `openai` | LLM (GPT direct) | [platform.openai.com](https://platform.openai.com) |
+| `deepseek` | LLM (DeepSeek direct) | [platform.deepseek.com](https://platform.deepseek.com) |
+| `qwen` | LLM (Qwen direct) | [dashscope.console.aliyun.com](https://dashscope.console.aliyun.com) |
+| `groq` | LLM + **Voice transcription** (Whisper) | [console.groq.com](https://console.groq.com) |
+| `cerebras` | LLM (Cerebras direct) | [cerebras.ai](https://cerebras.ai) |
+| `vivgrid` | LLM (Vivgrid direct) | [vivgrid.com](https://vivgrid.com) |
+| `moonshot` | LLM (Kimi/Moonshot direct) | [platform.moonshot.cn](https://platform.moonshot.cn) |
+| `minimax` | LLM (Minimax direct) | [platform.minimaxi.com](https://platform.minimaxi.com) |
+| `avian` | LLM (Avian direct) | [avian.io](https://avian.io) |
+| `mistral` | LLM (Mistral direct) | [console.mistral.ai](https://console.mistral.ai) |
+| `longcat` | LLM (Longcat direct) | [longcat.ai](https://longcat.ai) |
+| `modelscope` | LLM (ModelScope direct) | [modelscope.cn](https://modelscope.cn) |
+
+### Configuração de Modelos (model_list)
+
+> **Novidade?** O PicoClaw agora usa uma abordagem de configuração **centrada no modelo**. Basta especificar o formato `vendor/model` (ex.: `zhipu/glm-4.7`) para adicionar novos provedores — **sem necessidade de alteração de código!**
+
+Este design também permite **suporte multi-agente** com seleção flexível de provedores:
+
+- **Agentes diferentes, provedores diferentes**: Cada agente pode usar seu próprio provedor LLM
+- **Fallback de modelos**: Configure modelos primários e de fallback para resiliência
+- **Balanceamento de carga**: Distribua requisições entre múltiplos endpoints
+- **Configuração centralizada**: Gerencie todos os provedores em um só lugar
+
+#### 📋 Todos os Vendors Suportados
+
+| Vendor | `model` Prefix | Default API Base | Protocol | API Key |
+| ------------------- | ----------------- |-----------------------------------------------------| --------- | ---------------------------------------------------------------- |
+| **OpenAI** | `openai/` | `https://api.openai.com/v1` | OpenAI | [Get Key](https://platform.openai.com) |
+| **Anthropic** | `anthropic/` | `https://api.anthropic.com/v1` | Anthropic | [Get Key](https://console.anthropic.com) |
+| **智谱 AI (GLM)** | `zhipu/` | `https://open.bigmodel.cn/api/paas/v4` | OpenAI | [Get Key](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) |
+| **DeepSeek** | `deepseek/` | `https://api.deepseek.com/v1` | OpenAI | [Get Key](https://platform.deepseek.com) |
+| **Google Gemini** | `gemini/` | `https://generativelanguage.googleapis.com/v1beta` | OpenAI | [Get Key](https://aistudio.google.com/api-keys) |
+| **Groq** | `groq/` | `https://api.groq.com/openai/v1` | OpenAI | [Get Key](https://console.groq.com) |
+| **Moonshot** | `moonshot/` | `https://api.moonshot.cn/v1` | OpenAI | [Get Key](https://platform.moonshot.cn) |
+| **通义千问 (Qwen)** | `qwen/` | `https://dashscope.aliyuncs.com/compatible-mode/v1` | OpenAI | [Get Key](https://dashscope.console.aliyun.com) |
+| **NVIDIA** | `nvidia/` | `https://integrate.api.nvidia.com/v1` | OpenAI | [Get Key](https://build.nvidia.com) |
+| **Ollama** | `ollama/` | `http://localhost:11434/v1` | OpenAI | Local (no key needed) |
+| **OpenRouter** | `openrouter/` | `https://openrouter.ai/api/v1` | OpenAI | [Get Key](https://openrouter.ai/keys) |
+| **LiteLLM Proxy** | `litellm/` | `http://localhost:4000/v1` | OpenAI | Your LiteLLM proxy key |
+| **VLLM** | `vllm/` | `http://localhost:8000/v1` | OpenAI | Local |
+| **Cerebras** | `cerebras/` | `https://api.cerebras.ai/v1` | OpenAI | [Get Key](https://cerebras.ai) |
+| **VolcEngine (Doubao)** | `volcengine/` | `https://ark.cn-beijing.volces.com/api/v3` | OpenAI | [Get Key](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) |
+| **神算云** | `shengsuanyun/` | `https://router.shengsuanyun.com/api/v1` | OpenAI | - |
+| **BytePlus** | `byteplus/` | `https://ark.ap-southeast.bytepluses.com/api/v3` | OpenAI | [Get Key](https://www.byteplus.com) |
+| **Vivgrid** | `vivgrid/` | `https://api.vivgrid.com/v1` | OpenAI | [Get Key](https://vivgrid.com) |
+| **LongCat** | `longcat/` | `https://api.longcat.chat/openai` | OpenAI | [Get Key](https://longcat.chat/platform) |
+| **ModelScope (魔搭)**| `modelscope/` | `https://api-inference.modelscope.cn/v1` | OpenAI | [Get Token](https://modelscope.cn/my/tokens) |
+| **Antigravity** | `antigravity/` | Google Cloud | Custom | OAuth only |
+| **GitHub Copilot** | `github-copilot/` | `localhost:4321` | gRPC | - |
+
+#### Configuração Básica
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "ark-code-latest",
+ "model": "volcengine/ark-code-latest",
+ "api_key": "sk-your-api-key"
+ },
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_key": "sk-your-openai-key"
+ },
+ {
+ "model_name": "claude-sonnet-4.6",
+ "model": "anthropic/claude-sonnet-4.6",
+ "api_key": "sk-ant-your-key"
+ },
+ {
+ "model_name": "glm-4.7",
+ "model": "zhipu/glm-4.7",
+ "api_key": "your-zhipu-key"
+ }
+ ],
+ "agents": {
+ "defaults": {
+ "model_name": "gpt-5.4"
+ }
+ }
+}
+```
+
+#### Exemplos por Vendor
+
+**OpenAI**
+
+```json
+{
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_key": "sk-..."
+}
+```
+
+**VolcEngine (Doubao)**
+
+```json
+{
+ "model_name": "ark-code-latest",
+ "model": "volcengine/ark-code-latest",
+ "api_key": "sk-..."
+}
+```
+
+**智谱 AI (GLM)**
+
+```json
+{
+ "model_name": "glm-4.7",
+ "model": "zhipu/glm-4.7",
+ "api_key": "your-key"
+}
+```
+
+**DeepSeek**
+
+```json
+{
+ "model_name": "deepseek-chat",
+ "model": "deepseek/deepseek-chat",
+ "api_key": "sk-..."
+}
+```
+
+**Anthropic (com chave de API)**
+
+```json
+{
+ "model_name": "claude-sonnet-4.6",
+ "model": "anthropic/claude-sonnet-4.6",
+ "api_key": "sk-ant-your-key"
+}
+```
+
+> Execute `picoclaw auth login --provider anthropic` para colar seu token de API.
+
+**Anthropic Messages API (formato nativo)**
+
+Para acesso direto à API Anthropic ou endpoints personalizados que suportam apenas o formato de mensagem nativo da Anthropic:
+
+```json
+{
+ "model_name": "claude-opus-4-6",
+ "model": "anthropic-messages/claude-opus-4-6",
+ "api_key": "sk-ant-your-key",
+ "api_base": "https://api.anthropic.com"
+}
+```
+
+> Use o protocolo `anthropic-messages` quando:
+> - Usar proxies de terceiros que suportam apenas o endpoint nativo `/v1/messages` da Anthropic (não o compatível com OpenAI `/v1/chat/completions`)
+> - Conectar a serviços como MiniMax, Synthetic que requerem o formato de mensagem nativo da Anthropic
+> - O protocolo `anthropic` existente retorna erros 404 (indicando que o endpoint não suporta formato compatível com OpenAI)
+>
+> **Nota:** O protocolo `anthropic` usa formato compatível com OpenAI (`/v1/chat/completions`), enquanto `anthropic-messages` usa o formato nativo da Anthropic (`/v1/messages`). Escolha com base no formato suportado pelo seu endpoint.
+
+**Ollama (local)**
+
+```json
+{
+ "model_name": "llama3",
+ "model": "ollama/llama3"
+}
+```
+
+**Proxy/API Personalizado**
+
+```json
+{
+ "model_name": "my-custom-model",
+ "model": "openai/custom-model",
+ "api_base": "https://my-proxy.com/v1",
+ "api_key": "sk-...",
+ "request_timeout": 300
+}
+```
+
+**LiteLLM Proxy**
+
+```json
+{
+ "model_name": "lite-gpt4",
+ "model": "litellm/lite-gpt4",
+ "api_base": "http://localhost:4000/v1",
+ "api_key": "sk-..."
+}
+```
+
+O PicoClaw remove apenas o prefixo externo `litellm/` antes de enviar a requisição, então aliases de proxy como `litellm/lite-gpt4` enviam `lite-gpt4`, enquanto `litellm/openai/gpt-4o` envia `openai/gpt-4o`.
+
+#### Balanceamento de Carga
+
+Configure múltiplos endpoints para o mesmo nome de modelo — o PicoClaw fará automaticamente round-robin entre eles:
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_base": "https://api1.example.com/v1",
+ "api_key": "sk-key1"
+ },
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_base": "https://api2.example.com/v1",
+ "api_key": "sk-key2"
+ }
+ ]
+}
+```
+
+#### Migração da Configuração Legacy `providers`
+
+A configuração antiga `providers` está **descontinuada** mas ainda é suportada para compatibilidade retroativa.
+
+**Configuração Antiga (descontinuada):**
+
+```json
+{
+ "providers": {
+ "zhipu": {
+ "api_key": "your-key",
+ "api_base": "https://open.bigmodel.cn/api/paas/v4"
+ }
+ },
+ "agents": {
+ "defaults": {
+ "provider": "zhipu",
+ "model": "glm-4.7"
+ }
+ }
+}
+```
+
+**Configuração Nova (recomendada):**
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "glm-4.7",
+ "model": "zhipu/glm-4.7",
+ "api_key": "your-key"
+ }
+ ],
+ "agents": {
+ "defaults": {
+ "model_name": "glm-4.7"
+ }
+ }
+}
+```
+
+Para guia de migração detalhado, veja [migration/model-list-migration.md](../migration/model-list-migration.md).
+
+### Arquitetura de Provedores
+
+O PicoClaw roteia provedores por família de protocolo:
+
+- Protocolo compatível com OpenAI: OpenRouter, gateways compatíveis com OpenAI, Groq, Zhipu e endpoints estilo vLLM.
+- Protocolo Anthropic: Comportamento nativo da API Claude.
+- Caminho Codex/OAuth: Rota de autenticação OAuth/token da OpenAI.
+
+Isso mantém o runtime leve enquanto torna novos backends compatíveis com OpenAI basicamente uma operação de configuração (`api_base` + `api_key`).
+
+
+Zhipu
+
+**1. Obter chave de API e URL base**
+
+* Obtenha a [chave de API](https://bigmodel.cn/usercenter/proj-mgmt/apikeys)
+
+**2. Configurar**
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "workspace": "~/.picoclaw/workspace",
+ "model_name": "glm-4.7",
+ "max_tokens": 8192,
+ "temperature": 0.7,
+ "max_tool_iterations": 20
+ }
+ },
+ "providers": {
+ "zhipu": {
+ "api_key": "Your API Key",
+ "api_base": "https://open.bigmodel.cn/api/paas/v4"
+ }
+ }
+}
+```
+
+**3. Executar**
+
+```bash
+picoclaw agent -m "Hello"
+```
+
+
+
+
+Exemplo de configuração completa
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "model_name": "anthropic/claude-opus-4-5"
+ }
+ },
+ "session": {
+ "dm_scope": "per-channel-peer"
+ },
+ "providers": {
+ "openrouter": {
+ "api_key": "sk-or-v1-xxx"
+ },
+ "groq": {
+ "api_key": "gsk_xxx"
+ }
+ },
+ "channels": {
+ "telegram": {
+ "enabled": true,
+ "token": "123456:ABC...",
+ "allow_from": ["123456789"]
+ },
+ "discord": {
+ "enabled": true,
+ "token": "",
+ "allow_from": [""]
+ },
+ "whatsapp": {
+ "enabled": false,
+ "bridge_url": "ws://localhost:3001",
+ "use_native": false,
+ "session_store_path": "",
+ "allow_from": []
+ },
+ "feishu": {
+ "enabled": false,
+ "app_id": "cli_xxx",
+ "app_secret": "xxx",
+ "encrypt_key": "",
+ "verification_token": "",
+ "allow_from": []
+ },
+ "qq": {
+ "enabled": false,
+ "app_id": "",
+ "app_secret": "",
+ "allow_from": []
+ }
+ },
+ "tools": {
+ "web": {
+ "brave": {
+ "enabled": false,
+ "api_key": "BSA...",
+ "max_results": 5
+ },
+ "duckduckgo": {
+ "enabled": true,
+ "max_results": 5
+ },
+ "perplexity": {
+ "enabled": false,
+ "api_key": "",
+ "max_results": 5
+ },
+ "searxng": {
+ "enabled": false,
+ "base_url": "http://localhost:8888",
+ "max_results": 5
+ }
+ },
+ "cron": {
+ "exec_timeout_minutes": 5
+ }
+ },
+ "heartbeat": {
+ "enabled": true,
+ "interval": 30
+ }
+}
+```
+
+
+
+---
+
+## 📝 Comparação de Chaves de API
+
+| Service | Pricing | Use Case |
+| ---------------- | ------------------------ | ------------------------------------- |
+| **OpenRouter** | Free: 200K tokens/month | Multiple models (Claude, GPT-4, etc.) |
+| **Volcengine CodingPlan** | ¥9.9/first month | Best for Chinese users, multiple SOTA models (Doubao, DeepSeek, etc.) |
+| **Zhipu** | Free: 200K tokens/month | Suitable for Chinese users |
+| **Brave Search** | $5/1000 queries | Web search functionality |
+| **SearXNG** | Free (self-hosted) | Privacy-focused metasearch (70+ engines) |
+| **Groq** | Free tier available | Fast inference (Llama, Mixtral) |
+| **Cerebras** | Free tier available | Fast inference (Llama, Qwen, etc.) |
+| **LongCat** | Free: up to 5M tokens/day | Fast inference |
+| **ModelScope** | Free: 2000 requests/day | Inference (Qwen, GLM, DeepSeek, etc.) |
+
+---
+
+
+
+
diff --git a/docs/pt-br/spawn-tasks.md b/docs/pt-br/spawn-tasks.md
new file mode 100644
index 000000000..d6b539cb1
--- /dev/null
+++ b/docs/pt-br/spawn-tasks.md
@@ -0,0 +1,61 @@
+# 🔄 Tarefas Assíncronas e Spawn
+
+> Voltar ao [README](../../README.pt-br.md)
+
+## Tarefas Rápidas (resposta direta)
+
+- Informar a hora atual
+
+## Tarefas Longas (usar spawn para assíncrono)
+
+- Pesquisar na web notícias sobre IA e resumir
+- Verificar e-mail e relatar mensagens importantes
+```
+
+**Comportamentos principais:**
+
+| Feature | Description |
+| ----------------------- | --------------------------------------------------------- |
+| **spawn** | Creates async subagent, doesn't block heartbeat |
+| **Independent context** | Subagent has its own context, no session history |
+| **message tool** | Subagent communicates with user directly via message tool |
+| **Non-blocking** | After spawning, heartbeat continues to next task |
+
+#### Como Funciona a Comunicação do Subagente
+
+```
+Heartbeat é acionado
+ ↓
+Agente lê HEARTBEAT.md
+ ↓
+Para tarefa longa: spawn subagente
+ ↓ ↓
+Continua para próxima tarefa Subagente trabalha independentemente
+ ↓ ↓
+Todas as tarefas concluídas Subagente usa ferramenta "message"
+ ↓ ↓
+Responde HEARTBEAT_OK Usuário recebe resultado diretamente
+```
+
+O subagente tem acesso a ferramentas (message, web_search, etc.) e pode se comunicar com o usuário independentemente sem passar pelo agente principal.
+
+**Configuração:**
+
+```json
+{
+ "heartbeat": {
+ "enabled": true,
+ "interval": 30
+ }
+}
+```
+
+| Option | Default | Description |
+| ---------- | ------- | ---------------------------------- |
+| `enabled` | `true` | Enable/disable heartbeat |
+| `interval` | `30` | Check interval in minutes (min: 5) |
+
+**Variáveis de ambiente:**
+
+* `PICOCLAW_HEARTBEAT_ENABLED=false` para desabilitar
+* `PICOCLAW_HEARTBEAT_INTERVAL=60` para alterar o intervalo
diff --git a/docs/pt-br/tools_configuration.md b/docs/pt-br/tools_configuration.md
new file mode 100644
index 000000000..2cc4f3999
--- /dev/null
+++ b/docs/pt-br/tools_configuration.md
@@ -0,0 +1,360 @@
+# 🔧 Configuração de Ferramentas
+
+> Voltar ao [README](../../README.pt-br.md)
+
+A configuração de ferramentas do PicoClaw está localizada no campo `tools` do `config.json`.
+
+## Estrutura de diretórios
+
+```json
+{
+ "tools": {
+ "web": {
+ ...
+ },
+ "mcp": {
+ ...
+ },
+ "exec": {
+ ...
+ },
+ "cron": {
+ ...
+ },
+ "skills": {
+ ...
+ }
+ }
+}
+```
+
+## Ferramentas Web
+
+As ferramentas web são usadas para pesquisa e busca de páginas web.
+
+### Web Fetcher
+Configurações gerais para busca e processamento de conteúdo de páginas web.
+
+| Config | Tipo | Padrão | Descrição |
+|---------------------|--------|---------------|-----------------------------------------------------------------------------------------------|
+| `enabled` | bool | true | Habilitar a capacidade de busca de páginas web. |
+| `fetch_limit_bytes` | int | 10485760 | Tamanho máximo do payload da página web a ser buscado, em bytes (padrão é 10MB). |
+| `format` | string | "plaintext" | Formato de saída do conteúdo buscado. Opções: `plaintext` ou `markdown` (recomendado). |
+
+### Brave
+
+| Config | Tipo | Padrão | Descrição |
+|---------------|--------|--------|----------------------------|
+| `enabled` | bool | false | Habilitar pesquisa Brave |
+| `api_key` | string | - | Chave API do Brave Search |
+| `max_results` | int | 5 | Número máximo de resultados |
+
+### DuckDuckGo
+
+| Config | Tipo | Padrão | Descrição |
+|---------------|------|--------|--------------------------------|
+| `enabled` | bool | true | Habilitar pesquisa DuckDuckGo |
+| `max_results` | int | 5 | Número máximo de resultados |
+
+### Perplexity
+
+| Config | Tipo | Padrão | Descrição |
+|---------------|--------|--------|--------------------------------|
+| `enabled` | bool | false | Habilitar pesquisa Perplexity |
+| `api_key` | string | - | Chave API do Perplexity |
+| `max_results` | int | 5 | Número máximo de resultados |
+
+## Ferramenta Exec
+
+A ferramenta exec é usada para executar comandos shell.
+
+| Config | Tipo | Padrão | Descrição |
+|------------------------|-------|--------|-------------------------------------------------|
+| `enabled` | bool | true | Habilitar a ferramenta exec |
+| `enable_deny_patterns` | bool | true | Habilitar bloqueio padrão de comandos perigosos |
+| `custom_deny_patterns` | array | [] | Padrões de negação personalizados (expressões regulares) |
+
+### Desabilitando a Ferramenta Exec
+
+Para desabilitar completamente a ferramenta `exec`, defina `enabled` como `false`:
+
+**Via arquivo de configuração:**
+```json
+{
+ "tools": {
+ "exec": {
+ "enabled": false
+ }
+ }
+}
+```
+
+**Via variável de ambiente:**
+```bash
+PICOCLAW_TOOLS_EXEC_ENABLED=false
+```
+
+> **Nota:** Quando desabilitada, o agent não poderá executar comandos shell. Isso também afeta a capacidade da ferramenta Cron de executar comandos shell agendados.
+
+### Funcionalidade
+
+- **`enable_deny_patterns`**: Defina como `false` para desabilitar completamente os padrões de bloqueio de comandos perigosos padrão
+- **`custom_deny_patterns`**: Adicione padrões regex de negação personalizados; comandos correspondentes serão bloqueados
+
+### Padrões de comandos bloqueados por padrão
+
+Por padrão, o PicoClaw bloqueia os seguintes comandos perigosos:
+
+- Comandos de exclusão: `rm -rf`, `del /f/q`, `rmdir /s`
+- Operações de disco: `format`, `mkfs`, `diskpart`, `dd if=`, escrita em `/dev/sd*`
+- Operações do sistema: `shutdown`, `reboot`, `poweroff`
+- Substituição de comandos: `$()`, `${}`, crases
+- Pipe para shell: `| sh`, `| bash`
+- Escalação de privilégios: `sudo`, `chmod`, `chown`
+- Controle de processos: `pkill`, `killall`, `kill -9`
+- Operações remotas: `curl | sh`, `wget | sh`, `ssh`
+- Gerenciamento de pacotes: `apt`, `yum`, `dnf`, `npm install -g`, `pip install --user`
+- Contêineres: `docker run`, `docker exec`
+- Git: `git push`, `git force`
+- Outros: `eval`, `source *.sh`
+
+### Limitação arquitetural conhecida
+
+O guarda exec apenas valida o comando de nível superior enviado ao PicoClaw. Ele **não** inspeciona recursivamente processos filhos gerados por ferramentas de build ou scripts após o início desse comando.
+
+Exemplos de fluxos de trabalho que podem contornar o guarda de comando direto uma vez que o comando inicial é permitido:
+
+- `make run`
+- `go run ./cmd/...`
+- `cargo run`
+- `npm run build`
+
+Isso significa que o guarda é útil para bloquear comandos diretos obviamente perigosos, mas **não** é um sandbox completo para pipelines de build não revisados. Se seu modelo de ameaça inclui código não confiável no workspace, use isolamento mais forte, como contêineres, VMs ou um fluxo de aprovação em torno de comandos de build e execução.
+
+### Exemplo de configuração
+
+```json
+{
+ "tools": {
+ "exec": {
+ "enable_deny_patterns": true,
+ "custom_deny_patterns": [
+ "\\brm\\s+-r\\b",
+ "\\bkillall\\s+python"
+ ]
+ }
+ }
+}
+```
+
+## Ferramenta Cron
+
+A ferramenta cron é usada para agendar tarefas periódicas.
+
+| Config | Tipo | Padrão | Descrição |
+|------------------------|------|--------|-----------------------------------------------------|
+| `exec_timeout_minutes` | int | 5 | Tempo limite de execução em minutos, 0 significa sem limite |
+
+## Ferramenta MCP
+
+A ferramenta MCP permite a integração com servidores Model Context Protocol externos.
+
+### Descoberta de ferramentas (carregamento preguiçoso)
+
+Ao conectar a vários servidores MCP, expor centenas de ferramentas simultaneamente pode esgotar a janela de contexto do LLM e aumentar os custos de API. O recurso **Discovery** resolve isso mantendo as ferramentas MCP *ocultas* por padrão.
+
+Em vez de carregar todas as ferramentas, o LLM recebe uma ferramenta de pesquisa leve (usando correspondência de palavras-chave BM25 ou Regex). Quando o LLM precisa de uma capacidade específica, ele pesquisa a biblioteca oculta. As ferramentas correspondentes são então temporariamente "desbloqueadas" e injetadas no contexto por um número configurado de turnos (`ttl`).
+
+### Configuração global
+
+| Config | Tipo | Padrão | Descrição |
+|-------------|--------|--------|----------------------------------------------|
+| `enabled` | bool | false | Habilitar integração MCP globalmente |
+| `discovery` | object | `{}` | Configuração de descoberta de ferramentas (veja abaixo) |
+| `servers` | object | `{}` | Mapa de nome do servidor para configuração do servidor |
+
+### Configuração Discovery (`discovery`)
+
+| Config | Tipo | Padrão | Descrição |
+|----------------------|------|--------|-----------------------------------------------------------------------------------------------------------------------------------|
+| `enabled` | bool | false | Se true, as ferramentas MCP ficam ocultas e são carregadas sob demanda via pesquisa. Se false, todas as ferramentas são carregadas |
+| `ttl` | int | 5 | Número de turnos de conversa que uma ferramenta descoberta permanece desbloqueada |
+| `max_search_results` | int | 5 | Número máximo de ferramentas retornadas por consulta de pesquisa |
+| `use_bm25` | bool | true | Habilitar a ferramenta de pesquisa por linguagem natural/palavras-chave (`tool_search_tool_bm25`). **Aviso**: consome mais recursos que a pesquisa regex |
+| `use_regex` | bool | false | Habilitar a ferramenta de pesquisa por padrão regex (`tool_search_tool_regex`) |
+
+> **Nota:** Se `discovery.enabled` for `true`, você **deve** habilitar pelo menos um mecanismo de pesquisa (`use_bm25` ou `use_regex`),
+> caso contrário a aplicação falhará ao iniciar.
+
+### Configuração por servidor
+
+| Config | Tipo | Obrigatório | Descrição |
+|------------|--------|-------------|--------------------------------------------|
+| `enabled` | bool | sim | Habilitar este servidor MCP |
+| `type` | string | não | Tipo de transporte: `stdio`, `sse`, `http` |
+| `command` | string | stdio | Comando executável para transporte stdio |
+| `args` | array | não | Argumentos do comando para transporte stdio |
+| `env` | object | não | Variáveis de ambiente para processo stdio |
+| `env_file` | string | não | Caminho para arquivo de ambiente para processo stdio |
+| `url` | string | sse/http | URL do endpoint para transporte `sse`/`http` |
+| `headers` | object | não | Cabeçalhos HTTP para transporte `sse`/`http` |
+
+### Comportamento do transporte
+
+- Se `type` for omitido, o transporte é detectado automaticamente:
+ - `url` está definido → `sse`
+ - `command` está definido → `stdio`
+- `http` e `sse` ambos usam `url` + `headers` opcionais.
+- `env` e `env_file` são aplicados apenas a servidores `stdio`.
+
+### Exemplos de configuração
+
+#### 1) Servidor MCP Stdio
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "servers": {
+ "filesystem": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-filesystem",
+ "/tmp"
+ ]
+ }
+ }
+ }
+ }
+}
+```
+
+#### 2) Servidor MCP remoto SSE/HTTP
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "servers": {
+ "remote-mcp": {
+ "enabled": true,
+ "type": "sse",
+ "url": "https://example.com/mcp",
+ "headers": {
+ "Authorization": "Bearer YOUR_TOKEN"
+ }
+ }
+ }
+ }
+ }
+}
+```
+
+#### 3) Configuração MCP massiva com descoberta de ferramentas habilitada
+
+*Neste exemplo, o LLM verá apenas o `tool_search_tool_bm25`. Ele pesquisará e desbloqueará ferramentas do Github ou Postgres dinamicamente apenas quando solicitado pelo usuário.*
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "discovery": {
+ "enabled": true,
+ "ttl": 5,
+ "max_search_results": 5,
+ "use_bm25": true,
+ "use_regex": false
+ },
+ "servers": {
+ "github": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-github"
+ ],
+ "env": {
+ "GITHUB_PERSONAL_ACCESS_TOKEN": "YOUR_GITHUB_TOKEN"
+ }
+ },
+ "postgres": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-postgres",
+ "postgresql://user:password@localhost/dbname"
+ ]
+ },
+ "slack": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-slack"
+ ],
+ "env": {
+ "SLACK_BOT_TOKEN": "YOUR_SLACK_BOT_TOKEN",
+ "SLACK_TEAM_ID": "YOUR_SLACK_TEAM_ID"
+ }
+ }
+ }
+ }
+ }
+}
+```
+
+## Ferramenta Skills
+
+A ferramenta skills configura a descoberta e instalação de habilidades via registros como o ClawHub.
+
+### Registros
+
+| Config | Tipo | Padrão | Descrição |
+|------------------------------------|--------|-----------------------|----------------------------------------------|
+| `registries.clawhub.enabled` | bool | true | Habilitar registro ClawHub |
+| `registries.clawhub.base_url` | string | `https://clawhub.ai` | URL base do ClawHub |
+| `registries.clawhub.auth_token` | string | `""` | Token Bearer opcional para limites de taxa mais altos |
+| `registries.clawhub.search_path` | string | `/api/v1/search` | Caminho da API de pesquisa |
+| `registries.clawhub.skills_path` | string | `/api/v1/skills` | Caminho da API de Skills |
+| `registries.clawhub.download_path` | string | `/api/v1/download` | Caminho da API de download |
+
+### Exemplo de configuração
+
+```json
+{
+ "tools": {
+ "skills": {
+ "registries": {
+ "clawhub": {
+ "enabled": true,
+ "base_url": "https://clawhub.ai",
+ "auth_token": "",
+ "search_path": "/api/v1/search",
+ "skills_path": "/api/v1/skills",
+ "download_path": "/api/v1/download"
+ }
+ }
+ }
+ }
+}
+```
+
+## Variáveis de ambiente
+
+Todas as opções de configuração podem ser substituídas via variáveis de ambiente com o formato `PICOCLAW_TOOLS__`:
+
+Por exemplo:
+
+- `PICOCLAW_TOOLS_WEB_BRAVE_ENABLED=true`
+- `PICOCLAW_TOOLS_EXEC_ENABLED=false`
+- `PICOCLAW_TOOLS_EXEC_ENABLE_DENY_PATTERNS=false`
+- `PICOCLAW_TOOLS_CRON_EXEC_TIMEOUT_MINUTES=10`
+- `PICOCLAW_TOOLS_MCP_ENABLED=true`
+
+Nota: Configuração de tipo mapa aninhado (por exemplo `tools.mcp.servers..*`) é configurada no `config.json` em vez de variáveis de ambiente.
diff --git a/docs/pt-br/troubleshooting.md b/docs/pt-br/troubleshooting.md
new file mode 100644
index 000000000..286ad2ac8
--- /dev/null
+++ b/docs/pt-br/troubleshooting.md
@@ -0,0 +1,45 @@
+# 🐛 Solução de Problemas
+
+> Voltar ao [README](../../README.pt-br.md)
+
+## "model ... not found in model_list" ou OpenRouter "free is not a valid model ID"
+
+**Sintoma:** Você vê um dos seguintes erros:
+
+- `Error creating provider: model "openrouter/free" not found in model_list`
+- OpenRouter retorna 400: `"free is not a valid model ID"`
+
+**Causa:** O campo `model` na sua entrada `model_list` é o que é enviado para a API. Para o OpenRouter, você deve usar o ID de modelo **completo**, não uma abreviação.
+
+- **Errado:** `"model": "free"` → OpenRouter recebe `free` e rejeita.
+- **Correto:** `"model": "openrouter/free"` → OpenRouter recebe `openrouter/free` (roteamento automático do nível gratuito).
+
+**Correção:** Em `~/.picoclaw/config.json` (ou seu caminho de configuração):
+
+1. **agents.defaults.model_name** deve corresponder a um `model_name` em `model_list` (ex.: `"openrouter-free"`).
+2. O **model** dessa entrada deve ser um ID de modelo OpenRouter válido, por exemplo:
+ - `"openrouter/free"` – nível gratuito automático
+ - `"google/gemini-2.0-flash-exp:free"`
+ - `"meta-llama/llama-3.1-8b-instruct:free"`
+
+Exemplo:
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "model_name": "openrouter-free"
+ }
+ },
+ "model_list": [
+ {
+ "model_name": "openrouter-free",
+ "model": "openrouter/free",
+ "api_key": "sk-or-v1-YOUR_OPENROUTER_KEY",
+ "api_base": "https://openrouter.ai/api/v1"
+ }
+ ]
+}
+```
+
+Obtenha sua chave em [OpenRouter Keys](https://openrouter.ai/keys).
diff --git a/docs/spawn-tasks.md b/docs/spawn-tasks.md
new file mode 100644
index 000000000..05a5215d2
--- /dev/null
+++ b/docs/spawn-tasks.md
@@ -0,0 +1,70 @@
+# 🔄 Spawn & Async Tasks
+
+> Back to [README](../README.md)
+
+PicoClaw supports **asynchronous task execution** via the `spawn` tool. This is primarily used by the **Heartbeat** system to run long-running tasks without blocking the main agent loop.
+
+## Heartbeat
+
+The heartbeat system periodically checks `workspace/HEARTBEAT.md` for scheduled tasks. On first run, a default template is auto-generated. You can customize it to define quick tasks (handled inline) and long tasks (delegated via `spawn`).
+
+**Example `HEARTBEAT.md`:**
+
+```markdown
+## Quick Tasks (respond directly)
+
+- Report current time
+
+## Long Tasks (use spawn for async)
+
+- Search the web for AI news and summarize
+- Check email and report important messages
+```
+
+**Key behaviors:**
+
+| Feature | Description |
+| ----------------------- | --------------------------------------------------------- |
+| **spawn** | Creates async subagent, doesn't block heartbeat |
+| **Independent context** | Subagent has its own context, no session history |
+| **message tool** | Subagent communicates with user directly via message tool |
+| **Non-blocking** | After spawning, heartbeat continues to next task |
+
+#### How Subagent Communication Works
+
+```
+Heartbeat triggers
+ ↓
+Agent reads HEARTBEAT.md
+ ↓
+For long task: spawn subagent
+ ↓ ↓
+Continue to next task Subagent works independently
+ ↓ ↓
+All tasks done Subagent uses "message" tool
+ ↓ ↓
+Respond HEARTBEAT_OK User receives result directly
+```
+
+The subagent has access to tools (message, web_search, etc.) and can communicate with the user independently without going through the main agent.
+
+**Configuration:**
+
+```json
+{
+ "heartbeat": {
+ "enabled": true,
+ "interval": 30
+ }
+}
+```
+
+| Option | Default | Description |
+| ---------- | ------- | ---------------------------------- |
+| `enabled` | `true` | Enable/disable heartbeat |
+| `interval` | `30` | Check interval in minutes (min: 5) |
+
+**Environment variables:**
+
+* `PICOCLAW_HEARTBEAT_ENABLED=false` to disable
+* `PICOCLAW_HEARTBEAT_INTERVAL=60` to change interval
diff --git a/docs/subturn.md b/docs/subturn.md
new file mode 100644
index 000000000..b84c06627
--- /dev/null
+++ b/docs/subturn.md
@@ -0,0 +1,279 @@
+# 🔄 SubTurn Mechanism
+
+> Back to [README](../README.md)
+
+## Overview
+
+The `SubTurn` mechanism is a core feature in PicoClaw that allows tools to spawn isolated, nested agent loops to handle complex sub-tasks.
+
+By using a SubTurn, an agent can break down a problem and run a separate LLM invocation in an independent, ephemeral session. This ensures that intermediate reasoning, background tasks, or sub-agent outputs do not pollute the main conversation history.
+
+## Core Capabilities
+
+- **Context Isolation**: Each SubTurn uses an `ephemeralSessionStore`. Its message history does not leak into the parent task and is destroyed upon completion. The ephemeral session holds at most **50 messages**; older messages are automatically truncated when this limit is reached.
+- **Depth & Concurrency Limits**: Prevents infinite loops and resource exhaustion.
+ - **Maximum Depth**: Up to 3 nested levels.
+ - **Maximum Concurrency**: Up to 5 concurrent sub-turns per parent turn (managed via a semaphore with a 30-second timeout).
+- **Context Protection**: Supports soft context limits (`MaxContextRunes`). It proactively truncates old messages (while preserving system prompts and recent context) before hitting the provider's hard context window limit.
+- **Error Recovery**: Automatically detects and recovers from provider context length exceeded errors and truncation errors by compressing history and retrying.
+
+## Configuration (`SubTurnConfig`)
+
+When spawning a SubTurn, you must provide a `SubTurnConfig`:
+
+| Field | Type | Description |
+| :--- | :--- | :--- |
+| `Model` | `string` | The LLM model to use for the sub-turn (e.g., `gpt-4o-mini`). **Required.** |
+| `Tools` | `[]tools.Tool` | Tools granted to the sub-turn. If empty, it inherits the parent's tools. |
+| `SystemPrompt` | `string` | The task description for the sub-turn. Sent as the first user message to the LLM (not as a system prompt override). |
+| `ActualSystemPrompt` | `string` | Optional explicit system prompt to replace the agent's default. Leave empty to inherit the parent agent's system prompt. |
+| `MaxTokens` | `int` | Maximum tokens for the generated response. |
+| `Async` | `bool` | Controls the result delivery mode (Synchronous vs. Asynchronous). |
+| `Critical` | `bool` | If `true`, the sub-turn continues running even if the parent finishes gracefully. |
+| `Timeout` | `time.Duration` | Maximum execution time (default: 5 minutes). |
+| `MaxContextRunes`| `int` | Soft context limit. `0` = auto-calculate (75% of model's context window, recommended), `-1` = no limit (disable soft truncation, rely only on hard context error recovery), `>0` = use specified rune limit. |
+
+> **Note:** The `Async` flag does **not** make the call non-blocking. It only controls whether the result is also delivered to the parent's `pendingResults` channel. Both modes block the caller until the sub-turn completes. For true non-blocking execution, the caller must spawn the sub-turn in a separate goroutine.
+
+## Execution Modes
+
+### Synchronous (`Async: false`)
+
+This is the standard mode where the caller needs the result immediately to proceed.
+
+- The caller blocks until the sub-turn completes.
+- The result is **only** returned directly via the function return value.
+- It is **not** delivered to the parent's pending results channel.
+
+**Example:**
+```go
+cfg := agent.SubTurnConfig{
+ Model: "gpt-4o-mini",
+ SystemPrompt: "Analyze the provided codebase...",
+ Async: false,
+}
+result, err := agent.SpawnSubTurn(ctx, cfg)
+// Process result immediately
+```
+
+### Asynchronous (`Async: true`)
+
+Used for "fire-and-forget" operations or parallel processing where the parent turn collects results later.
+
+- The result is delivered to the parent turn's `pendingResults` channel.
+- The result is **also** returned via the function return value (for consistency).
+- The parent's Agent Loop will poll this channel in subsequent iterations and automatically inject the results into the ongoing conversation context as `[SubTurn Result]`.
+
+**Example:**
+```go
+cfg := agent.SubTurnConfig{
+ Model: "gpt-4o-mini",
+ SystemPrompt: "Run a background security scan...",
+ Async: true,
+}
+result, err := agent.SpawnSubTurn(ctx, cfg)
+// The result will also be injected into the parent loop later via channel
+```
+
+## Error Recovery and Retries
+
+SubTurns implement automatic retry mechanisms for transient errors:
+
+| Error Type | Max Retries | Recovery Action |
+|:-----------|:------------|:----------------|
+| Context Length Exceeded | 2 | Force compress history and retry |
+| Response Truncated (`finish_reason="truncated"`) | 2 | Inject recovery prompt and retry |
+
+### Truncation Recovery
+When the LLM response is truncated (`finish_reason="truncated"`), SubTurn automatically:
+1. Detects the truncation from `turnState.lastFinishReason`
+2. Injects a recovery prompt: "Your previous response was truncated due to length. Please provide a shorter, complete response..."
+3. Retries up to 2 times
+
+### Context Error Recovery
+When the provider returns a context length error (e.g., `context_length_exceeded`):
+1. Force compresses the message history (drops oldest 50% of conversation)
+2. Retries with the compressed context
+3. Up to 2 retries before failing
+
+## Lifecycle and Cancellation
+
+SubTurns operate within an independent context but maintain a structural link to their parent `turnState`.
+
+### Graceful Parent Finish
+When the parent task finishes naturally (`Finish(false)`):
+- **Non-critical** sub-turns receive a signal to exit gracefully without throwing an error.
+- **Critical** (`Critical: true`) sub-turns continue running in the background. Once finished, their results are emitted as **Orphan Results** so the data is not lost.
+
+### Hard Abort
+When the parent task is forcefully aborted (e.g., user interrupts with `/stop`):
+- A cascading cancellation is triggered, instantly terminating all child and grandchild sub-turns.
+- The root turn's session history rolls back to the snapshot taken at turn start (`initialHistoryLength`), preventing dirty context. SubTurns are not affected by this rollback as they use ephemeral sessions that are discarded anyway.
+
+## Agent Loop Integration
+
+### Bus Draining During Processing
+
+When a message enters the `Run()` loop, the agent starts a `drainBusToSteering` goroutine before calling `processMessage`. This goroutine runs concurrently with the entire processing lifecycle and continuously consumes any new inbound messages from the bus, redirecting them into the **steering queue** instead of dropping them.
+
+This ensures that if a user sends a follow-up message while the agent is processing (including during SubTurn execution), the message is not lost — it will be picked up between tool call iterations via `dequeueSteeringMessages`.
+
+The drain goroutine stops automatically when `processMessage` returns (via a cancellable context).
+
+### Pending Result Polling
+
+The agent loop polls for async SubTurn results at two points per iteration:
+1. **Before the LLM call**: injects any arrived results as `[SubTurn Result]` messages into the conversation context.
+2. **After all tool executions**: polls again during the tool loop to catch results that arrived during tool execution.
+3. **After the final iteration**: one last poll before the turn ends to avoid losing late-arriving results.
+
+### Turn State Tracking
+
+All active root turns are registered in `AgentLoop.activeTurnStates` (`sync.Map`, keyed by session key). This allows `HardAbort` and `/subagents` observability commands to find and operate on active turns.
+
+## Event Bus Integration
+
+SubTurns emit specific events to the PicoClaw `EventBus` for observability and debugging:
+
+| Event Kind | When Emitted | Payload |
+|:------|:-------------|:--------|
+| `subturn_spawn` | Sub-turn successfully initialized | `SubTurnSpawnPayload{AgentID, Label, ParentTurnID}` |
+| `subturn_end` | Sub-turn finishes (success or error) | `SubTurnEndPayload{AgentID, Status}` |
+| `subturn_result_delivered` | Async result successfully delivered to parent | `SubTurnResultDeliveredPayload{TargetChannel, TargetChatID, ContentLen}` |
+| `subturn_orphan` | Result cannot be delivered (parent finished or channel full) | `SubTurnOrphanPayload{ParentTurnID, ChildTurnID, Reason}` |
+
+## API Reference
+
+### SpawnSubTurn (Public Entry Point)
+
+```go
+func SpawnSubTurn(ctx context.Context, cfg SubTurnConfig) (*tools.ToolResult, error)
+```
+
+This is the exported package-level entry point for agent-internal code (e.g., tests, direct invocations). It retrieves `AgentLoop` and `turnState` from context and delegates to the internal `spawnSubTurn`.
+
+**Requirements:**
+- `AgentLoop` must be injected into context via `WithAgentLoop()`
+- Parent `turnState` must exist in context (automatically set when called from tools)
+
+**Returns:**
+- `*tools.ToolResult`: Contains `ForLLM` field with the sub-turn's output
+- `error`: One of the defined error types or context errors
+
+### AgentLoopSpawner (Interface Implementation)
+
+```go
+type AgentLoopSpawner struct { al *AgentLoop }
+
+func (s *AgentLoopSpawner) SpawnSubTurn(ctx context.Context, cfg tools.SubTurnConfig) (*tools.ToolResult, error)
+```
+
+This implements the `tools.SubTurnSpawner` interface for use by tools that need to spawn sub-turns without a direct import of the `agent` package (avoiding circular dependencies). It converts `tools.SubTurnConfig` → `agent.SubTurnConfig` before delegating to the internal `spawnSubTurn`.
+
+### NewSubTurnSpawner
+
+```go
+func NewSubTurnSpawner(al *AgentLoop) *AgentLoopSpawner
+```
+
+Creates a new spawner instance for the given AgentLoop. Pass the returned value to `SpawnTool.SetSpawner()` or `SubagentTool.SetSpawner()` during tool registration.
+
+### Continue
+
+```go
+func (al *AgentLoop) Continue(ctx context.Context, sessionKey string) error
+```
+
+Resumes an idle agent turn by injecting any queued steering messages as a new LLM iteration. Used when the agent is waiting and a deferred steering message needs to be processed without a new inbound message arriving.
+
+## Context Propagation
+
+SubTurn relies on context values for proper operation:
+
+| Context Key | Purpose |
+|:------------|:--------|
+| `agentLoopKey` | Stores `*AgentLoop` for tool access and SubTurn spawning |
+| `turnStateKey` | Stores `*turnState` for hierarchy tracking and result delivery |
+
+### Injecting Dependencies
+
+```go
+// Before calling tools that may spawn SubTurns
+ctx = WithAgentLoop(ctx, agentLoop)
+ctx = withTurnState(ctx, turnState)
+```
+
+### Independent Child Context
+
+**Important**: The child SubTurn uses an **independent context** derived from `context.Background()`, not from the parent context. This design choice:
+
+- Allows critical SubTurns to continue after parent cancellation
+- Prevents parent timeout from affecting child execution
+- Child has its own timeout for self-protection (`Timeout` config or 5 minutes default)
+
+## Error Types
+
+| Error | Condition |
+|:------|:----------|
+| `ErrDepthLimitExceeded` | SubTurn depth exceeds 3 levels |
+| `ErrInvalidSubTurnConfig` | Required field `Model` is empty |
+| `ErrConcurrencyTimeout` | All 5 concurrency slots occupied for 30+ seconds |
+| Context errors | Parent context cancelled during semaphore acquisition |
+
+## Thread Safety
+
+SubTurns are designed for concurrent execution:
+
+- **Parent-child relationships**: Managed under mutex (`parentTS.mu.Lock()`)
+- **Active turn tracking**: Uses `sync.Map` for concurrent access to `activeTurnStates`
+- **ID generation**: Uses `atomic.Int64` for unique SubTurn IDs (format: `subturn-N`, globally monotonic per `AgentLoop` instance)
+- **Result delivery**: Reads parent state under lock, releases before channel send (small race window acceptable)
+
+## Orphan Results
+
+An orphan result occurs when:
+1. Parent turn finishes before the SubTurn completes
+2. The `pendingResults` channel is full (buffer size: 16)
+
+When a result becomes orphan:
+- `SubTurnOrphanResultEvent` is emitted to EventBus
+- The result is **NOT** delivered to the LLM context
+- External systems can listen to this event for custom handling
+
+### Preventing Orphan Results
+- Use `Critical: true` for important SubTurns that must complete
+- Monitor `SubTurnOrphanResultEvent` for observability
+- Consider the 16-buffer limit when spawning many async SubTurns
+
+## Tool Inheritance
+
+### When `cfg.Tools` is empty:
+- SubTurn inherits **all** tools from the parent agent
+- Tools are registered in a new `ToolRegistry` instance
+- Tool TTL is managed independently from parent
+
+### When `cfg.Tools` is specified:
+- Only the specified tools are available to the SubTurn
+- Parent tools are **NOT** merged
+- Use this to restrict SubTurn capabilities for security or focus
+
+**Example - Restricted SubTurn:**
+```go
+cfg := agent.SubTurnConfig{
+ Model: "gpt-4o-mini",
+ Tools: []tools.Tool{readOnlyTool}, // Only read-only access
+ SystemPrompt: "Analyze the file structure...",
+}
+```
+
+## Reference
+
+| Constant | Value |
+|:---------|:------|
+| `maxSubTurnDepth` | 3 |
+| `maxConcurrentSubTurns` | 5 |
+| `concurrencyTimeout` | 30s |
+| `defaultSubTurnTimeout` | 5m |
+| `maxEphemeralHistorySize` | 50 messages |
+| `pendingResults` buffer | 16 |
+| `MaxContextRunes` default | 75% of model context window |
diff --git a/docs/tools_configuration.md b/docs/tools_configuration.md
index 8c8eb31f0..d0160050d 100644
--- a/docs/tools_configuration.md
+++ b/docs/tools_configuration.md
@@ -30,13 +30,23 @@ PicoClaw's tools configuration is located in the `tools` field of `config.json`.
Web tools are used for web search and fetching.
+### Web Fetcher
+General settings for fetching and processing webpage content.
+
+| Config | Type | Default | Description |
+|---------------------|--------|---------------|-----------------------------------------------------------------------------------------------|
+| `enabled` | bool | true | Enable the webpage fetching capability. |
+| `fetch_limit_bytes` | int | 10485760 | Maximum size of the webpage payload to fetch, in bytes (default is 10MB). |
+| `format` | string | "plaintext" | Output format of the fetched content. Options: `plaintext` or `markdown` (recommended). |
+
### Brave
-| Config | Type | Default | Description |
-|---------------|--------|---------|---------------------------|
-| `enabled` | bool | false | Enable Brave search |
-| `api_key` | string | - | Brave Search API key |
-| `max_results` | int | 5 | Maximum number of results |
+| Config | Type | Default | Description |
+|---------------|----------|---------|------------------------------------------------|
+| `enabled` | bool | false | Enable Brave search |
+| `api_key` | string | - | Brave Search API key |
+| `api_keys` | string[] | - | Multiple API keys for rotation (takes priority over `api_key`) |
+| `max_results` | int | 5 | Maximum number of results |
### DuckDuckGo
@@ -47,11 +57,46 @@ Web tools are used for web search and fetching.
### Perplexity
+| Config | Type | Default | Description |
+|---------------|----------|---------|------------------------------------------------|
+| `enabled` | bool | false | Enable Perplexity search |
+| `api_key` | string | - | Perplexity API key |
+| `api_keys` | string[] | - | Multiple API keys for rotation (takes priority over `api_key`) |
+| `max_results` | int | 5 | Maximum number of results |
+
+### Tavily
+
| Config | Type | Default | Description |
|---------------|--------|---------|---------------------------|
-| `enabled` | bool | false | Enable Perplexity search |
-| `api_key` | string | - | Perplexity API key |
-| `max_results` | int | 5 | Maximum number of results |
+| `enabled` | bool | false | Enable Tavily search |
+| `api_key` | string | - | Tavily API key |
+| `base_url` | string | - | Custom Tavily API base URL |
+| `max_results` | int | 0 | Maximum number of results (0 = default) |
+
+### SearXNG
+
+| Config | Type | Default | Description |
+|---------------|--------|--------------------------|---------------------------|
+| `enabled` | bool | false | Enable SearXNG search |
+| `base_url` | string | `http://localhost:8888` | SearXNG instance URL |
+| `max_results` | int | 5 | Maximum number of results |
+
+### GLM Search
+
+| Config | Type | Default | Description |
+|-----------------|--------|------------------------------------------------------|---------------------------|
+| `enabled` | bool | false | Enable GLM Search |
+| `api_key` | string | - | GLM API key |
+| `base_url` | string | `https://open.bigmodel.cn/api/paas/v4/web_search` | GLM Search API URL |
+| `search_engine` | string | `search_std` | Search engine type |
+| `max_results` | int | 5 | Maximum number of results |
+
+### Additional Web Settings
+
+| Config | Type | Default | Description |
+|--------------------------|----------|---------|----------------------------------------------------------------|
+| `prefer_native` | bool | true | Prefer provider's native search over configured search engines |
+| `private_host_whitelist` | string[] | `[]` | Private/internal hosts allowed for web fetching |
## Exec Tool
@@ -59,9 +104,32 @@ The exec tool is used to execute shell commands.
| Config | Type | Default | Description |
|------------------------|-------|---------|--------------------------------------------|
+| `enabled` | bool | true | Enable the exec tool |
| `enable_deny_patterns` | bool | true | Enable default dangerous command blocking |
| `custom_deny_patterns` | array | [] | Custom deny patterns (regular expressions) |
+### Disabling the Exec Tool
+
+To completely disable the `exec` tool, set `enabled` to `false`:
+
+**Via config file:**
+```json
+{
+ "tools": {
+ "exec": {
+ "enabled": false
+ }
+ }
+}
+```
+
+**Via environment variable:**
+```bash
+PICOCLAW_TOOLS_EXEC_ENABLED=false
+```
+
+> **Note:** When disabled, the agent will not be able to execute shell commands. This also affects the Cron tool's ability to run scheduled shell commands.
+
### Functionality
- **`enable_deny_patterns`**: Set to `false` to completely disable the default dangerous command blocking patterns
@@ -84,6 +152,22 @@ By default, PicoClaw blocks the following dangerous commands:
- Git: `git push`, `git force`
- Other: `eval`, `source *.sh`
+### Known Architectural Limitation
+
+The exec guard only validates the top-level command sent to PicoClaw. It does **not** recursively inspect child
+processes spawned by build tools or scripts after that command starts running.
+
+Examples of workflows that can bypass the direct command guard once the initial command is allowed:
+
+- `make run`
+- `go run ./cmd/...`
+- `cargo run`
+- `npm run build`
+
+This means the guard is useful for blocking obviously dangerous direct commands, but it is **not** a full sandbox for
+unreviewed build pipelines. If your threat model includes untrusted code in the workspace, use stronger isolation such
+as containers, VMs, or an approval flow around build-and-run commands.
+
### Configuration Example
```json
@@ -107,6 +191,7 @@ The cron tool is used for scheduling periodic tasks.
| Config | Type | Default | Description |
|------------------------|------|---------|------------------------------------------------|
| `exec_timeout_minutes` | int | 5 | Execution timeout in minutes, 0 means no limit |
+| `allow_command` | bool | false | Allow cron tasks to execute shell commands |
## MCP Tool
@@ -133,7 +218,7 @@ and injected into the context for a configured number of turns (`ttl`).
| Config | Type | Default | Description |
|----------------------|------|---------|-----------------------------------------------------------------------------------------------------------------------------------|
-| `enabled` | bool | false | If true, MCP tools are hidden and loaded on-demand via search. If false, all tools are loaded |
+| `enabled` | bool | false | Global default: if `true`, all MCP tools are hidden and loaded on-demand via search; if `false`, all tools are loaded into context. Individual servers can override this with the per-server `deferred` field. |
| `ttl` | int | 5 | Number of conversational turns a discovered tool remains unlocked |
| `max_search_results` | int | 5 | Maximum number of tools returned per search query |
| `use_bm25` | bool | true | Enable the natural language/keyword search tool (`tool_search_tool_bm25`). **Warning**: consumes more resources than regex search |
@@ -144,16 +229,17 @@ and injected into the context for a configured number of turns (`ttl`).
### Per-Server Config
-| Config | Type | Required | Description |
-|------------|--------|----------|--------------------------------------------|
-| `enabled` | bool | yes | Enable this MCP server |
-| `type` | string | no | Transport type: `stdio`, `sse`, `http` |
-| `command` | string | stdio | Executable command for stdio transport |
-| `args` | array | no | Command arguments for stdio transport |
-| `env` | object | no | Environment variables for stdio process |
-| `env_file` | string | no | Path to environment file for stdio process |
-| `url` | string | sse/http | Endpoint URL for `sse`/`http` transport |
-| `headers` | object | no | HTTP headers for `sse`/`http` transport |
+| Config | Type | Required | Description |
+|------------|---------|----------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------|
+| `enabled` | bool | yes | Enable this MCP server |
+| `deferred` | bool | no | Override deferred mode for this server only. `true` = tools are hidden and discoverable via search; `false` = tools are always visible in context. When omitted, the global `discovery.enabled` value applies. |
+| `type` | string | no | Transport type: `stdio`, `sse`, `http` |
+| `command` | string | stdio | Executable command for stdio transport |
+| `args` | array | no | Command arguments for stdio transport |
+| `env` | object | no | Environment variables for stdio process |
+| `env_file` | string | no | Path to environment file for stdio process |
+| `url` | string | sse/http | Endpoint URL for `sse`/`http` transport |
+| `headers` | object | no | HTTP headers for `sse`/`http` transport |
### Transport Behavior
@@ -266,6 +352,50 @@ dynamically only when requested by the user.*
}
```
+#### 4) Mixed setup: per-server deferred override
+
+*Discovery is enabled globally, but `filesystem` is pinned as always-visible while `context7` follows the global
+default (deferred). `aws` explicitly opts in to deferred mode even though it is the same as the global default.*
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "discovery": {
+ "enabled": true,
+ "ttl": 5,
+ "max_search_results": 5,
+ "use_bm25": true
+ },
+ "servers": {
+ "filesystem": {
+ "enabled": true,
+ "command": "npx",
+ "args": ["-y", "@modelcontextprotocol/server-filesystem", "/workspace"],
+ "deferred": false
+ },
+ "context7": {
+ "enabled": true,
+ "command": "npx",
+ "args": ["-y", "@upstash/context7-mcp"]
+ },
+ "aws": {
+ "enabled": true,
+ "command": "npx",
+ "args": ["-y", "aws-mcp-server"],
+ "deferred": true
+ }
+ }
+ }
+ }
+}
+```
+
+> **Tip:** `deferred` on a per-server basis is independent of `discovery.enabled`. You can keep
+> `discovery.enabled: false` globally (all tools visible by default) and still mark individual
+> high-volume servers as `"deferred": true` to avoid polluting the context with their tools.
+
## Skills Tool
The skills tool configures skill discovery and installation via registries like ClawHub.
@@ -277,9 +407,27 @@ The skills tool configures skill discovery and installation via registries like
| `registries.clawhub.enabled` | bool | true | Enable ClawHub registry |
| `registries.clawhub.base_url` | string | `https://clawhub.ai` | ClawHub base URL |
| `registries.clawhub.auth_token` | string | `""` | Optional Bearer token for higher rate limits |
-| `registries.clawhub.search_path` | string | `/api/v1/search` | Search API path |
-| `registries.clawhub.skills_path` | string | `/api/v1/skills` | Skills API path |
-| `registries.clawhub.download_path` | string | `/api/v1/download` | Download API path |
+| `registries.clawhub.search_path` | string | `""` | Search API path |
+| `registries.clawhub.skills_path` | string | `""` | Skills API path |
+| `registries.clawhub.download_path` | string | `""` | Download API path |
+| `registries.clawhub.timeout` | int | 0 | Request timeout in seconds (0 = default) |
+| `registries.clawhub.max_zip_size` | int | 0 | Max skill zip size in bytes (0 = default) |
+| `registries.clawhub.max_response_size` | int | 0 | Max API response size in bytes (0 = default) |
+
+### GitHub Integration
+
+| Config | Type | Default | Description |
+|------------------|--------|---------|--------------------------------------|
+| `github.proxy` | string | `""` | HTTP proxy for GitHub API requests |
+| `github.token` | string | `""` | GitHub personal access token |
+
+### Search Settings
+
+| Config | Type | Default | Description |
+|---------------------------|------|---------|--------------------------------------------|
+| `max_concurrent_searches` | int | 2 | Max concurrent skill search requests |
+| `search_cache.max_size` | int | 50 | Max cached search results |
+| `search_cache.ttl_seconds`| int | 300 | Cache TTL in seconds |
### Configuration Example
@@ -291,11 +439,17 @@ The skills tool configures skill discovery and installation via registries like
"clawhub": {
"enabled": true,
"base_url": "https://clawhub.ai",
- "auth_token": "",
- "search_path": "/api/v1/search",
- "skills_path": "/api/v1/skills",
- "download_path": "/api/v1/download"
+ "auth_token": ""
}
+ },
+ "github": {
+ "proxy": "",
+ "token": ""
+ },
+ "max_concurrent_searches": 2,
+ "search_cache": {
+ "max_size": 50,
+ "ttl_seconds": 300
}
}
}
@@ -309,6 +463,7 @@ All configuration options can be overridden via environment variables with the f
For example:
- `PICOCLAW_TOOLS_WEB_BRAVE_ENABLED=true`
+- `PICOCLAW_TOOLS_EXEC_ENABLED=false`
- `PICOCLAW_TOOLS_EXEC_ENABLE_DENY_PATTERNS=false`
- `PICOCLAW_TOOLS_CRON_EXEC_TIMEOUT_MINUTES=10`
- `PICOCLAW_TOOLS_MCP_ENABLED=true`
diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md
index 219d2c6e3..096beec78 100644
--- a/docs/troubleshooting.md
+++ b/docs/troubleshooting.md
@@ -14,7 +14,7 @@
**Fix:** In `~/.picoclaw/config.json` (or your config path):
-1. **agents.defaults.model** must match a `model_name` in `model_list` (e.g. `"openrouter-free"`).
+1. **agents.defaults.model_name** must match a `model_name` in `model_list` (e.g. `"openrouter-free"`).
2. That entry’s **model** must be a valid OpenRouter model ID, for example:
- `"openrouter/free"` – auto free-tier
- `"google/gemini-2.0-flash-exp:free"`
@@ -26,7 +26,7 @@ Example snippet:
{
"agents": {
"defaults": {
- "model": "openrouter-free"
+ "model_name": "openrouter-free"
}
},
"model_list": [
diff --git a/docs/vi/ANTIGRAVITY_AUTH.md b/docs/vi/ANTIGRAVITY_AUTH.md
new file mode 100644
index 000000000..783dc5181
--- /dev/null
+++ b/docs/vi/ANTIGRAVITY_AUTH.md
@@ -0,0 +1,807 @@
+> Quay lại [README](../../README.vi.md)
+
+# Hướng dẫn Xác thực và Tích hợp Antigravity
+
+## Tổng quan
+
+**Antigravity** (Google Cloud Code Assist) là nhà cung cấp mô hình AI được Google hỗ trợ, cung cấp quyền truy cập vào các mô hình như Claude Opus 4.6 và Gemini thông qua hạ tầng đám mây của Google. Tài liệu này cung cấp hướng dẫn đầy đủ về cách xác thực hoạt động, cách lấy danh sách mô hình và cách triển khai nhà cung cấp mới trong PicoClaw.
+
+---
+
+## Mục lục
+
+1. [Luồng xác thực](#luồng-xác-thực)
+2. [Chi tiết triển khai OAuth](#chi-tiết-triển-khai-oauth)
+3. [Quản lý token](#quản-lý-token)
+4. [Lấy danh sách mô hình](#lấy-danh-sách-mô-hình)
+5. [Theo dõi mức sử dụng](#theo-dõi-mức-sử-dụng)
+6. [Cấu trúc plugin nhà cung cấp](#cấu-trúc-plugin-nhà-cung-cấp)
+7. [Yêu cầu tích hợp](#yêu-cầu-tích-hợp)
+8. [Các endpoint API](#các-endpoint-api)
+9. [Cấu hình](#cấu-hình)
+10. [Tạo nhà cung cấp mới trong PicoClaw](#tạo-nhà-cung-cấp-mới-trong-picoclaw)
+
+---
+
+## Luồng xác thực
+
+### 1. OAuth 2.0 với PKCE
+
+Antigravity sử dụng **OAuth 2.0 với PKCE (Proof Key for Code Exchange)** để xác thực an toàn:
+
+```
+┌─────────────┐ ┌─────────────────┐
+│ Client │ ───(1) Generate PKCE Pair────────> │ │
+│ │ ───(2) Open Auth URL─────────────> │ Google OAuth │
+│ │ │ Server │
+│ │ <──(3) Redirect with Code───────── │ │
+│ │ └─────────────────┘
+│ │ ───(4) Exchange Code for Tokens──> │ Token URL │
+│ │ │ │
+│ │ <──(5) Access + Refresh Tokens──── │ │
+└─────────────┘ └─────────────────┘
+```
+
+### 2. Các bước chi tiết
+
+#### Bước 1: Tạo tham số PKCE
+```typescript
+function generatePkce(): { verifier: string; challenge: string } {
+ const verifier = randomBytes(32).toString("hex");
+ const challenge = createHash("sha256").update(verifier).digest("base64url");
+ return { verifier, challenge };
+}
+```
+
+#### Bước 2: Xây dựng URL ủy quyền
+```typescript
+const AUTH_URL = "https://accounts.google.com/o/oauth2/v2/auth";
+const REDIRECT_URI = "http://localhost:51121/oauth-callback";
+
+function buildAuthUrl(params: { challenge: string; state: string }): string {
+ const url = new URL(AUTH_URL);
+ url.searchParams.set("client_id", CLIENT_ID);
+ url.searchParams.set("response_type", "code");
+ url.searchParams.set("redirect_uri", REDIRECT_URI);
+ url.searchParams.set("scope", SCOPES.join(" "));
+ url.searchParams.set("code_challenge", params.challenge);
+ url.searchParams.set("code_challenge_method", "S256");
+ url.searchParams.set("state", params.state);
+ url.searchParams.set("access_type", "offline");
+ url.searchParams.set("prompt", "consent");
+ return url.toString();
+}
+```
+
+**Các phạm vi quyền cần thiết:**
+```typescript
+const SCOPES = [
+ "https://www.googleapis.com/auth/cloud-platform",
+ "https://www.googleapis.com/auth/userinfo.email",
+ "https://www.googleapis.com/auth/userinfo.profile",
+ "https://www.googleapis.com/auth/cclog",
+ "https://www.googleapis.com/auth/experimentsandconfigs",
+];
+```
+
+#### Bước 3: Xử lý callback OAuth
+
+**Chế độ tự động (Phát triển cục bộ):**
+- Khởi động máy chủ HTTP cục bộ trên cổng 51121
+- Chờ chuyển hướng từ Google
+- Trích xuất mã ủy quyền từ tham số truy vấn
+
+**Chế độ thủ công (Từ xa/Không có giao diện):**
+- Hiển thị URL ủy quyền cho người dùng
+- Người dùng hoàn tất xác thực trong trình duyệt
+- Người dùng dán URL chuyển hướng đầy đủ vào terminal
+- Phân tích mã từ URL đã dán
+
+#### Bước 4: Đổi mã lấy token
+```typescript
+const TOKEN_URL = "https://oauth2.googleapis.com/token";
+
+async function exchangeCode(params: {
+ code: string;
+ verifier: string;
+}): Promise<{ access: string; refresh: string; expires: number }> {
+ const response = await fetch(TOKEN_URL, {
+ method: "POST",
+ headers: { "Content-Type": "application/x-www-form-urlencoded" },
+ body: new URLSearchParams({
+ client_id: CLIENT_ID,
+ client_secret: CLIENT_SECRET,
+ code: params.code,
+ grant_type: "authorization_code",
+ redirect_uri: REDIRECT_URI,
+ code_verifier: params.verifier,
+ }),
+ });
+
+ const data = await response.json();
+
+ return {
+ access: data.access_token,
+ refresh: data.refresh_token,
+ expires: Date.now() + data.expires_in * 1000 - 5 * 60 * 1000, // 5 min buffer
+ };
+}
+```
+
+#### Bước 5: Lấy dữ liệu người dùng bổ sung
+
+**Email người dùng:**
+```typescript
+async function fetchUserEmail(accessToken: string): Promise {
+ const response = await fetch(
+ "https://www.googleapis.com/oauth2/v1/userinfo?alt=json",
+ { headers: { Authorization: `Bearer ${accessToken}` } }
+ );
+ const data = await response.json();
+ return data.email;
+}
+```
+
+**ID dự án (Bắt buộc cho các lệnh gọi API):**
+```typescript
+async function fetchProjectId(accessToken: string): Promise {
+ const headers = {
+ Authorization: `Bearer ${accessToken}`,
+ "Content-Type": "application/json",
+ "User-Agent": "google-api-nodejs-client/9.15.1",
+ "X-Goog-Api-Client": "google-cloud-sdk vscode_cloudshelleditor/0.1",
+ "Client-Metadata": JSON.stringify({
+ ideType: "IDE_UNSPECIFIED",
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+ }),
+ };
+
+ const response = await fetch(
+ "https://cloudcode-pa.googleapis.com/v1internal:loadCodeAssist",
+ {
+ method: "POST",
+ headers,
+ body: JSON.stringify({
+ metadata: {
+ ideType: "IDE_UNSPECIFIED",
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+ },
+ }),
+ }
+ );
+
+ const data = await response.json();
+ return data.cloudaicompanionProject || "rising-fact-p41fc"; // Giá trị mặc định dự phòng
+}
+```
+
+---
+
+## Chi tiết triển khai OAuth
+
+### Thông tin xác thực client
+
+**Quan trọng:** Các giá trị này được mã hóa base64 trong mã nguồn để đồng bộ với pi-ai:
+
+```typescript
+const decode = (s: string) => Buffer.from(s, "base64").toString();
+
+const CLIENT_ID = decode(
+ "MTA3MTAwNjA2MDU5MS10bWhzc2luMmgyMWxjcmUyMzV2dG9sb2poNGc0MDNlcC5hcHBzLmdvb2dsZXVzZXJjb250ZW50LmNvbQ=="
+);
+const CLIENT_SECRET = decode("R09DU1BYLUs1OEZXUjQ4NkxkTEoxbUxCOHNYQzR6NnFEQWY=");
+```
+
+### Các chế độ luồng OAuth
+
+1. **Luồng tự động** (Máy cục bộ có trình duyệt):
+ - Tự động mở trình duyệt
+ - Máy chủ callback cục bộ bắt chuyển hướng
+ - Không cần tương tác người dùng sau xác thực ban đầu
+
+2. **Luồng thủ công** (Từ xa/Không có giao diện/WSL2):
+ - Hiển thị URL để sao chép-dán thủ công
+ - Người dùng hoàn tất xác thực trong trình duyệt bên ngoài
+ - Người dùng dán lại URL chuyển hướng đầy đủ
+
+```typescript
+function shouldUseManualOAuthFlow(isRemote: boolean): boolean {
+ return isRemote || isWSL2Sync();
+}
+```
+
+---
+
+## Quản lý token
+
+### Cấu trúc hồ sơ xác thực
+
+```typescript
+type OAuthCredential = {
+ type: "oauth";
+ provider: "google-antigravity";
+ access: string; // Token truy cập
+ refresh: string; // Token làm mới
+ expires: number; // Dấu thời gian hết hạn (ms kể từ epoch)
+ email?: string; // Email người dùng
+ projectId?: string; // ID dự án Google Cloud
+};
+```
+
+### Làm mới token
+
+Thông tin xác thực bao gồm token làm mới có thể được sử dụng để lấy token truy cập mới khi token hiện tại hết hạn. Thời gian hết hạn được đặt với bộ đệm 5 phút để tránh điều kiện tranh chấp.
+
+---
+
+## Lấy danh sách mô hình
+
+### Lấy các mô hình khả dụng
+
+```typescript
+const BASE_URL = "https://cloudcode-pa.googleapis.com";
+
+async function fetchAvailableModels(
+ accessToken: string,
+ projectId: string
+): Promise {
+ const headers = {
+ Authorization: `Bearer ${accessToken}`,
+ "Content-Type": "application/json",
+ "User-Agent": "antigravity",
+ "X-Goog-Api-Client": "google-cloud-sdk vscode_cloudshelleditor/0.1",
+ };
+
+ const response = await fetch(
+ `${BASE_URL}/v1internal:fetchAvailableModels`,
+ {
+ method: "POST",
+ headers,
+ body: JSON.stringify({ project: projectId }),
+ }
+ );
+
+ const data = await response.json();
+
+ // Trả về các mô hình kèm thông tin hạn mức
+ return Object.entries(data.models).map(([modelId, modelInfo]) => ({
+ id: modelId,
+ displayName: modelInfo.displayName,
+ quotaInfo: {
+ remainingFraction: modelInfo.quotaInfo?.remainingFraction,
+ resetTime: modelInfo.quotaInfo?.resetTime,
+ isExhausted: modelInfo.quotaInfo?.isExhausted,
+ },
+ }));
+}
+```
+
+### Định dạng phản hồi
+
+```typescript
+type FetchAvailableModelsResponse = {
+ models?: Record;
+};
+```
+
+---
+
+## Theo dõi mức sử dụng
+
+### Lấy dữ liệu sử dụng
+
+```typescript
+export async function fetchAntigravityUsage(
+ token: string,
+ timeoutMs: number
+): Promise {
+ // 1. Lấy thông tin tín dụng và gói dịch vụ
+ const loadCodeAssistRes = await fetch(
+ `${BASE_URL}/v1internal:loadCodeAssist`,
+ {
+ method: "POST",
+ headers: {
+ Authorization: `Bearer ${token}`,
+ "Content-Type": "application/json",
+ },
+ body: JSON.stringify({
+ metadata: {
+ ideType: "ANTIGRAVITY",
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+ },
+ }),
+ }
+ );
+
+ // Trích xuất thông tin tín dụng
+ const { availablePromptCredits, planInfo, currentTier } = data;
+
+ // 2. Lấy hạn mức mô hình
+ const modelsRes = await fetch(
+ `${BASE_URL}/v1internal:fetchAvailableModels`,
+ {
+ method: "POST",
+ headers: { Authorization: `Bearer ${token}` },
+ body: JSON.stringify({ project: projectId }),
+ }
+ );
+
+ // Xây dựng cửa sổ sử dụng
+ return {
+ provider: "google-antigravity",
+ displayName: "Google Antigravity",
+ windows: [
+ { label: "Credits", usedPercent: calculateUsedPercent(available, monthly) },
+ // Hạn mức từng mô hình...
+ ],
+ plan: currentTier?.name || planType,
+ };
+}
+```
+
+### Cấu trúc phản hồi sử dụng
+
+```typescript
+type ProviderUsageSnapshot = {
+ provider: "google-antigravity";
+ displayName: string;
+ windows: UsageWindow[];
+ plan?: string;
+ error?: string;
+};
+
+type UsageWindow = {
+ label: string; // "Credits" hoặc ID mô hình
+ usedPercent: number; // 0-100
+ resetAt?: number; // Dấu thời gian khi hạn mức được đặt lại
+};
+```
+
+---
+
+## Cấu trúc plugin nhà cung cấp
+
+### Định nghĩa plugin
+
+```typescript
+const antigravityPlugin = {
+ id: "google-antigravity-auth",
+ name: "Google Antigravity Auth",
+ description: "OAuth flow for Google Antigravity (Cloud Code Assist)",
+ configSchema: emptyPluginConfigSchema(),
+
+ register(api: PicoClawPluginApi) {
+ api.registerProvider({
+ id: "google-antigravity",
+ label: "Google Antigravity",
+ docsPath: "/providers/models",
+ aliases: ["antigravity"],
+
+ auth: [
+ {
+ id: "oauth",
+ label: "Google OAuth",
+ hint: "PKCE + localhost callback",
+ kind: "oauth",
+ run: async (ctx: ProviderAuthContext) => {
+ // Triển khai OAuth tại đây
+ },
+ },
+ ],
+ });
+ },
+};
+```
+
+### ProviderAuthContext
+
+```typescript
+type ProviderAuthContext = {
+ config: PicoClawConfig;
+ agentDir?: string;
+ workspaceDir?: string;
+ prompter: WizardPrompter; // Lời nhắc/thông báo UI
+ runtime: RuntimeEnv; // Ghi log, v.v.
+ isRemote: boolean; // Có đang chạy từ xa không
+ openUrl: (url: string) => Promise; // Mở trình duyệt
+ oauth: {
+ createVpsAwareHandlers: Function;
+ };
+};
+```
+
+### ProviderAuthResult
+
+```typescript
+type ProviderAuthResult = {
+ profiles: Array<{
+ profileId: string;
+ credential: AuthProfileCredential;
+ }>;
+ configPatch?: Partial;
+ defaultModel?: string;
+ notes?: string[];
+};
+```
+
+---
+
+## Yêu cầu tích hợp
+
+### 1. Môi trường/Phụ thuộc cần thiết
+
+- Go ≥ 1.25
+- Mã nguồn PicoClaw (`pkg/providers/` và `pkg/auth/`)
+- Các gói thư viện chuẩn `crypto` và `net/http`
+
+### 2. Các header bắt buộc cho lệnh gọi API
+
+```typescript
+const REQUIRED_HEADERS = {
+ "Authorization": `Bearer ${accessToken}`,
+ "Content-Type": "application/json",
+ "User-Agent": "antigravity", // hoặc "google-api-nodejs-client/9.15.1"
+ "X-Goog-Api-Client": "google-cloud-sdk vscode_cloudshelleditor/0.1",
+};
+
+// Đối với các lệnh gọi loadCodeAssist, cũng bao gồm:
+const CLIENT_METADATA = {
+ ideType: "ANTIGRAVITY", // hoặc "IDE_UNSPECIFIED"
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+};
+```
+
+### 3. Làm sạch schema mô hình
+
+Antigravity sử dụng các mô hình tương thích Gemini, vì vậy schema công cụ phải được làm sạch:
+
+```typescript
+const GOOGLE_SCHEMA_UNSUPPORTED_KEYWORDS = new Set([
+ "patternProperties",
+ "additionalProperties",
+ "$schema",
+ "$id",
+ "$ref",
+ "$defs",
+ "definitions",
+ "examples",
+ "minLength",
+ "maxLength",
+ "minimum",
+ "maximum",
+ "multipleOf",
+ "pattern",
+ "format",
+ "minItems",
+ "maxItems",
+ "uniqueItems",
+ "minProperties",
+ "maxProperties",
+]);
+
+// Làm sạch schema trước khi gửi
+function cleanToolSchemaForGemini(schema: Record): unknown {
+ // Xóa các từ khóa không được hỗ trợ
+ // Đảm bảo cấp cao nhất có type: "object"
+ // Làm phẳng các union anyOf/oneOf
+}
+```
+
+### 4. Xử lý khối suy nghĩ (Mô hình Claude)
+
+Đối với các mô hình Claude qua Antigravity, khối suy nghĩ cần xử lý đặc biệt:
+
+```typescript
+const ANTIGRAVITY_SIGNATURE_RE = /^[A-Za-z0-9+/]+={0,2}$/;
+
+export function sanitizeAntigravityThinkingBlocks(
+ messages: AgentMessage[]
+): AgentMessage[] {
+ // Xác thực chữ ký suy nghĩ
+ // Chuẩn hóa các trường chữ ký
+ // Loại bỏ các khối suy nghĩ chưa ký
+}
+```
+
+---
+
+## Các endpoint API
+
+### Endpoint xác thực
+
+| Endpoint | Phương thức | Mục đích |
+|----------|------------|----------|
+| `https://accounts.google.com/o/oauth2/v2/auth` | GET | Ủy quyền OAuth |
+| `https://oauth2.googleapis.com/token` | POST | Trao đổi token |
+| `https://www.googleapis.com/oauth2/v1/userinfo` | GET | Thông tin người dùng (email) |
+
+### Endpoint Cloud Code Assist
+
+| Endpoint | Phương thức | Mục đích |
+|----------|------------|----------|
+| `https://cloudcode-pa.googleapis.com/v1internal:loadCodeAssist` | POST | Tải thông tin dự án, tín dụng, gói dịch vụ |
+| `https://cloudcode-pa.googleapis.com/v1internal:fetchAvailableModels` | POST | Liệt kê các mô hình khả dụng kèm hạn mức |
+| `https://cloudcode-pa.googleapis.com/v1internal:streamGenerateContent?alt=sse` | POST | Endpoint streaming chat |
+
+**Định dạng yêu cầu API (Chat):**
+Endpoint `v1internal:streamGenerateContent` yêu cầu một envelope bao bọc yêu cầu Gemini tiêu chuẩn:
+
+```json
+{
+ "project": "your-project-id",
+ "model": "model-id",
+ "request": {
+ "contents": [...],
+ "systemInstruction": {...},
+ "generationConfig": {...},
+ "tools": [...]
+ },
+ "requestType": "agent",
+ "userAgent": "antigravity",
+ "requestId": "agent-timestamp-random"
+}
+```
+
+**Định dạng phản hồi API (SSE):**
+Mỗi thông điệp SSE (`data: {...}`) được bao bọc trong trường `response`:
+
+```json
+{
+ "response": {
+ "candidates": [...],
+ "usageMetadata": {...},
+ "modelVersion": "...",
+ "responseId": "..."
+ },
+ "traceId": "...",
+ "metadata": {}
+}
+```
+
+---
+
+## Cấu hình
+
+### Cấu hình config.json
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "gemini-flash",
+ "model": "antigravity/gemini-3-flash",
+ "auth_method": "oauth"
+ }
+ ],
+ "agents": {
+ "defaults": {
+ "model_name": "gemini-flash"
+ }
+ }
+}
+```
+
+### Lưu trữ hồ sơ xác thực
+
+Hồ sơ xác thực được lưu trữ trong `~/.picoclaw/auth.json`:
+
+```json
+{
+ "credentials": {
+ "google-antigravity": {
+ "access_token": "ya29...",
+ "refresh_token": "1//...",
+ "expires_at": "2026-01-01T00:00:00Z",
+ "provider": "google-antigravity",
+ "auth_method": "oauth",
+ "email": "user@example.com",
+ "project_id": "my-project-id"
+ }
+ }
+}
+```
+
+---
+
+## Tạo nhà cung cấp mới trong PicoClaw
+
+Các nhà cung cấp PicoClaw được triển khai dưới dạng gói Go trong `pkg/providers/`. Để thêm nhà cung cấp mới:
+
+### Triển khai từng bước
+
+#### 1. Tạo file nhà cung cấp
+
+Tạo file Go mới trong `pkg/providers/`:
+
+```
+pkg/providers/
+└── your_provider.go
+```
+
+#### 2. Triển khai interface Provider
+
+Nhà cung cấp của bạn phải triển khai interface `Provider` được định nghĩa trong `pkg/providers/types.go`:
+
+```go
+package providers
+
+type YourProvider struct {
+ apiKey string
+ apiBase string
+}
+
+func NewYourProvider(apiKey, apiBase, proxy string) *YourProvider {
+ if apiBase == "" {
+ apiBase = "https://api.your-provider.com/v1"
+ }
+ return &YourProvider{apiKey: apiKey, apiBase: apiBase}
+}
+
+func (p *YourProvider) Chat(ctx context.Context, messages []Message, tools []Tool, cb StreamCallback) error {
+ // Triển khai hoàn thành chat với streaming
+}
+```
+
+#### 3. Đăng ký trong factory
+
+Thêm nhà cung cấp của bạn vào switch giao thức trong `pkg/providers/factory.go`:
+
+```go
+case "your-provider":
+ return NewYourProvider(sel.apiKey, sel.apiBase, sel.proxy), nil
+```
+
+#### 4. Thêm cấu hình mặc định (Tùy chọn)
+
+Thêm mục mặc định trong `pkg/config/defaults.go`:
+
+```go
+{
+ ModelName: "your-model",
+ Model: "your-provider/model-name",
+ APIKey: "",
+},
+```
+
+#### 5. Thêm hỗ trợ xác thực (Tùy chọn)
+
+Nếu nhà cung cấp của bạn yêu cầu OAuth hoặc xác thực đặc biệt, thêm case vào `cmd/picoclaw/internal/auth/helpers.go`:
+
+```go
+case "your-provider":
+ authLoginYourProvider()
+```
+
+#### 6. Cấu hình qua `config.json`
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "your-model",
+ "model": "your-provider/model-name",
+ "api_key": "your-api-key",
+ "api_base": "https://api.your-provider.com/v1"
+ }
+ ]
+}
+```
+
+---
+
+## Kiểm thử triển khai của bạn
+
+### Lệnh CLI
+
+```bash
+# Xác thực với nhà cung cấp
+picoclaw auth login --provider your-provider
+
+# Liệt kê mô hình (cho Antigravity)
+picoclaw auth models
+
+# Khởi động gateway
+picoclaw gateway
+
+# Chạy agent với mô hình cụ thể
+picoclaw agent -m "Hello" --model your-model
+```
+
+### Biến môi trường cho kiểm thử
+
+```bash
+# Ghi đè mô hình mặc định
+export PICOCLAW_AGENTS_DEFAULTS_MODEL=your-model
+
+# Ghi đè cài đặt nhà cung cấp
+export PICOCLAW_MODEL_LIST='[{"model_name":"your-model","model":"your-provider/model-name","api_key":"..."}]'
+```
+
+---
+
+## Tài liệu tham khảo
+
+- **File nguồn:**
+ - `pkg/providers/antigravity_provider.go` - Triển khai nhà cung cấp Antigravity
+ - `pkg/auth/oauth.go` - Triển khai luồng OAuth
+ - `pkg/auth/store.go` - Lưu trữ thông tin xác thực (`~/.picoclaw/auth.json`)
+ - `pkg/providers/factory.go` - Factory nhà cung cấp và định tuyến giao thức
+ - `pkg/providers/types.go` - Định nghĩa interface nhà cung cấp
+ - `cmd/picoclaw/internal/auth/helpers.go` - Lệnh CLI xác thực
+
+- **Tài liệu:**
+ - `docs/ANTIGRAVITY_USAGE.md` - Hướng dẫn sử dụng Antigravity
+ - `docs/migration/model-list-migration.md` - Hướng dẫn di chuyển
+
+---
+
+## Lưu ý
+
+1. **Dự án Google Cloud:** Antigravity yêu cầu Gemini for Google Cloud được bật trên dự án Google Cloud của bạn
+2. **Hạn mức:** Sử dụng hạn mức dự án Google Cloud (không tính phí riêng)
+3. **Truy cập mô hình:** Các mô hình khả dụng phụ thuộc vào cấu hình dự án Google Cloud của bạn
+4. **Khối suy nghĩ:** Mô hình Claude qua Antigravity yêu cầu xử lý đặc biệt khối suy nghĩ có chữ ký
+5. **Làm sạch schema:** Schema công cụ phải được làm sạch để loại bỏ các từ khóa JSON Schema không được hỗ trợ
+
+---
+
+## Xử lý lỗi thường gặp
+
+### 1. Giới hạn tốc độ (HTTP 429)
+
+Antigravity trả về lỗi 429 khi hạn mức dự án/mô hình đã cạn kiệt. Phản hồi lỗi thường chứa `quotaResetDelay` trong trường `details`.
+
+**Ví dụ lỗi 429:**
+```json
+{
+ "error": {
+ "code": 429,
+ "message": "You have exhausted your capacity on this model. Your quota will reset after 4h30m28s.",
+ "status": "RESOURCE_EXHAUSTED",
+ "details": [
+ {
+ "@type": "type.googleapis.com/google.rpc.ErrorInfo",
+ "metadata": {
+ "quotaResetDelay": "4h30m28.060903746s"
+ }
+ }
+ ]
+ }
+}
+```
+
+### 2. Phản hồi trống (Mô hình bị hạn chế)
+
+Một số mô hình có thể xuất hiện trong danh sách mô hình khả dụng nhưng trả về phản hồi trống (200 OK nhưng luồng SSE trống). Điều này thường xảy ra với các mô hình xem trước hoặc bị hạn chế mà dự án hiện tại không có quyền sử dụng.
+
+**Cách xử lý:** Coi phản hồi trống là lỗi, thông báo cho người dùng rằng mô hình có thể bị hạn chế hoặc không hợp lệ cho dự án của họ.
+
+---
+
+## Khắc phục sự cố
+
+### "Token expired" (Token đã hết hạn)
+- Làm mới token OAuth: `picoclaw auth login --provider antigravity`
+
+### "Gemini for Google Cloud is not enabled" (Gemini for Google Cloud chưa được bật)
+- Bật API trong Google Cloud Console của bạn
+
+### "Project not found" (Không tìm thấy dự án)
+- Đảm bảo dự án Google Cloud của bạn đã bật các API cần thiết
+- Kiểm tra xem ID dự án có được lấy chính xác trong quá trình xác thực không
+
+### Mô hình không xuất hiện trong danh sách
+- Xác minh xác thực OAuth đã hoàn tất thành công
+- Kiểm tra lưu trữ hồ sơ xác thực: `~/.picoclaw/auth.json`
+- Chạy lại `picoclaw auth login --provider antigravity`
diff --git a/docs/vi/ANTIGRAVITY_USAGE.md b/docs/vi/ANTIGRAVITY_USAGE.md
new file mode 100644
index 000000000..4a696f770
--- /dev/null
+++ b/docs/vi/ANTIGRAVITY_USAGE.md
@@ -0,0 +1,72 @@
+> Quay lại [README](../../README.vi.md)
+
+# Sử dụng nhà cung cấp Antigravity trong PicoClaw
+
+Hướng dẫn này giải thích cách thiết lập và sử dụng nhà cung cấp **Antigravity** (Google Cloud Code Assist) trong PicoClaw.
+
+## Điều kiện tiên quyết
+
+1. Một tài khoản Google.
+2. Đã kích hoạt Google Cloud Code Assist (thường có sẵn thông qua quy trình giới thiệu "Gemini for Google Cloud").
+
+## 1. Xác thực
+
+Để xác thực với Antigravity, chạy lệnh sau:
+
+```bash
+picoclaw auth login --provider antigravity
+```
+
+### Xác thực thủ công (Headless/VPS)
+Nếu bạn đang chạy trên máy chủ (Coolify/Docker) và không thể truy cập `localhost`, hãy làm theo các bước sau:
+1. Chạy lệnh ở trên.
+2. Sao chép URL được cung cấp và mở nó trong trình duyệt cục bộ của bạn.
+3. Hoàn tất đăng nhập.
+4. Trình duyệt của bạn sẽ chuyển hướng đến URL `localhost:51121` (trang sẽ không tải được).
+5. **Sao chép URL cuối cùng đó** từ thanh địa chỉ trình duyệt.
+6. **Dán nó vào terminal** nơi PicoClaw đang chờ.
+
+PicoClaw sẽ tự động trích xuất mã ủy quyền và hoàn tất quy trình.
+
+## 2. Quản lý mô hình
+
+### Liệt kê các mô hình khả dụng
+Để xem dự án của bạn có quyền truy cập vào những mô hình nào và kiểm tra hạn mức của chúng:
+
+```bash
+picoclaw auth models
+```
+
+### Chuyển đổi mô hình
+Bạn có thể thay đổi mô hình mặc định trong `~/.picoclaw/config.json` hoặc ghi đè qua CLI:
+
+```bash
+# Ghi đè cho một lệnh duy nhất
+picoclaw agent -m "Hello" --model claude-opus-4-6-thinking
+```
+
+## 3. Sử dụng thực tế (Coolify/Docker)
+
+Nếu bạn đang triển khai qua Coolify hoặc Docker, hãy làm theo các bước sau để kiểm tra:
+
+1. **Biến môi trường**:
+ * `PICOCLAW_AGENTS_DEFAULTS_MODEL=gemini-flash`
+2. **Lưu trữ xác thực**:
+ Nếu bạn đã đăng nhập cục bộ, bạn có thể sao chép thông tin xác thực lên máy chủ:
+ ```bash
+ scp ~/.picoclaw/auth.json user@your-server:~/.picoclaw/
+ ```
+ *Hoặc*, chạy lệnh `auth login` một lần trên máy chủ nếu bạn có quyền truy cập terminal.
+
+## 4. Khắc phục sự cố
+
+* **Phản hồi trống**: Nếu một mô hình trả về phản hồi trống, nó có thể bị hạn chế cho dự án của bạn. Hãy thử `gemini-3-flash` hoặc `claude-opus-4-6-thinking`.
+* **429 Giới hạn tốc độ**: Antigravity có hạn mức nghiêm ngặt. PicoClaw sẽ hiển thị "thời gian đặt lại" trong thông báo lỗi nếu bạn đạt đến giới hạn.
+* **404 Không tìm thấy**: Đảm bảo bạn đang sử dụng ID mô hình từ danh sách `picoclaw auth models`. Sử dụng ID ngắn (ví dụ: `gemini-3-flash`) thay vì đường dẫn đầy đủ.
+
+## 5. Tóm tắt các mô hình hoạt động tốt
+
+Dựa trên kiểm tra, các mô hình sau đáng tin cậy nhất:
+* `gemini-3-flash` (Nhanh, khả dụng cao)
+* `gemini-2.5-flash-lite` (Nhẹ)
+* `claude-opus-4-6-thinking` (Mạnh mẽ, bao gồm khả năng suy luận)
diff --git a/docs/vi/chat-apps.md b/docs/vi/chat-apps.md
new file mode 100644
index 000000000..3680fed69
--- /dev/null
+++ b/docs/vi/chat-apps.md
@@ -0,0 +1,625 @@
+# 💬 Cấu Hình Ứng Dụng Chat
+
+> Quay lại [README](../../README.vi.md)
+
+## 💬 Ứng Dụng Chat
+
+Trò chuyện với picoclaw của bạn qua Telegram, Discord, WhatsApp, Matrix, QQ, DingTalk, LINE, WeCom, Feishu, Slack, IRC, OneBot hoặc MaixCam
+
+> **Lưu ý**: Tất cả các kênh dựa trên webhook (LINE, WeCom, v.v.) được phục vụ trên một máy chủ HTTP Gateway chung (`gateway.host`:`gateway.port`, mặc định `127.0.0.1:18790`). Không có port riêng cho từng kênh. Lưu ý: Feishu sử dụng chế độ WebSocket/SDK và không sử dụng máy chủ HTTP webhook chung.
+
+| Kênh | Độ khó | Mô tả | Tài liệu |
+| -------------------- | ------------------ | ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- |
+| **Telegram** | ⭐ Dễ | Khuyến nghị, chuyển giọng nói thành văn bản, long polling (không cần IP công khai) | [Tài liệu](../channels/telegram/README.vi.md) |
+| **Discord** | ⭐ Dễ | Socket Mode, hỗ trợ nhóm/DM, hệ sinh thái bot phong phú | [Tài liệu](../channels/discord/README.vi.md) |
+| **WhatsApp** | ⭐ Dễ | Bản địa (quét QR) hoặc Bridge URL | [Tài liệu](#whatsapp) |
+| **Slack** | ⭐ Dễ | **Socket Mode** (không cần IP công khai), doanh nghiệp | [Tài liệu](../channels/slack/README.vi.md) |
+| **Matrix** | ⭐⭐ Trung bình | Giao thức liên kết, hỗ trợ tự lưu trữ | [Tài liệu](../channels/matrix/README.vi.md) |
+| **QQ** | ⭐⭐ Trung bình | API bot chính thức, cộng đồng Trung Quốc | [Tài liệu](../channels/qq/README.vi.md) |
+| **DingTalk** | ⭐⭐ Trung bình | Chế độ Stream (không cần IP công khai), doanh nghiệp | [Tài liệu](../channels/dingtalk/README.vi.md) |
+| **LINE** | ⭐⭐⭐ Nâng cao | Yêu cầu HTTPS Webhook | [Tài liệu](../channels/line/README.vi.md) |
+| **WeCom (企业微信)** | ⭐⭐⭐ Nâng cao | Bot nhóm (Webhook), ứng dụng tùy chỉnh (API), AI Bot | [Bot](../channels/wecom/wecom_bot/README.vi.md) / [App](../channels/wecom/wecom_app/README.vi.md) / [AI Bot](../channels/wecom/wecom_aibot/README.vi.md) |
+| **Feishu (飞书)** | ⭐⭐⭐ Nâng cao | Cộng tác doanh nghiệp, nhiều tính năng | [Tài liệu](../channels/feishu/README.vi.md) |
+| **IRC** | ⭐⭐ Trung bình | Máy chủ + cấu hình TLS | - |
+| **OneBot** | ⭐⭐ Trung bình | Tương thích NapCat/Go-CQHTTP, hệ sinh thái cộng đồng | [Tài liệu](../channels/onebot/README.vi.md) |
+| **MaixCam** | ⭐ Dễ | Kênh tích hợp phần cứng cho camera AI Sipeed | [Tài liệu](../channels/maixcam/README.vi.md) |
+| **Pico** | ⭐ Dễ | Kênh giao thức bản địa PicoClaw | |
+
+
+Telegram (Khuyến nghị)
+
+**1. Tạo bot**
+
+* Mở Telegram, tìm `@BotFather`
+* Gửi `/newbot`, làm theo hướng dẫn
+* Sao chép token
+
+**2. Cấu hình**
+
+```json
+{
+ "channels": {
+ "telegram": {
+ "enabled": true,
+ "token": "YOUR_BOT_TOKEN",
+ "allow_from": ["YOUR_USER_ID"]
+ }
+ }
+}
+```
+
+> Lấy user ID của bạn từ `@userinfobot` trên Telegram.
+
+**3. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+**4. Menu lệnh Telegram (tự động đăng ký khi khởi động)**
+
+PicoClaw hiện lưu trữ định nghĩa lệnh trong một registry chung. Khi khởi động, Telegram sẽ tự động đăng ký các lệnh bot được hỗ trợ (ví dụ `/start`, `/help`, `/show`, `/list`) để menu lệnh và hành vi runtime luôn đồng bộ.
+Đăng ký menu lệnh Telegram vẫn là UX khám phá cục bộ của kênh; thực thi lệnh chung được xử lý tập trung trong vòng lặp agent qua commands executor.
+
+Nếu đăng ký lệnh thất bại (lỗi tạm thời mạng/API), kênh vẫn khởi động và PicoClaw thử lại đăng ký trong nền.
+
+
+
+
+Discord
+
+**1. Tạo bot**
+
+* Truy cập
+* Tạo ứng dụng → Bot → Add Bot
+* Sao chép bot token
+
+**2. Bật intents**
+
+* Trong cài đặt Bot, bật **MESSAGE CONTENT INTENT**
+* (Tùy chọn) Bật **SERVER MEMBERS INTENT** nếu bạn muốn sử dụng danh sách cho phép dựa trên dữ liệu thành viên
+
+**3. Lấy User ID**
+* Cài đặt Discord → Nâng cao → bật **Developer Mode**
+* Nhấp chuột phải vào avatar → **Copy User ID**
+
+**4. Cấu hình**
+
+```json
+{
+ "channels": {
+ "discord": {
+ "enabled": true,
+ "token": "YOUR_BOT_TOKEN",
+ "allow_from": ["YOUR_USER_ID"]
+ }
+ }
+}
+```
+
+**5. Mời bot**
+
+* OAuth2 → URL Generator
+* Scopes: `bot`
+* Bot Permissions: `Send Messages`, `Read Message History`
+* Mở URL mời được tạo và thêm bot vào server của bạn
+
+**Tùy chọn: Chế độ kích hoạt nhóm**
+
+Mặc định bot phản hồi tất cả tin nhắn trong kênh server. Để giới hạn phản hồi chỉ khi @mention, thêm:
+
+```json
+{
+ "channels": {
+ "discord": {
+ "group_trigger": { "mention_only": true }
+ }
+ }
+}
+```
+
+Bạn cũng có thể kích hoạt bằng tiền tố từ khóa (ví dụ: `!bot`):
+
+```json
+{
+ "channels": {
+ "discord": {
+ "group_trigger": { "prefixes": ["!bot"] }
+ }
+ }
+}
+```
+
+**6. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+WhatsApp (native qua whatsmeow)
+
+PicoClaw có thể kết nối WhatsApp theo hai cách:
+
+- **Native (khuyến nghị):** In-process sử dụng [whatsmeow](https://github.com/tulir/whatsmeow). Không cần bridge riêng. Đặt `"use_native": true` và để trống `bridge_url`. Lần chạy đầu tiên, quét mã QR bằng WhatsApp (Thiết bị liên kết). Phiên được lưu trong workspace (ví dụ: `workspace/whatsapp/`). Kênh native là **tùy chọn** để giữ binary mặc định nhỏ; build với `-tags whatsapp_native` (ví dụ: `make build-whatsapp-native` hoặc `go build -tags whatsapp_native ./cmd/...`).
+- **Bridge:** Kết nối đến bridge WebSocket bên ngoài. Đặt `bridge_url` (ví dụ: `ws://localhost:3001`) và giữ `use_native` là false.
+
+**Cấu hình (native)**
+
+```json
+{
+ "channels": {
+ "whatsapp": {
+ "enabled": true,
+ "use_native": true,
+ "session_store_path": "",
+ "allow_from": []
+ }
+ }
+}
+```
+
+Nếu `session_store_path` trống, phiên được lưu tại `/whatsapp/`. Chạy `picoclaw gateway`; lần chạy đầu tiên, quét mã QR hiển thị trong terminal bằng WhatsApp → Thiết bị liên kết.
+
+
+
+
+QQ
+
+**Thiết lập nhanh (khuyến nghị)**
+
+QQ Open Platform cung cấp trang thiết lập một chạm cho bot tương thích OpenClaw:
+
+1. Mở [QQ Bot Quick Start](https://q.qq.com/qqbot/openclaw/index.html) và quét mã QR để đăng nhập
+2. Bot được tạo tự động — sao chép **App ID** và **App Secret**
+3. Cấu hình PicoClaw:
+
+```json
+{
+ "channels": {
+ "qq": {
+ "enabled": true,
+ "app_id": "YOUR_APP_ID",
+ "app_secret": "YOUR_APP_SECRET",
+ "allow_from": []
+ }
+ }
+}
+```
+
+4. Chạy `picoclaw gateway` và mở QQ để trò chuyện với bot của bạn
+
+> App Secret chỉ hiển thị một lần. Lưu ngay lập tức — xem lại sẽ buộc phải đặt lại.
+>
+> Bot được tạo qua trang thiết lập nhanh ban đầu chỉ dành cho người tạo và không hỗ trợ chat nhóm. Để bật quyền truy cập nhóm, cấu hình chế độ sandbox trên [QQ Open Platform](https://q.qq.com/).
+
+**Thiết lập thủ công**
+
+Nếu bạn muốn tạo bot thủ công:
+
+* Đăng nhập tại [QQ Open Platform](https://q.qq.com/) để đăng ký làm nhà phát triển
+* Tạo bot QQ — tùy chỉnh avatar và tên
+* Sao chép **App ID** và **App Secret** từ cài đặt bot
+* Cấu hình như trên và chạy `picoclaw gateway`
+
+
+
+
+DingTalk
+
+**1. Tạo bot**
+
+* Truy cập [Open Platform](https://open.dingtalk.com/)
+* Tạo ứng dụng nội bộ
+* Sao chép Client ID và Client Secret
+
+**2. Cấu hình**
+
+```json
+{
+ "channels": {
+ "dingtalk": {
+ "enabled": true,
+ "client_id": "YOUR_CLIENT_ID",
+ "client_secret": "YOUR_CLIENT_SECRET",
+ "allow_from": []
+ }
+ }
+}
+```
+
+> Đặt `allow_from` trống để cho phép tất cả người dùng, hoặc chỉ định DingTalk user ID để giới hạn truy cập.
+
+**3. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+MaixCam
+
+Kênh tích hợp được thiết kế đặc biệt cho phần cứng camera AI Sipeed.
+
+```json
+{
+ "channels": {
+ "maixcam": {
+ "enabled": true
+ }
+ }
+}
+```
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+
+Matrix
+
+**1. Chuẩn bị tài khoản bot**
+
+* Sử dụng homeserver ưa thích (ví dụ: `https://matrix.org` hoặc tự host)
+* Tạo user bot và lấy access token
+
+**2. Cấu hình**
+
+```json
+{
+ "channels": {
+ "matrix": {
+ "enabled": true,
+ "homeserver": "https://matrix.org",
+ "user_id": "@your-bot:matrix.org",
+ "access_token": "YOUR_MATRIX_ACCESS_TOKEN",
+ "allow_from": []
+ }
+ }
+}
+```
+
+**3. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+Để xem đầy đủ các tùy chọn (`device_id`, `join_on_invite`, `group_trigger`, `placeholder`, `reasoning_channel_id`), xem [Hướng Dẫn Cấu Hình Kênh Matrix](../channels/matrix/README.md).
+
+
+
+
+LINE
+
+**1. Tạo Tài Khoản LINE Official**
+
+- Truy cập [LINE Developers Console](https://developers.line.biz/)
+- Tạo provider → Tạo kênh Messaging API
+- Sao chép **Channel Secret** và **Channel Access Token**
+
+**2. Cấu hình**
+
+```json
+{
+ "channels": {
+ "line": {
+ "enabled": true,
+ "channel_secret": "YOUR_CHANNEL_SECRET",
+ "channel_access_token": "YOUR_CHANNEL_ACCESS_TOKEN",
+ "webhook_path": "/webhook/line",
+ "allow_from": []
+ }
+ }
+}
+```
+
+> Webhook LINE được phục vụ trên máy chủ Gateway chung (`gateway.host`:`gateway.port`, mặc định `127.0.0.1:18790`).
+
+**3. Thiết lập Webhook URL**
+
+LINE yêu cầu HTTPS cho webhook. Sử dụng reverse proxy hoặc tunnel:
+
+```bash
+# Ví dụ với ngrok (port mặc định gateway là 18790)
+ngrok http 18790
+```
+
+Sau đó đặt Webhook URL trong LINE Developers Console thành `https://your-domain/webhook/line` và bật **Use webhook**.
+
+**4. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+> Trong chat nhóm, bot chỉ phản hồi khi được @mention. Phản hồi trích dẫn tin nhắn gốc.
+
+
+
+
+WeCom (企业微信)
+
+PicoClaw hỗ trợ ba loại tích hợp WeCom:
+
+**Tùy chọn 1: WeCom Bot (Bot)** - Thiết lập dễ hơn, hỗ trợ chat nhóm
+**Tùy chọn 2: WeCom App (App Tùy chỉnh)** - Nhiều tính năng hơn, nhắn tin chủ động, chỉ chat riêng
+**Tùy chọn 3: WeCom AI Bot (AI Bot)** - AI Bot chính thức, phản hồi streaming, hỗ trợ chat nhóm & riêng
+
+Xem [Hướng Dẫn Cấu Hình WeCom AI Bot](../channels/wecom/wecom_aibot/README.vi.md) để biết hướng dẫn thiết lập chi tiết.
+
+**Thiết Lập Nhanh - WeCom Bot:**
+
+**1. Tạo bot**
+
+* Truy cập Console Quản Trị WeCom → Chat Nhóm → Thêm Bot Nhóm
+* Sao chép URL webhook (định dạng: `https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=xxx`)
+
+**2. Cấu hình**
+
+```json
+{
+ "channels": {
+ "wecom": {
+ "enabled": true,
+ "token": "YOUR_TOKEN",
+ "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
+ "webhook_url": "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=YOUR_KEY",
+ "webhook_path": "/webhook/wecom",
+ "allow_from": []
+ }
+ }
+}
+```
+
+> Webhook WeCom được phục vụ trên máy chủ Gateway chung (`gateway.host`:`gateway.port`, mặc định `127.0.0.1:18790`).
+
+**Thiết Lập Nhanh - WeCom App:**
+
+**1. Tạo ứng dụng**
+
+* Truy cập Console Quản Trị WeCom → Quản Lý App → Tạo App
+* Sao chép **AgentId** và **Secret**
+* Truy cập trang "Công Ty Của Tôi", sao chép **CorpID**
+
+**2. Cấu hình nhận tin nhắn**
+
+* Trong chi tiết App, nhấp "Nhận Tin Nhắn" → "Cấu Hình API"
+* Đặt URL thành `http://your-server:18790/webhook/wecom-app`
+* Tạo **Token** và **EncodingAESKey**
+
+**3. Cấu hình**
+
+```json
+{
+ "channels": {
+ "wecom_app": {
+ "enabled": true,
+ "corp_id": "wwxxxxxxxxxxxxxxxx",
+ "corp_secret": "YOUR_CORP_SECRET",
+ "agent_id": 1000002,
+ "token": "YOUR_TOKEN",
+ "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
+ "webhook_path": "/webhook/wecom-app",
+ "allow_from": []
+ }
+ }
+}
+```
+
+**4. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+> **Lưu ý**: Callback webhook WeCom được phục vụ trên port Gateway (mặc định 18790). Sử dụng reverse proxy cho HTTPS.
+
+**Thiết Lập Nhanh - WeCom AI Bot:**
+
+**1. Tạo AI Bot**
+
+* Truy cập Console Quản Trị WeCom → Quản Lý App → AI Bot
+* Trong cài đặt AI Bot, cấu hình callback URL: `http://your-server:18790/webhook/wecom-aibot`
+* Sao chép **Token** và nhấp "Tạo Ngẫu Nhiên" cho **EncodingAESKey**
+
+**2. Cấu hình**
+
+```json
+{
+ "channels": {
+ "wecom_aibot": {
+ "enabled": true,
+ "token": "YOUR_TOKEN",
+ "encoding_aes_key": "YOUR_43_CHAR_ENCODING_AES_KEY",
+ "webhook_path": "/webhook/wecom-aibot",
+ "allow_from": [],
+ "welcome_message": "Hello! How can I help you?",
+ "processing_message": "⏳ Processing, please wait. The results will be sent shortly."
+ }
+ }
+}
+```
+
+**3. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+> **Lưu ý**: WeCom AI Bot sử dụng giao thức streaming pull — không lo timeout phản hồi. Tác vụ dài (>30 giây) tự động chuyển sang gửi qua `response_url` push.
+
+
+
+
+Feishu (Lark)
+
+PicoClaw kết nối với Feishu qua chế độ WebSocket/SDK — không cần URL webhook công khai hay máy chủ callback.
+
+**1. Tạo ứng dụng**
+
+* Truy cập [Feishu Open Platform](https://open.feishu.cn/) và tạo ứng dụng
+* Trong cài đặt ứng dụng, bật khả năng **Bot**
+* Tạo phiên bản và xuất bản ứng dụng (ứng dụng phải được xuất bản mới có hiệu lực)
+* Sao chép **App ID** (bắt đầu bằng `cli_`) và **App Secret**
+
+**2. Cấu hình**
+
+```json
+{
+ "channels": {
+ "feishu": {
+ "enabled": true,
+ "app_id": "cli_xxx",
+ "app_secret": "YOUR_APP_SECRET",
+ "allow_from": []
+ }
+ }
+}
+```
+
+Tùy chọn: `encrypt_key` và `verification_token` để mã hóa sự kiện (khuyến nghị cho môi trường production).
+
+**3. Chạy và trò chuyện**
+
+```bash
+picoclaw gateway
+```
+
+Mở Feishu, tìm tên bot của bạn và bắt đầu trò chuyện. Bạn cũng có thể thêm bot vào nhóm — sử dụng `group_trigger.mention_only: true` để chỉ phản hồi khi được @mention.
+
+Để xem đầy đủ các tùy chọn, xem [Hướng Dẫn Cấu Hình Kênh Feishu](../channels/feishu/README.vi.md).
+
+
+
+
+Slack
+
+**1. Tạo ứng dụng Slack**
+
+* Truy cập [Slack API](https://api.slack.com/apps) và tạo ứng dụng mới
+* Trong **OAuth & Permissions**, thêm các scope bot: `chat:write`, `app_mentions:read`, `im:history`, `im:read`, `im:write`
+* Cài đặt ứng dụng vào workspace của bạn
+* Sao chép **Bot Token** (`xoxb-...`) và **App-Level Token** (`xapp-...`, bật Socket Mode để lấy token này)
+
+**2. Cấu hình**
+
+```json
+{
+ "channels": {
+ "slack": {
+ "enabled": true,
+ "bot_token": "xoxb-YOUR-BOT-TOKEN",
+ "app_token": "xapp-YOUR-APP-TOKEN",
+ "allow_from": []
+ }
+ }
+}
+```
+
+**3. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+IRC
+
+**1. Cấu hình**
+
+```json
+{
+ "channels": {
+ "irc": {
+ "enabled": true,
+ "server": "irc.libera.chat:6697",
+ "tls": true,
+ "nick": "picoclaw-bot",
+ "channels": ["#your-channel"],
+ "password": "",
+ "allow_from": []
+ }
+ }
+}
+```
+
+Tùy chọn: `nickserv_password` để xác thực NickServ, `sasl_user`/`sasl_password` để xác thực SASL.
+
+**2. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+Bot sẽ kết nối đến máy chủ IRC và tham gia các kênh đã chỉ định.
+
+
+
+
+OneBot (QQ qua giao thức OneBot)
+
+OneBot là giao thức mở cho bot QQ. PicoClaw kết nối với bất kỳ triển khai tương thích OneBot v11 nào (ví dụ: [Lagrange](https://github.com/LagrangeDev/Lagrange.Core), [NapCat](https://github.com/NapNeko/NapCatQQ)) qua WebSocket.
+
+**1. Thiết lập triển khai OneBot**
+
+Cài đặt và chạy framework bot QQ tương thích OneBot v11. Bật máy chủ WebSocket của nó.
+
+**2. Cấu hình**
+
+```json
+{
+ "channels": {
+ "onebot": {
+ "enabled": true,
+ "ws_url": "ws://127.0.0.1:8080",
+ "access_token": "",
+ "allow_from": []
+ }
+ }
+}
+```
+
+| Trường | Mô tả |
+|--------|-------|
+| `ws_url` | URL WebSocket của triển khai OneBot |
+| `access_token` | Token truy cập để xác thực (nếu đã cấu hình trong OneBot) |
+| `reconnect_interval` | Khoảng thời gian kết nối lại tính bằng giây (mặc định: 5) |
+
+**3. Chạy**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+MaixCam
+
+Kênh tích hợp được thiết kế đặc biệt cho phần cứng camera AI Sipeed.
+
+```json
+{
+ "channels": {
+ "maixcam": {
+ "enabled": true
+ }
+ }
+}
+```
+
+```bash
+picoclaw gateway
+```
+
+
diff --git a/docs/vi/configuration.md b/docs/vi/configuration.md
new file mode 100644
index 000000000..a21929359
--- /dev/null
+++ b/docs/vi/configuration.md
@@ -0,0 +1,219 @@
+# ⚙️ Hướng Dẫn Cấu Hình
+
+> Quay lại [README](../../README.vi.md)
+
+## ⚙️ Cấu Hình
+
+File cấu hình: `~/.picoclaw/config.json`
+
+### Biến Môi Trường
+
+Bạn có thể ghi đè các đường dẫn mặc định bằng biến môi trường. Điều này hữu ích cho cài đặt portable, triển khai container, hoặc chạy picoclaw như dịch vụ hệ thống. Các biến này độc lập và kiểm soát các đường dẫn khác nhau.
+
+| Biến | Mô tả | Đường Dẫn Mặc Định |
+|-------------------|-----------------------------------------------------------------------------------------------------------------------------------------|---------------------------|
+| `PICOCLAW_CONFIG` | Ghi đè đường dẫn đến file cấu hình. Chỉ định trực tiếp cho picoclaw file `config.json` nào cần tải, bỏ qua tất cả vị trí khác. | `~/.picoclaw/config.json` |
+| `PICOCLAW_HOME` | Ghi đè thư mục gốc cho dữ liệu picoclaw. Thay đổi vị trí mặc định của `workspace` và các thư mục dữ liệu khác. | `~/.picoclaw` |
+
+**Ví dụ:**
+
+```bash
+# Chạy picoclaw với file cấu hình cụ thể
+# Đường dẫn workspace sẽ được đọc từ trong file cấu hình đó
+PICOCLAW_CONFIG=/etc/picoclaw/production.json picoclaw gateway
+
+# Chạy picoclaw với tất cả dữ liệu lưu tại /opt/picoclaw
+# Cấu hình sẽ được tải từ mặc định ~/.picoclaw/config.json
+# Workspace sẽ được tạo tại /opt/picoclaw/workspace
+PICOCLAW_HOME=/opt/picoclaw picoclaw agent
+
+# Sử dụng cả hai cho thiết lập tùy chỉnh hoàn toàn
+PICOCLAW_HOME=/srv/picoclaw PICOCLAW_CONFIG=/srv/picoclaw/main.json picoclaw gateway
+```
+
+### Bố Cục Workspace
+
+PicoClaw lưu trữ dữ liệu trong workspace đã cấu hình (mặc định: `~/.picoclaw/workspace`):
+
+```
+~/.picoclaw/workspace/
+├── sessions/ # Phiên hội thoại và lịch sử
+├── memory/ # Bộ nhớ dài hạn (MEMORY.md)
+├── state/ # Trạng thái bền vững (kênh cuối, v.v.)
+├── cron/ # Cơ sở dữ liệu tác vụ lên lịch
+├── skills/ # Skill tùy chỉnh
+├── AGENT.md # Hướng dẫn hành vi agent
+├── HEARTBEAT.md # Prompt tác vụ định kỳ (kiểm tra mỗi 30 phút)
+├── IDENTITY.md # Danh tính agent
+├── SOUL.md # Linh hồn agent
+└── USER.md # Tùy chọn người dùng
+```
+
+> **Lưu ý:** Các thay đổi đối với `AGENT.md`, `SOUL.md`, `USER.md` và `memory/MEMORY.md` được tự động phát hiện trong thời gian chạy thông qua theo dõi thời gian sửa đổi file (mtime). **Không cần khởi động lại gateway** sau khi chỉnh sửa các file này — agent sẽ tải nội dung mới vào yêu cầu tiếp theo.
+
+### Nguồn Skill
+
+Mặc định, skill được tải từ:
+
+1. `~/.picoclaw/workspace/skills` (workspace)
+2. `~/.picoclaw/skills` (global)
+3. `<đường-dẫn-nhúng-khi-build>/skills` (tích hợp)
+
+Cho thiết lập nâng cao/test, bạn có thể ghi đè thư mục gốc skill builtin với:
+
+```bash
+export PICOCLAW_BUILTIN_SKILLS=/path/to/skills
+```
+
+### Chính Sách Thực Thi Lệnh Thống Nhất
+
+- Lệnh slash chung được thực thi qua một đường dẫn duy nhất trong `pkg/agent/loop.go` qua `commands.Executor`.
+- Adapter kênh không còn xử lý lệnh chung cục bộ; chúng chuyển tiếp văn bản đầu vào đến đường dẫn bus/agent. Telegram vẫn tự động đăng ký lệnh được hỗ trợ khi khởi động.
+- Lệnh slash không xác định (ví dụ `/foo`) được chuyển sang xử lý LLM bình thường.
+- Lệnh đã đăng ký nhưng không được hỗ trợ trên kênh hiện tại (ví dụ `/show` trên WhatsApp) trả về lỗi rõ ràng cho người dùng và dừng xử lý tiếp.
+
+### 🔒 Sandbox Bảo Mật
+
+PicoClaw chạy trong môi trường sandbox mặc định. Agent chỉ có thể truy cập file và thực thi lệnh trong workspace đã cấu hình.
+
+#### Cấu Hình Mặc Định
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "workspace": "~/.picoclaw/workspace",
+ "restrict_to_workspace": true
+ }
+ }
+}
+```
+
+| Tùy chọn | Mặc định | Mô tả |
+| ----------------------- | ----------------------- | ----------------------------------------- |
+| `workspace` | `~/.picoclaw/workspace` | Thư mục làm việc của agent |
+| `restrict_to_workspace` | `true` | Giới hạn truy cập file/lệnh trong workspace |
+
+#### Công Cụ Được Bảo Vệ
+
+Khi `restrict_to_workspace: true`, các công cụ sau được sandbox:
+
+| Công cụ | Chức năng | Giới hạn |
+| ------------- | ---------------- | -------------------------------------- |
+| `read_file` | Đọc file | Chỉ file trong workspace |
+| `write_file` | Ghi file | Chỉ file trong workspace |
+| `list_dir` | Liệt kê thư mục | Chỉ thư mục trong workspace |
+| `edit_file` | Sửa file | Chỉ file trong workspace |
+| `append_file` | Nối vào file | Chỉ file trong workspace |
+| `exec` | Thực thi lệnh | Đường dẫn lệnh phải trong workspace |
+
+#### Bảo Vệ Exec Bổ Sung
+
+Ngay cả khi `restrict_to_workspace: false`, công cụ `exec` chặn các lệnh nguy hiểm sau:
+
+* `rm -rf`, `del /f`, `rmdir /s` — Xóa hàng loạt
+* `format`, `mkfs`, `diskpart` — Định dạng đĩa
+* `dd if=` — Tạo ảnh đĩa
+* Ghi vào `/dev/sd[a-z]` — Ghi trực tiếp đĩa
+* `shutdown`, `reboot`, `poweroff` — Tắt hệ thống
+* Fork bomb `:(){ :|:& };:`
+
+### Kiểm Soát Truy Cập File
+
+| Config Key | Type | Default | Description |
+|------------|------|---------|-------------|
+| `tools.allow_read_paths` | string[] | `[]` | Additional paths allowed for reading outside workspace |
+| `tools.allow_write_paths` | string[] | `[]` | Additional paths allowed for writing outside workspace |
+
+### Bảo Mật Exec
+
+| Config Key | Type | Default | Description |
+|------------|------|---------|-------------|
+| `tools.exec.allow_remote` | bool | `false` | Allow exec tool from remote channels (Telegram/Discord etc.) |
+| `tools.exec.enable_deny_patterns` | bool | `true` | Enable dangerous command interception |
+| `tools.exec.custom_deny_patterns` | string[] | `[]` | Custom regex patterns to block |
+| `tools.exec.custom_allow_patterns` | string[] | `[]` | Custom regex patterns to allow |
+
+> **Lưu ý Bảo Mật:** Bảo vệ symlink được bật mặc định — tất cả đường dẫn file được giải quyết qua `filepath.EvalSymlinks` trước khi so khớp whitelist, ngăn chặn tấn công thoát qua symlink.
+
+#### Hạn Chế Đã Biết: Tiến Trình Con Từ Công Cụ Build
+
+Guard bảo mật exec chỉ kiểm tra dòng lệnh mà PicoClaw khởi chạy trực tiếp. Nó không kiểm tra đệ quy các tiến trình con được tạo bởi công cụ phát triển được phép như `make`, `go run`, `cargo`, `npm run`, hoặc script build tùy chỉnh.
+
+Điều này có nghĩa là lệnh cấp cao nhất vẫn có thể biên dịch hoặc khởi chạy binary khác sau khi vượt qua kiểm tra guard ban đầu. Trong thực tế, hãy coi script build, Makefile, script package, và binary được tạo như mã thực thi cần cùng mức độ review như lệnh shell trực tiếp.
+
+Cho môi trường rủi ro cao hơn:
+
+* Review script build trước khi thực thi.
+* Ưu tiên phê duyệt/review thủ công cho quy trình biên dịch và chạy.
+* Chạy PicoClaw trong container hoặc VM nếu bạn cần cách ly mạnh hơn guard tích hợp.
+
+#### Ví Dụ Lỗi
+
+```
+[ERROR] tool: Tool execution failed
+{tool=exec, error=Command blocked by safety guard (path outside working dir)}
+```
+
+```
+[ERROR] tool: Tool execution failed
+{tool=exec, error=Command blocked by safety guard (dangerous pattern detected)}
+```
+
+#### Tắt Giới Hạn (Rủi Ro Bảo Mật)
+
+Nếu bạn cần agent truy cập đường dẫn ngoài workspace:
+
+**Phương pháp 1: File cấu hình**
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "restrict_to_workspace": false
+ }
+ }
+}
+```
+
+**Phương pháp 2: Biến môi trường**
+
+```bash
+export PICOCLAW_AGENTS_DEFAULTS_RESTRICT_TO_WORKSPACE=false
+```
+
+> ⚠️ **Cảnh báo**: Tắt giới hạn này cho phép agent truy cập bất kỳ đường dẫn nào trên hệ thống. Chỉ sử dụng cẩn thận trong môi trường được kiểm soát.
+
+#### Tính Nhất Quán Ranh Giới Bảo Mật
+
+Cài đặt `restrict_to_workspace` áp dụng nhất quán trên tất cả đường dẫn thực thi:
+
+| Đường Dẫn Thực Thi | Ranh Giới Bảo Mật |
+| -------------------- | ---------------------------- |
+| Main Agent | `restrict_to_workspace` ✅ |
+| Subagent / Spawn | Kế thừa cùng giới hạn ✅ |
+| Heartbeat tasks | Kế thừa cùng giới hạn ✅ |
+
+Tất cả đường dẫn chia sẻ cùng giới hạn workspace — không có cách nào vượt qua ranh giới bảo mật qua subagent hoặc tác vụ lên lịch.
+
+### Heartbeat (Tác Vụ Định Kỳ)
+
+PicoClaw có thể thực hiện tác vụ định kỳ tự động. Tạo file `HEARTBEAT.md` trong workspace:
+
+```markdown
+# Tác Vụ Định Kỳ
+
+- Kiểm tra email cho tin nhắn quan trọng
+- Xem lịch cho sự kiện sắp tới
+- Kiểm tra dự báo thời tiết
+```
+
+Agent sẽ đọc file này mỗi 30 phút (có thể cấu hình) và thực thi các tác vụ sử dụng công cụ có sẵn.
+
+#### Tác Vụ Bất Đồng Bộ Với Spawn
+
+Cho tác vụ chạy lâu (tìm kiếm web, gọi API), sử dụng công cụ `spawn` để tạo **subagent**:
+
+```markdown
+# Tác Vụ Định Kỳ
+```
diff --git a/docs/vi/credential_encryption.md b/docs/vi/credential_encryption.md
new file mode 100644
index 000000000..9ba24588b
--- /dev/null
+++ b/docs/vi/credential_encryption.md
@@ -0,0 +1,159 @@
+> Quay lại [README](../../README.vi.md)
+
+# Mã hóa Thông tin Xác thực
+
+PicoClaw hỗ trợ mã hóa các giá trị `api_key` trong các mục cấu hình `model_list`.
+Các khóa đã mã hóa được lưu trữ dưới dạng chuỗi `enc://` và được giải mã tự động khi khởi động.
+
+---
+
+## Bắt đầu Nhanh
+
+**1. Đặt cụm mật khẩu**
+
+```bash
+export PICOCLAW_KEY_PASSPHRASE="your-passphrase"
+```
+
+**2. Mã hóa khóa API**
+
+Chạy `picoclaw onboard` — nó yêu cầu nhập cụm mật khẩu và tạo khóa SSH,
+sau đó tự động mã hóa lại tất cả các mục `api_key` dạng văn bản thuần trong cấu hình
+ở lần gọi `SaveConfig` tiếp theo. Giá trị `enc://` kết quả sẽ có dạng:
+
+```
+enc://AAAA...base64...
+```
+
+**3. Dán kết quả vào cấu hình**
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "gpt-4o",
+ "model": "openai/gpt-4o",
+ "api_key": "enc://AAAA...base64...",
+ "api_base": "https://api.openai.com/v1"
+ }
+ ]
+}
+```
+
+---
+
+## Các Định dạng `api_key` được Hỗ trợ
+
+| Định dạng | Ví dụ | Hành vi |
+|-----------|-------|---------|
+| Văn bản thuần | `sk-abc123` | Sử dụng nguyên trạng |
+| Tham chiếu tệp | `file://openai.key` | Nội dung được đọc từ cùng thư mục với tệp cấu hình |
+| Đã mã hóa | `enc://` | Giải mã khi khởi động bằng `PICOCLAW_KEY_PASSPHRASE` |
+| Trống | `""` | Truyền qua không thay đổi (dùng với `auth_method: oauth`) |
+
+---
+
+## Thiết kế Mật mã
+
+### Dẫn xuất Khóa
+
+Mã hóa sử dụng **HKDF-SHA256** với khóa riêng SSH làm yếu tố thứ hai.
+
+```
+sshHash = SHA256(ssh_private_key_file_bytes)
+ikm = HMAC-SHA256(key=sshHash, message=passphrase)
+aes_key = HKDF-SHA256(ikm, salt, info="picoclaw-credential-v1", 32 bytes)
+```
+
+### Mã hóa
+
+```
+AES-256-GCM(key=aes_key, nonce=random[12], plaintext=api_key)
+```
+
+### Định dạng Truyền tải
+
+```
+enc://
+```
+
+| Trường | Kích thước | Mô tả |
+|--------|-----------|-------|
+| `salt` | 16 byte | Ngẫu nhiên mỗi lần mã hóa; đưa vào HKDF |
+| `nonce` | 12 byte | Ngẫu nhiên mỗi lần mã hóa; IV của AES-GCM |
+| `ciphertext` | thay đổi | Bản mã AES-256-GCM + thẻ xác thực 16 byte |
+
+Thẻ xác thực GCM được tự động nối vào bản mã. Bất kỳ sự giả mạo nào đều khiến giải mã thất bại với lỗi thay vì trả về văn bản thuần bị hỏng.
+
+### Hiệu suất
+
+| Thao tác | Thời gian (ARM Cortex-A) |
+|----------|--------------------------|
+| Dẫn xuất khóa (HKDF) | < 1 ms |
+| Giải mã AES-256-GCM | < 1 ms |
+| **Tổng chi phí khởi động** | **< 2 ms mỗi khóa** |
+
+---
+
+## Bảo mật Hai Yếu tố với Khóa SSH
+
+Khi khóa riêng SSH được cung cấp, việc phá vỡ mã hóa yêu cầu **cả hai**:
+
+1. **Cụm mật khẩu** (`PICOCLAW_KEY_PASSPHRASE`)
+2. **Tệp khóa riêng SSH**
+
+Điều này có nghĩa là chỉ rò rỉ tệp cấu hình không đủ để khôi phục khóa API, ngay cả khi cụm mật khẩu yếu. Khóa SSH đóng góp 256 bit entropy (Ed25519) bất kể độ mạnh của cụm mật khẩu.
+
+### Mô hình Mối đe dọa
+
+| Kẻ tấn công có | Có thể giải mã? |
+|----------------|-----------------|
+| Chỉ tệp cấu hình | Không — cần cụm mật khẩu + khóa SSH |
+| Chỉ khóa SSH | Không — cần cụm mật khẩu |
+| Chỉ cụm mật khẩu | Không — cần khóa SSH |
+| Tệp cấu hình + khóa SSH + cụm mật khẩu | Có — xâm phạm hoàn toàn |
+
+---
+
+## Biến Môi trường
+
+| Biến | Bắt buộc | Mô tả |
+|------|----------|-------|
+| `PICOCLAW_KEY_PASSPHRASE` | Có (cho `enc://`) | Cụm mật khẩu dùng để dẫn xuất khóa |
+| `PICOCLAW_SSH_KEY_PATH` | Không | Đường dẫn đến khóa riêng SSH. Nếu không đặt, tự động phát hiện từ `~/.ssh/picoclaw_ed25519.key` |
+
+### Tự động Phát hiện Khóa SSH
+
+Nếu `PICOCLAW_SSH_KEY_PATH` không được đặt, PicoClaw tìm khóa chuyên dụng:
+
+```
+~/.ssh/picoclaw_ed25519.key
+```
+
+Tệp chuyên dụng này tránh xung đột với các khóa SSH hiện có của người dùng.
+Chạy `picoclaw onboard` để tạo tự động.
+
+`os.UserHomeDir()` được sử dụng để phân giải thư mục home đa nền tảng (đọc `USERPROFILE` trên Windows, `HOME` trên Unix/macOS).
+
+> **Lưu ý:** Tệp khóa SSH là bắt buộc cho mã hóa thông tin xác thực. Nếu không tìm thấy khóa và `PICOCLAW_SSH_KEY_PATH` không được đặt, mã hóa/giải mã sẽ thất bại. Chạy `picoclaw onboard` để tạo khóa tự động.
+
+---
+
+## Di chuyển
+
+Vì tài liệu bí mật duy nhất là `PICOCLAW_KEY_PASSPHRASE` và tệp khóa riêng SSH, việc di chuyển rất đơn giản:
+
+1. Sao chép tệp cấu hình sang máy mới.
+2. Đặt `PICOCLAW_KEY_PASSPHRASE` với cùng giá trị.
+3. Sao chép tệp khóa riêng SSH đến cùng đường dẫn (hoặc đặt `PICOCLAW_SSH_KEY_PATH` đến vị trí mới).
+
+Không cần mã hóa lại.
+
+---
+
+## Lưu ý về Bảo mật
+
+- **Cả cụm mật khẩu và khóa SSH đều bắt buộc.** Khóa SSH đóng vai trò yếu tố thứ hai — không có nó, mã hóa/giải mã sẽ thất bại. Chạy `picoclaw onboard` để tạo khóa nếu chưa tồn tại.
+- **Khóa SSH chỉ đọc khi chạy.** PicoClaw không bao giờ ghi hoặc sửa đổi tệp khóa SSH.
+- **Khóa văn bản thuần vẫn được hỗ trợ.** Các cấu hình hiện có không dùng `enc://` không bị ảnh hưởng.
+- **Định dạng `enc://` được quản lý phiên bản** thông qua trường `info` của HKDF (`picoclaw-credential-v1`), cho phép nâng cấp thuật toán trong tương lai mà không làm hỏng các giá trị đã mã hóa hiện có.
diff --git a/docs/vi/debug.md b/docs/vi/debug.md
new file mode 100644
index 000000000..69583d486
--- /dev/null
+++ b/docs/vi/debug.md
@@ -0,0 +1,36 @@
+# Gỡ lỗi PicoClaw
+
+> Quay lại [README](../../README.vi.md)
+
+PicoClaw thực hiện nhiều tương tác phức tạp ở hậu trường cho mỗi yêu cầu nhận được — từ định tuyến tin nhắn và đánh giá độ phức tạp, đến thực thi công cụ và thích ứng với lỗi mô hình. Khả năng xem chính xác những gì đang xảy ra là rất quan trọng, không chỉ để khắc phục các sự cố tiềm ẩn, mà còn để thực sự hiểu cách agent hoạt động.
+
+## Khởi động PicoClaw ở chế độ gỡ lỗi
+
+Để nhận thông tin chi tiết về những gì agent đang thực hiện (yêu cầu LLM, lệnh gọi công cụ, định tuyến tin nhắn), bạn có thể khởi động gateway PicoClaw với cờ gỡ lỗi:
+
+```bash
+picoclaw gateway --debug
+# or
+picoclaw gateway -d
+```
+
+Ở chế độ này, hệ thống sẽ định dạng log chi tiết và hiển thị bản xem trước của prompt hệ thống và kết quả thực thi công cụ.
+
+## Tắt cắt ngắn log (log đầy đủ)
+
+Theo mặc định, PicoClaw cắt ngắn các chuỗi rất dài (như *Prompt Hệ thống* hoặc kết quả JSON lớn) trong log gỡ lỗi để giữ cho console dễ đọc.
+
+Nếu bạn cần kiểm tra đầu ra đầy đủ của một lệnh hoặc payload chính xác được gửi đến mô hình LLM, bạn có thể sử dụng cờ `--no-truncate`.
+
+**Lưu ý:** Cờ này *chỉ* hoạt động khi kết hợp với chế độ `--debug`.
+
+```bash
+picoclaw gateway --debug --no-truncate
+
+```
+
+Khi cờ này được kích hoạt, chức năng cắt ngắn toàn cục sẽ bị vô hiệu hóa. Điều này cực kỳ hữu ích để:
+
+* Xác minh cú pháp chính xác của các tin nhắn được gửi đến nhà cung cấp.
+* Đọc đầu ra đầy đủ của các công cụ như `exec`, `web_fetch` hoặc `read_file`.
+* Gỡ lỗi lịch sử phiên được lưu trong bộ nhớ.
diff --git a/docs/vi/docker.md b/docs/vi/docker.md
new file mode 100644
index 000000000..eddc20a75
--- /dev/null
+++ b/docs/vi/docker.md
@@ -0,0 +1,167 @@
+# 🐳 Docker và Bắt Đầu Nhanh
+
+> Quay lại [README](../../README.vi.md)
+
+## 🐳 Docker Compose
+
+Bạn cũng có thể chạy PicoClaw bằng Docker Compose mà không cần cài đặt gì trên máy.
+
+```bash
+# 1. Clone repo này
+git clone https://github.com/sipeed/picoclaw.git
+cd picoclaw
+
+# 2. Lần chạy đầu tiên — tự động tạo docker/data/config.json rồi thoát
+# (chỉ kích hoạt khi cả config.json và workspace/ đều không tồn tại)
+docker compose -f docker/docker-compose.yml --profile gateway up
+# Container hiển thị "First-run setup complete." và dừng lại.
+
+# 3. Cấu hình API key của bạn
+vim docker/data/config.json # Set provider API keys, bot tokens, etc.
+
+# 4. Khởi động
+docker compose -f docker/docker-compose.yml --profile gateway up -d
+```
+
+> [!TIP]
+> **Người dùng Docker**: Mặc định, Gateway lắng nghe trên `127.0.0.1`, không thể truy cập từ host. Nếu bạn cần truy cập các health endpoint hoặc mở port, hãy đặt `PICOCLAW_GATEWAY_HOST=0.0.0.0` trong môi trường hoặc cập nhật `config.json`.
+
+```bash
+# 5. Kiểm tra log
+docker compose -f docker/docker-compose.yml logs -f picoclaw-gateway
+
+# 6. Dừng
+docker compose -f docker/docker-compose.yml --profile gateway down
+```
+
+### Chế Độ Launcher (Web Console)
+
+Image `launcher` bao gồm cả ba binary (`picoclaw`, `picoclaw-launcher`, `picoclaw-launcher-tui`) và khởi động web console mặc định, cung cấp giao diện trình duyệt để cấu hình và chat.
+
+```bash
+docker compose -f docker/docker-compose.yml --profile launcher up -d
+```
+
+Mở http://localhost:18800 trong trình duyệt. Launcher tự động quản lý tiến trình gateway.
+
+> [!WARNING]
+> Web console chưa hỗ trợ xác thực. Tránh để lộ ra internet công cộng.
+
+### Chế Độ Agent (One-shot)
+
+```bash
+# Đặt câu hỏi
+docker compose -f docker/docker-compose.yml run --rm picoclaw-agent -m "What is 2+2?"
+
+# Chế độ tương tác
+docker compose -f docker/docker-compose.yml run --rm picoclaw-agent
+```
+
+### Cập Nhật
+
+```bash
+docker compose -f docker/docker-compose.yml pull
+docker compose -f docker/docker-compose.yml --profile gateway up -d
+```
+
+### 🚀 Bắt Đầu Nhanh
+
+> [!TIP]
+> Cấu hình API Key trong `~/.picoclaw/config.json`. Lấy API Key: [Volcengine (CodingPlan)](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) (LLM) · [OpenRouter](https://openrouter.ai/keys) (LLM) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) (LLM). Tìm kiếm web là tùy chọn — lấy miễn phí [Tavily API](https://tavily.com) (1000 truy vấn miễn phí/tháng) hoặc [Brave Search API](https://brave.com/search/api) (2000 truy vấn miễn phí/tháng).
+
+**1. Khởi tạo**
+
+```bash
+picoclaw onboard
+```
+
+**2. Cấu hình** (`~/.picoclaw/config.json`)
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "workspace": "~/.picoclaw/workspace",
+ "model_name": "gpt-5.4",
+ "max_tokens": 8192,
+ "temperature": 0.7,
+ "max_tool_iterations": 20
+ }
+ },
+ "model_list": [
+ {
+ "model_name": "ark-code-latest",
+ "model": "volcengine/ark-code-latest",
+ "api_key": "sk-your-api-key",
+ "api_base":"https://ark.cn-beijing.volces.com/api/coding/v3"
+ },
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_key": "your-api-key",
+ "request_timeout": 300
+ },
+ {
+ "model_name": "claude-sonnet-4.6",
+ "model": "anthropic/claude-sonnet-4.6",
+ "api_key": "your-anthropic-key"
+ }
+ ],
+ "tools": {
+ "web": {
+ "enabled": true,
+ "fetch_limit_bytes": 10485760,
+ "format": "plaintext",
+ "brave": {
+ "enabled": false,
+ "api_key": "YOUR_BRAVE_API_KEY",
+ "max_results": 5
+ },
+ "tavily": {
+ "enabled": false,
+ "api_key": "YOUR_TAVILY_API_KEY",
+ "max_results": 5
+ },
+ "duckduckgo": {
+ "enabled": true,
+ "max_results": 5
+ },
+ "perplexity": {
+ "enabled": false,
+ "api_key": "YOUR_PERPLEXITY_API_KEY",
+ "max_results": 5
+ },
+ "searxng": {
+ "enabled": false,
+ "base_url": "http://your-searxng-instance:8888",
+ "max_results": 5
+ }
+ }
+ }
+}
+```
+
+> **Mới**: Định dạng cấu hình `model_list` cho phép thêm provider mà không cần thay đổi code. Xem [Cấu Hình Mô Hình](#cấu-hình-mô-hình-model_list) để biết chi tiết.
+> `request_timeout` là tùy chọn và tính bằng giây. Nếu bỏ qua hoặc đặt `<= 0`, PicoClaw sử dụng timeout mặc định (120s).
+
+**3. Lấy API Key**
+
+* **Nhà cung cấp LLM**: [OpenRouter](https://openrouter.ai/keys) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) · [Anthropic](https://console.anthropic.com) · [OpenAI](https://platform.openai.com) · [Gemini](https://aistudio.google.com/api-keys)
+* **Tìm kiếm Web** (tùy chọn):
+ * [Brave Search](https://brave.com/search/api) - Trả phí ($5/1000 truy vấn, ~$5-6/tháng)
+ * [Perplexity](https://www.perplexity.ai) - Tìm kiếm bằng AI với giao diện chat
+ * [SearXNG](https://github.com/searxng/searxng) - Công cụ tìm kiếm tổng hợp tự host (miễn phí, không cần API key)
+ * [Tavily](https://tavily.com) - Tối ưu cho AI Agent (1000 yêu cầu/tháng)
+ * DuckDuckGo - Fallback tích hợp (không cần API key)
+
+> **Lưu ý**: Xem `config.example.json` để có mẫu cấu hình đầy đủ.
+
+**4. Chat**
+
+```bash
+picoclaw agent -m "What is 2+2?"
+```
+
+Vậy là xong! Bạn có một trợ lý AI hoạt động trong 2 phút.
+
+---
diff --git a/docs/vi/hardware-compatibility.md b/docs/vi/hardware-compatibility.md
new file mode 100644
index 000000000..8315c049e
--- /dev/null
+++ b/docs/vi/hardware-compatibility.md
@@ -0,0 +1,152 @@
+> Quay lại [README](../../README.vi.md)
+
+# 🖥️ PicoClaw Danh sách tương thích phần cứng
+
+PicoClaw chạy được trên hầu hết mọi thiết bị Linux. Trang này ghi nhận các chip, sản phẩm và bo mạch phát triển đã được xác minh.
+
+**Phần cứng của bạn chưa có trong danh sách?** Gửi PR để thêm vào! Các nhà sản xuất phần cứng được hoan nghênh đóng góp và đồng quảng bá.
+
+---
+
+## 1. Hỗ trợ chip đã xác minh
+
+### x86
+
+| Nhà sản xuất | Chip | Ghi chú |
+|--------------|------|---------|
+| Intel | Any x86 CPU (i386+) | Tất cả bộ xử lý desktop/server/laptop |
+| AMD | Any x86 CPU | Tất cả bộ xử lý desktop/server/laptop |
+
+### ARM
+
+| Kiến trúc phụ | Chip tiêu biểu | Ghi chú |
+|----------------|----------------|---------|
+| ARMv6 | [BCM2835](https://www.raspberrypi.com/documentation/computers/processors.html#bcm2835) (Raspberry Pi 1/Zero) | Đơn nhân ARM1176JZF-S |
+| ARMv7 | [Allwinner V3s](https://linux-sunxi.org/V3s) | Đơn nhân Cortex-A7, dùng trong LicheePi Zero |
+| ARM64 | [Allwinner H618](https://linux-sunxi.org/H618) | Bốn nhân Cortex-A53, dùng trong Orange Pi Zero 3 |
+| ARM64 | [BCM2711](https://www.raspberrypi.com/documentation/computers/processors.html#bcm2711) (Raspberry Pi 4) | Bốn nhân Cortex-A72 |
+| ARM64 | [BCM2712](https://www.raspberrypi.com/documentation/computers/processors.html#bcm2712) (Raspberry Pi 5) | Bốn nhân Cortex-A76 |
+| ARM64 | [AX630C](https://www.axera-tech.com/) (爱芯元智) | Hai nhân Cortex-A53 + NPU, dùng trong NanoKVM-Pro / MaixCAM2 |
+
+### RISC-V (riscv64)
+
+| Nhà sản xuất | Chip | Lõi | Ghi chú |
+|--------------|------|-----|---------|
+| [SOPHGO (算能)](https://www.sophgo.com/) | SG2002 | C906 @ 1GHz | 256MB DDR3 tích hợp, dùng trong LicheeRV-Nano / NanoKVM / MaixCAM |
+| [Allwinner (全志)](https://www.allwinnertech.com/) | V861 | Dual C907 | 128MB DDR3L tích hợp, 1 TOPS NPU, camera AI 4K SiP |
+| [Allwinner (全志)](https://www.allwinnertech.com/) | V881 | C907 | Dòng camera AI RISC-V |
+| [Arterytek (匠芯创)](https://www.arterytek.com/) | D213 | RISC-V | Dùng trong HaaS506-LD1 RTU công nghiệp |
+| [SpacemiT (进迭)](https://www.spacemit.com/) | K1 | 8x X60 @ 1.8GHz | Dùng trong Milk-V Jupiter, BananaPi BPI-F3 |
+| [SpacemiT (进迭)](https://www.spacemit.com/) | K3 | 8x X100 @ 2.5GHz | Tuân thủ RVA23, RVV 1024-bit, suy luận AI FP8 |
+| [Zhihe (知合)](https://www.zhihe-tech.com/) | A210 | High-perf RISC-V | 8 lõi, 16MB cache L3, cấp desktop |
+| [Canaan (嘉楠)](https://www.canaan-creative.com/) | K230 | Dual C908 @ 1.6GHz | 6 TOPS KPU, dùng trong CanMV-K230 |
+
+### MIPS
+
+| Nhà sản xuất | Chip | Ghi chú |
+|--------------|------|---------|
+| MediaTek | [MT7620](https://www.mediatek.com/products/home-networking/mt7620) | MIPS24KEc @ 580MHz, dùng trong nhiều router OpenWrt (vd. Xiaomi Router 3G) |
+
+### LoongArch (loong64)
+
+| Nhà sản xuất | Chip | Ghi chú |
+|--------------|------|---------|
+| [Loongson (龙芯)](https://www.loongson.cn/) | 3A5000 | Bốn nhân LA464 @ 2.5GHz, desktop/máy trạm |
+| [Loongson (龙芯)](https://www.loongson.cn/) | 3A6000 | Bốn nhân 4C/8T @ 2.5GHz, IPC tương đương Intel thế hệ 10 |
+| [Loongson (龙芯)](https://www.loongson.cn/) | 2K1000LA | Hai nhân @ 1GHz, ứng dụng công nghiệp/IoT |
+
+---
+
+## 2. Sản phẩm đã xác minh (theo ngày phát hành)
+
+Sản phẩm tiêu dùng, router và thiết bị công nghiệp đã được kiểm thử với PicoClaw.
+
+| Năm | Sản phẩm | Kiến trúc | SoC | RAM | Danh mục |
+|-----|----------|-----------|-----|-----|----------|
+| 2009 | Nokia N900 | ARM (A8) | OMAP3430 | 256MB | Điện thoại thông minh |
+| 2012 | Samsung Galaxy Note 10.1 (N8000) | ARM (A9) | Exynos 4412 | 2GB | Máy tính bảng |
+| 2016 | Xiaomi Router 3G (小米路由器3G) | MIPS | MT7620 | 256MB | Router (OpenWrt) |
+| 2018 | Phicomm N1 (斐讯N1) | ARM64 (A53) | S905D | 2GB | TV Box / Máy chủ gia đình |
+| 2019 | Xiaomi AI Speaker (小爱音箱) | ARM64 (A53) | — | 256MB | Loa thông minh |
+| 2024 | [NanoKVM](https://wiki.sipeed.com/hardware/en/kvm/NanoKVM/introduction.html) | RISC-V | SG2002 | 256MB | IP-KVM |
+| 2025 | HaaS506-LD1 | RISC-V | D213 | 128MB | RTU công nghiệp |
+| 2025 | [NanoKVM-Pro](https://wiki.sipeed.com/hardware/en/kvm/NanoKVM_Pro/introduction.html) | ARM64 (A53) | AX630C | 1GB | IP-KVM Pro |
+| 2026 | [MaixCAM2](https://wiki.sipeed.com/hardware/en/maixcam/index.html) | ARM64 (A53) | AX630C | 1/4GB | Camera AI 4K |
+
+---
+
+## 3. Bo mạch phát triển đã xác minh (theo ngày phát hành)
+
+| Năm | Bo mạch | Kiến trúc | SoC | RAM | Liên kết mua |
+|-----|---------|-----------|-----|-----|--------------|
+| 2012 | [Raspberry Pi 1 Model B](https://www.raspberrypi.com/products/) | ARMv6 | BCM2835 | 512MB | — |
+| 2015 | [Raspberry Pi 2 Model B](https://www.raspberrypi.com/products/raspberry-pi-2-model-b/) | ARMv7 (A7) | BCM2836 | 1GB | — |
+| 2015 | [Raspberry Pi Zero](https://www.raspberrypi.com/products/raspberry-pi-zero/) | ARMv6 | BCM2835 | 512MB | — |
+| 2016 | [Raspberry Pi 3 Model B](https://www.raspberrypi.com/products/raspberry-pi-3-model-b/) | ARM64 (A53) | BCM2837 | 1GB | — |
+| 2017 | [LicheePi Zero](https://wiki.sipeed.com/hardware/en/lichee/Zero/Zero.html) | ARMv7 (A7) | Allwinner V3s | 64MB | [Sipeed](https://sipeed.com/) |
+| 2019 | [Raspberry Pi 4 Model B](https://www.raspberrypi.com/products/raspberry-pi-4-model-b/) | ARM64 (A72) | BCM2711 | 1~8GB | [RPi](https://www.raspberrypi.com/) |
+| 2023 | [Raspberry Pi 5](https://www.raspberrypi.com/products/raspberry-pi-5/) | ARM64 (A76) | BCM2712 | 2~8GB | [RPi](https://www.raspberrypi.com/) |
+| 2024 | [LicheeRV-Nano](https://wiki.sipeed.com/hardware/en/lichee/RV_Nano/1_intro.html) | RISC-V | SG2002 | 256MB | [AliExpress](https://www.aliexpress.com/item/1005006519668532.html) |
+| 2024 | [MaixCAM-Pro](https://wiki.sipeed.com/hardware/en/maixcam/index.html) | RISC-V | SG2002 | 256MB | [Sipeed](https://sipeed.com/) |
+| 2024 | [Milk-V Duo 64M](https://milkv.io/docs/duo/getting-started/duo) | RISC-V | CV1800B | 64MB | [Milk-V](https://milkv.io/) |
+| 2024 | [CanMV-K230](https://developer.canaan-creative.com/k230_canmv/en/main/) | RISC-V | K230 | 512MB | [Canaan](https://www.canaan-creative.com/) |
+
+---
+
+## 4. Cũng hoạt động trên
+
+### Điện thoại Android (qua Termux)
+
+Bất kỳ điện thoại Android ARM64 nào (2015+) với 1GB+ RAM. Cài đặt [Termux](https://github.com/termux/termux-app), sử dụng `proot` để chạy PicoClaw.
+
+> Xem [README: Chạy trên điện thoại Android cũ](../../README.vi.md#-run-on-old-android-phones) để biết hướng dẫn cài đặt.
+
+### Desktop / Máy chủ / Đám mây
+
+| Nền tảng | Ghi chú |
+|----------|---------|
+| x86_64 Linux | Binary gốc, không phụ thuộc |
+| x86_64 Windows | Binary gốc |
+| macOS (Intel / Apple Silicon) | Binary gốc |
+| Docker (any platform) | `docker compose` một dòng lệnh, xem [Hướng dẫn Docker](docker.md) |
+| OpenWrt routers | Bản dựng MIPS/ARM, yêu cầu >32MB RAM trống |
+| FreeBSD / NetBSD | Có bản dựng x86_64 và arm64 |
+
+---
+
+## 5. Yêu cầu tối thiểu
+
+| Tài nguyên | Tối thiểu | Khuyến nghị |
+|------------|-----------|-------------|
+| RAM | 10MB trống | 32MB+ trống |
+| Lưu trữ | 20MB (binary) | 50MB+ (với workspace) |
+| CPU | Bất kỳ (đơn nhân 0.6GHz+) | — |
+| OS | Linux (kernel 3.x+) | Linux 5.x+ |
+| Mạng | Bắt buộc (cho các lệnh gọi API LLM) | Ethernet hoặc WiFi |
+
+---
+
+## 6. Cách kiểm thử và đóng góp
+
+```bash
+# 1. Tải xuống cho kiến trúc của bạn
+wget https://github.com/sipeed/picoclaw/releases/latest/download/picoclaw_Linux_arm64.tar.gz
+tar xzf picoclaw_Linux_arm64.tar.gz
+
+# 2. Khởi tạo
+./picoclaw onboard
+
+# 3. Kiểm thử
+./picoclaw agent -m "Hello, what board am I running on?"
+```
+
+Các bản dựng có sẵn: `linux-amd64`, `linux-arm64`, `linux-arm`, `linux-riscv64`, `linux-loong64`, `linux-mipsle`
+
+### Thêm phần cứng của bạn
+
+1. Fork kho lưu trữ này
+2. Thêm chip / sản phẩm / bo mạch của bạn vào bảng tương ứng
+3. Bao gồm: tên, kiến trúc, SoC, RAM, năm và liên kết nếu có
+4. Gửi PR
+
+Nhà sản xuất phần cứng: muốn thêm hỗ trợ chính thức hoặc đồng quảng bá? Mở issue hoặc liên hệ qua [Discord](https://discord.gg/V4sAZ9XWpN).
diff --git a/docs/vi/providers.md b/docs/vi/providers.md
new file mode 100644
index 000000000..09b51c56b
--- /dev/null
+++ b/docs/vi/providers.md
@@ -0,0 +1,433 @@
+# 🔌 Nhà Cung Cấp và Cấu Hình Mô Hình
+
+> Quay lại [README](../../README.vi.md)
+
+### Nhà Cung Cấp
+
+> [!NOTE]
+> Groq cung cấp chuyển đổi giọng nói miễn phí qua Whisper. Nếu được cấu hình, tin nhắn âm thanh từ bất kỳ kênh nào sẽ được tự động chuyển đổi ở cấp agent.
+
+| Provider | Purpose | Get API Key |
+| ------------ | --------------------------------------- | ------------------------------------------------------------ |
+| `gemini` | LLM (Gemini direct) | [aistudio.google.com](https://aistudio.google.com) |
+| `zhipu` | LLM (Zhipu direct) | [bigmodel.cn](https://bigmodel.cn) |
+| `volcengine` | LLM(Volcengine direct) | [volcengine.com](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) |
+| `openrouter` | LLM (recommended, access to all models) | [openrouter.ai](https://openrouter.ai) |
+| `anthropic` | LLM (Claude direct) | [console.anthropic.com](https://console.anthropic.com) |
+| `openai` | LLM (GPT direct) | [platform.openai.com](https://platform.openai.com) |
+| `deepseek` | LLM (DeepSeek direct) | [platform.deepseek.com](https://platform.deepseek.com) |
+| `qwen` | LLM (Qwen direct) | [dashscope.console.aliyun.com](https://dashscope.console.aliyun.com) |
+| `groq` | LLM + **Voice transcription** (Whisper) | [console.groq.com](https://console.groq.com) |
+| `cerebras` | LLM (Cerebras direct) | [cerebras.ai](https://cerebras.ai) |
+| `vivgrid` | LLM (Vivgrid direct) | [vivgrid.com](https://vivgrid.com) |
+| `moonshot` | LLM (Kimi/Moonshot direct) | [platform.moonshot.cn](https://platform.moonshot.cn) |
+| `minimax` | LLM (Minimax direct) | [platform.minimaxi.com](https://platform.minimaxi.com) |
+| `avian` | LLM (Avian direct) | [avian.io](https://avian.io) |
+| `mistral` | LLM (Mistral direct) | [console.mistral.ai](https://console.mistral.ai) |
+| `longcat` | LLM (Longcat direct) | [longcat.ai](https://longcat.ai) |
+| `modelscope` | LLM (ModelScope direct) | [modelscope.cn](https://modelscope.cn) |
+
+### Cấu Hình Mô Hình (model_list)
+
+> **Có gì mới?** PicoClaw hiện sử dụng cách tiếp cận cấu hình **tập trung vào mô hình**. Chỉ cần chỉ định định dạng `vendor/model` (ví dụ: `zhipu/glm-4.7`) để thêm provider mới — **không cần thay đổi code!**
+
+Thiết kế này cũng cho phép **hỗ trợ đa agent** với lựa chọn provider linh hoạt:
+
+- **Agent khác nhau, provider khác nhau**: Mỗi agent có thể sử dụng provider LLM riêng
+- **Fallback mô hình**: Cấu hình mô hình chính và dự phòng cho khả năng phục hồi
+- **Cân bằng tải**: Phân phối yêu cầu qua nhiều endpoint
+- **Cấu hình tập trung**: Quản lý tất cả provider tại một nơi
+
+#### 📋 Tất Cả Vendor Được Hỗ Trợ
+
+| Vendor | `model` Prefix | Default API Base | Protocol | API Key |
+| ------------------- | ----------------- |-----------------------------------------------------| --------- | ---------------------------------------------------------------- |
+| **OpenAI** | `openai/` | `https://api.openai.com/v1` | OpenAI | [Get Key](https://platform.openai.com) |
+| **Anthropic** | `anthropic/` | `https://api.anthropic.com/v1` | Anthropic | [Get Key](https://console.anthropic.com) |
+| **智谱 AI (GLM)** | `zhipu/` | `https://open.bigmodel.cn/api/paas/v4` | OpenAI | [Get Key](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) |
+| **DeepSeek** | `deepseek/` | `https://api.deepseek.com/v1` | OpenAI | [Get Key](https://platform.deepseek.com) |
+| **Google Gemini** | `gemini/` | `https://generativelanguage.googleapis.com/v1beta` | OpenAI | [Get Key](https://aistudio.google.com/api-keys) |
+| **Groq** | `groq/` | `https://api.groq.com/openai/v1` | OpenAI | [Get Key](https://console.groq.com) |
+| **Moonshot** | `moonshot/` | `https://api.moonshot.cn/v1` | OpenAI | [Get Key](https://platform.moonshot.cn) |
+| **通义千问 (Qwen)** | `qwen/` | `https://dashscope.aliyuncs.com/compatible-mode/v1` | OpenAI | [Get Key](https://dashscope.console.aliyun.com) |
+| **NVIDIA** | `nvidia/` | `https://integrate.api.nvidia.com/v1` | OpenAI | [Get Key](https://build.nvidia.com) |
+| **Ollama** | `ollama/` | `http://localhost:11434/v1` | OpenAI | Local (no key needed) |
+| **OpenRouter** | `openrouter/` | `https://openrouter.ai/api/v1` | OpenAI | [Get Key](https://openrouter.ai/keys) |
+| **LiteLLM Proxy** | `litellm/` | `http://localhost:4000/v1` | OpenAI | Your LiteLLM proxy key |
+| **VLLM** | `vllm/` | `http://localhost:8000/v1` | OpenAI | Local |
+| **Cerebras** | `cerebras/` | `https://api.cerebras.ai/v1` | OpenAI | [Get Key](https://cerebras.ai) |
+| **VolcEngine (Doubao)** | `volcengine/` | `https://ark.cn-beijing.volces.com/api/v3` | OpenAI | [Get Key](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) |
+| **神算云** | `shengsuanyun/` | `https://router.shengsuanyun.com/api/v1` | OpenAI | - |
+| **BytePlus** | `byteplus/` | `https://ark.ap-southeast.bytepluses.com/api/v3` | OpenAI | [Get Key](https://www.byteplus.com) |
+| **Vivgrid** | `vivgrid/` | `https://api.vivgrid.com/v1` | OpenAI | [Get Key](https://vivgrid.com) |
+| **LongCat** | `longcat/` | `https://api.longcat.chat/openai` | OpenAI | [Get Key](https://longcat.chat/platform) |
+| **ModelScope (魔搭)**| `modelscope/` | `https://api-inference.modelscope.cn/v1` | OpenAI | [Get Token](https://modelscope.cn/my/tokens) |
+| **Antigravity** | `antigravity/` | Google Cloud | Custom | OAuth only |
+| **GitHub Copilot** | `github-copilot/` | `localhost:4321` | gRPC | - |
+
+#### Cấu Hình Cơ Bản
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "ark-code-latest",
+ "model": "volcengine/ark-code-latest",
+ "api_key": "sk-your-api-key"
+ },
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_key": "sk-your-openai-key"
+ },
+ {
+ "model_name": "claude-sonnet-4.6",
+ "model": "anthropic/claude-sonnet-4.6",
+ "api_key": "sk-ant-your-key"
+ },
+ {
+ "model_name": "glm-4.7",
+ "model": "zhipu/glm-4.7",
+ "api_key": "your-zhipu-key"
+ }
+ ],
+ "agents": {
+ "defaults": {
+ "model_name": "gpt-5.4"
+ }
+ }
+}
+```
+
+#### Ví Dụ Theo Vendor
+
+**OpenAI**
+
+```json
+{
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_key": "sk-..."
+}
+```
+
+**VolcEngine (Doubao)**
+
+```json
+{
+ "model_name": "ark-code-latest",
+ "model": "volcengine/ark-code-latest",
+ "api_key": "sk-..."
+}
+```
+
+**智谱 AI (GLM)**
+
+```json
+{
+ "model_name": "glm-4.7",
+ "model": "zhipu/glm-4.7",
+ "api_key": "your-key"
+}
+```
+
+**DeepSeek**
+
+```json
+{
+ "model_name": "deepseek-chat",
+ "model": "deepseek/deepseek-chat",
+ "api_key": "sk-..."
+}
+```
+
+**Anthropic (với API key)**
+
+```json
+{
+ "model_name": "claude-sonnet-4.6",
+ "model": "anthropic/claude-sonnet-4.6",
+ "api_key": "sk-ant-your-key"
+}
+```
+
+> Chạy `picoclaw auth login --provider anthropic` để dán API token.
+
+**Anthropic Messages API (định dạng native)**
+
+Để truy cập trực tiếp API Anthropic hoặc endpoint tùy chỉnh chỉ hỗ trợ định dạng message native của Anthropic:
+
+```json
+{
+ "model_name": "claude-opus-4-6",
+ "model": "anthropic-messages/claude-opus-4-6",
+ "api_key": "sk-ant-your-key",
+ "api_base": "https://api.anthropic.com"
+}
+```
+
+> Sử dụng giao thức `anthropic-messages` khi:
+> - Sử dụng proxy bên thứ ba chỉ hỗ trợ endpoint native `/v1/messages` của Anthropic (không tương thích OpenAI `/v1/chat/completions`)
+> - Kết nối đến dịch vụ như MiniMax, Synthetic yêu cầu định dạng message native của Anthropic
+> - Giao thức `anthropic` hiện tại trả về lỗi 404 (cho thấy endpoint không hỗ trợ định dạng tương thích OpenAI)
+>
+> **Lưu ý:** Giao thức `anthropic` sử dụng định dạng tương thích OpenAI (`/v1/chat/completions`), trong khi `anthropic-messages` sử dụng định dạng native của Anthropic (`/v1/messages`). Chọn dựa trên định dạng endpoint hỗ trợ.
+
+**Ollama (local)**
+
+```json
+{
+ "model_name": "llama3",
+ "model": "ollama/llama3"
+}
+```
+
+**Proxy/API Tùy Chỉnh**
+
+```json
+{
+ "model_name": "my-custom-model",
+ "model": "openai/custom-model",
+ "api_base": "https://my-proxy.com/v1",
+ "api_key": "sk-...",
+ "request_timeout": 300
+}
+```
+
+**LiteLLM Proxy**
+
+```json
+{
+ "model_name": "lite-gpt4",
+ "model": "litellm/lite-gpt4",
+ "api_base": "http://localhost:4000/v1",
+ "api_key": "sk-..."
+}
+```
+
+PicoClaw chỉ loại bỏ tiền tố ngoài `litellm/` trước khi gửi yêu cầu, nên alias proxy như `litellm/lite-gpt4` gửi `lite-gpt4`, trong khi `litellm/openai/gpt-4o` gửi `openai/gpt-4o`.
+
+#### Cân Bằng Tải
+
+Cấu hình nhiều endpoint cho cùng tên mô hình — PicoClaw sẽ tự động round-robin giữa chúng:
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_base": "https://api1.example.com/v1",
+ "api_key": "sk-key1"
+ },
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_base": "https://api2.example.com/v1",
+ "api_key": "sk-key2"
+ }
+ ]
+}
+```
+
+#### Di Chuyển Từ Cấu Hình Legacy `providers`
+
+Cấu hình `providers` cũ đã **ngừng hỗ trợ** nhưng vẫn được hỗ trợ để tương thích ngược.
+
+**Cấu hình cũ (ngừng hỗ trợ):**
+
+```json
+{
+ "providers": {
+ "zhipu": {
+ "api_key": "your-key",
+ "api_base": "https://open.bigmodel.cn/api/paas/v4"
+ }
+ },
+ "agents": {
+ "defaults": {
+ "provider": "zhipu",
+ "model": "glm-4.7"
+ }
+ }
+}
+```
+
+**Cấu hình mới (khuyến nghị):**
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "glm-4.7",
+ "model": "zhipu/glm-4.7",
+ "api_key": "your-key"
+ }
+ ],
+ "agents": {
+ "defaults": {
+ "model_name": "glm-4.7"
+ }
+ }
+}
+```
+
+Để xem hướng dẫn di chuyển chi tiết, xem [migration/model-list-migration.md](../migration/model-list-migration.md).
+
+### Kiến Trúc Provider
+
+PicoClaw định tuyến provider theo họ giao thức:
+
+- Giao thức tương thích OpenAI: OpenRouter, gateway tương thích OpenAI, Groq, Zhipu, và endpoint kiểu vLLM.
+- Giao thức Anthropic: Hành vi API native của Claude.
+- Đường dẫn Codex/OAuth: Tuyến xác thực OAuth/token của OpenAI.
+
+Điều này giữ runtime nhẹ trong khi làm cho backend tương thích OpenAI mới chủ yếu là thao tác cấu hình (`api_base` + `api_key`).
+
+
+Zhipu
+
+**1. Lấy API key và URL base**
+
+* Lấy [API key](https://bigmodel.cn/usercenter/proj-mgmt/apikeys)
+
+**2. Cấu hình**
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "workspace": "~/.picoclaw/workspace",
+ "model_name": "glm-4.7",
+ "max_tokens": 8192,
+ "temperature": 0.7,
+ "max_tool_iterations": 20
+ }
+ },
+ "providers": {
+ "zhipu": {
+ "api_key": "Your API Key",
+ "api_base": "https://open.bigmodel.cn/api/paas/v4"
+ }
+ }
+}
+```
+
+**3. Chạy**
+
+```bash
+picoclaw agent -m "Hello"
+```
+
+
+
+
+Ví dụ cấu hình đầy đủ
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "model_name": "anthropic/claude-opus-4-5"
+ }
+ },
+ "session": {
+ "dm_scope": "per-channel-peer"
+ },
+ "providers": {
+ "openrouter": {
+ "api_key": "sk-or-v1-xxx"
+ },
+ "groq": {
+ "api_key": "gsk_xxx"
+ }
+ },
+ "channels": {
+ "telegram": {
+ "enabled": true,
+ "token": "123456:ABC...",
+ "allow_from": ["123456789"]
+ },
+ "discord": {
+ "enabled": true,
+ "token": "",
+ "allow_from": [""]
+ },
+ "whatsapp": {
+ "enabled": false,
+ "bridge_url": "ws://localhost:3001",
+ "use_native": false,
+ "session_store_path": "",
+ "allow_from": []
+ },
+ "feishu": {
+ "enabled": false,
+ "app_id": "cli_xxx",
+ "app_secret": "xxx",
+ "encrypt_key": "",
+ "verification_token": "",
+ "allow_from": []
+ },
+ "qq": {
+ "enabled": false,
+ "app_id": "",
+ "app_secret": "",
+ "allow_from": []
+ }
+ },
+ "tools": {
+ "web": {
+ "brave": {
+ "enabled": false,
+ "api_key": "BSA...",
+ "max_results": 5
+ },
+ "duckduckgo": {
+ "enabled": true,
+ "max_results": 5
+ },
+ "perplexity": {
+ "enabled": false,
+ "api_key": "",
+ "max_results": 5
+ },
+ "searxng": {
+ "enabled": false,
+ "base_url": "http://localhost:8888",
+ "max_results": 5
+ }
+ },
+ "cron": {
+ "exec_timeout_minutes": 5
+ }
+ },
+ "heartbeat": {
+ "enabled": true,
+ "interval": 30
+ }
+}
+```
+
+
+
+---
+
+## 📝 So Sánh API Key
+
+| Service | Pricing | Use Case |
+| ---------------- | ------------------------ | ------------------------------------- |
+| **OpenRouter** | Free: 200K tokens/month | Multiple models (Claude, GPT-4, etc.) |
+| **Volcengine CodingPlan** | ¥9.9/first month | Best for Chinese users, multiple SOTA models (Doubao, DeepSeek, etc.) |
+| **Zhipu** | Free: 200K tokens/month | Suitable for Chinese users |
+| **Brave Search** | $5/1000 queries | Web search functionality |
+| **SearXNG** | Free (self-hosted) | Privacy-focused metasearch (70+ engines) |
+| **Groq** | Free tier available | Fast inference (Llama, Mixtral) |
+| **Cerebras** | Free tier available | Fast inference (Llama, Qwen, etc.) |
+| **LongCat** | Free: up to 5M tokens/day | Fast inference |
+| **ModelScope** | Free: 2000 requests/day | Inference (Qwen, GLM, DeepSeek, etc.) |
+
+---
+
+
+
+
diff --git a/docs/vi/spawn-tasks.md b/docs/vi/spawn-tasks.md
new file mode 100644
index 000000000..78f728040
--- /dev/null
+++ b/docs/vi/spawn-tasks.md
@@ -0,0 +1,61 @@
+# 🔄 Tác Vụ Bất Đồng Bộ và Spawn
+
+> Quay lại [README](../../README.vi.md)
+
+## Tác Vụ Nhanh (phản hồi trực tiếp)
+
+- Báo cáo thời gian hiện tại
+
+## Tác Vụ Dài (sử dụng spawn cho bất đồng bộ)
+
+- Tìm kiếm web tin tức AI và tóm tắt
+- Kiểm tra email và báo cáo tin nhắn quan trọng
+```
+
+**Hành vi chính:**
+
+| Feature | Description |
+| ----------------------- | --------------------------------------------------------- |
+| **spawn** | Creates async subagent, doesn't block heartbeat |
+| **Independent context** | Subagent has its own context, no session history |
+| **message tool** | Subagent communicates with user directly via message tool |
+| **Non-blocking** | After spawning, heartbeat continues to next task |
+
+#### Cách Giao Tiếp Subagent Hoạt Động
+
+```
+Heartbeat được kích hoạt
+ ↓
+Agent đọc HEARTBEAT.md
+ ↓
+Cho tác vụ dài: spawn subagent
+ ↓ ↓
+Tiếp tục tác vụ tiếp theo Subagent làm việc độc lập
+ ↓ ↓
+Tất cả tác vụ hoàn thành Subagent sử dụng công cụ "message"
+ ↓ ↓
+Phản hồi HEARTBEAT_OK Người dùng nhận kết quả trực tiếp
+```
+
+Subagent có quyền truy cập công cụ (message, web_search, v.v.) và có thể giao tiếp với người dùng độc lập mà không cần qua agent chính.
+
+**Cấu hình:**
+
+```json
+{
+ "heartbeat": {
+ "enabled": true,
+ "interval": 30
+ }
+}
+```
+
+| Option | Default | Description |
+| ---------- | ------- | ---------------------------------- |
+| `enabled` | `true` | Enable/disable heartbeat |
+| `interval` | `30` | Check interval in minutes (min: 5) |
+
+**Biến môi trường:**
+
+* `PICOCLAW_HEARTBEAT_ENABLED=false` để tắt
+* `PICOCLAW_HEARTBEAT_INTERVAL=60` để thay đổi khoảng thời gian
diff --git a/docs/vi/tools_configuration.md b/docs/vi/tools_configuration.md
new file mode 100644
index 000000000..76a336186
--- /dev/null
+++ b/docs/vi/tools_configuration.md
@@ -0,0 +1,360 @@
+# 🔧 Cấu Hình Công Cụ
+
+> Quay lại [README](../../README.vi.md)
+
+Cấu hình công cụ của PicoClaw nằm trong trường `tools` của `config.json`.
+
+## Cấu trúc thư mục
+
+```json
+{
+ "tools": {
+ "web": {
+ ...
+ },
+ "mcp": {
+ ...
+ },
+ "exec": {
+ ...
+ },
+ "cron": {
+ ...
+ },
+ "skills": {
+ ...
+ }
+ }
+}
+```
+
+## Công cụ Web
+
+Các công cụ web được sử dụng để tìm kiếm và tải nội dung web.
+
+### Web Fetcher
+Cài đặt chung để tải và xử lý nội dung trang web.
+
+| Cấu hình | Kiểu | Mặc định | Mô tả |
+|----------------------|--------|---------------|-----------------------------------------------------------------------------------------------|
+| `enabled` | bool | true | Bật khả năng tải trang web. |
+| `fetch_limit_bytes` | int | 10485760 | Kích thước tối đa của payload trang web cần tải, tính bằng byte (mặc định là 10MB). |
+| `format` | string | "plaintext" | Định dạng đầu ra của nội dung đã tải. Tùy chọn: `plaintext` hoặc `markdown` (khuyến nghị). |
+
+### Brave
+
+| Cấu hình | Kiểu | Mặc định | Mô tả |
+|----------------|--------|----------|----------------------------|
+| `enabled` | bool | false | Bật tìm kiếm Brave |
+| `api_key` | string | - | Khóa API Brave Search |
+| `max_results` | int | 5 | Số kết quả tối đa |
+
+### DuckDuckGo
+
+| Cấu hình | Kiểu | Mặc định | Mô tả |
+|----------------|------|----------|-------------------------------|
+| `enabled` | bool | true | Bật tìm kiếm DuckDuckGo |
+| `max_results` | int | 5 | Số kết quả tối đa |
+
+### Perplexity
+
+| Cấu hình | Kiểu | Mặc định | Mô tả |
+|----------------|--------|----------|-------------------------------|
+| `enabled` | bool | false | Bật tìm kiếm Perplexity |
+| `api_key` | string | - | Khóa API Perplexity |
+| `max_results` | int | 5 | Số kết quả tối đa |
+
+## Công cụ Exec
+
+Công cụ exec được sử dụng để thực thi các lệnh shell.
+
+| Cấu hình | Kiểu | Mặc định | Mô tả |
+|--------------------------|-------|----------|------------------------------------------------|
+| `enabled` | bool | true | Bật công cụ exec |
+| `enable_deny_patterns` | bool | true | Bật chặn lệnh nguy hiểm mặc định |
+| `custom_deny_patterns` | array | [] | Mẫu từ chối tùy chỉnh (biểu thức chính quy) |
+
+### Vô hiệu hóa Công cụ Exec
+
+Để hoàn toàn vô hiệu hóa công cụ `exec`, đặt `enabled` thành `false`:
+
+**Qua tệp cấu hình:**
+```json
+{
+ "tools": {
+ "exec": {
+ "enabled": false
+ }
+ }
+}
+```
+
+**Qua biến môi trường:**
+```bash
+PICOCLAW_TOOLS_EXEC_ENABLED=false
+```
+
+> **Lưu ý:** Khi bị vô hiệu hóa, agent sẽ không thể thực thi lệnh shell. Điều này cũng ảnh hưởng đến khả năng chạy lệnh shell theo lịch của công cụ Cron.
+
+### Chức năng
+
+- **`enable_deny_patterns`**: Đặt thành `false` để tắt hoàn toàn các mẫu chặn lệnh nguy hiểm mặc định
+- **`custom_deny_patterns`**: Thêm các mẫu regex từ chối tùy chỉnh; các lệnh khớp sẽ bị chặn
+
+### Các mẫu lệnh bị chặn mặc định
+
+Theo mặc định, PicoClaw chặn các lệnh nguy hiểm sau:
+
+- Lệnh xóa: `rm -rf`, `del /f/q`, `rmdir /s`
+- Thao tác đĩa: `format`, `mkfs`, `diskpart`, `dd if=`, ghi vào `/dev/sd*`
+- Thao tác hệ thống: `shutdown`, `reboot`, `poweroff`
+- Thay thế lệnh: `$()`, `${}`, dấu backtick
+- Pipe đến shell: `| sh`, `| bash`
+- Leo thang đặc quyền: `sudo`, `chmod`, `chown`
+- Điều khiển tiến trình: `pkill`, `killall`, `kill -9`
+- Thao tác từ xa: `curl | sh`, `wget | sh`, `ssh`
+- Quản lý gói: `apt`, `yum`, `dnf`, `npm install -g`, `pip install --user`
+- Container: `docker run`, `docker exec`
+- Git: `git push`, `git force`
+- Khác: `eval`, `source *.sh`
+
+### Hạn chế kiến trúc đã biết
+
+Bộ bảo vệ exec chỉ xác thực lệnh cấp cao nhất được gửi đến PicoClaw. Nó **không** kiểm tra đệ quy các tiến trình con được tạo bởi các công cụ build hoặc script sau khi lệnh đó bắt đầu chạy.
+
+Ví dụ về các quy trình có thể bỏ qua bộ bảo vệ lệnh trực tiếp sau khi lệnh ban đầu được cho phép:
+
+- `make run`
+- `go run ./cmd/...`
+- `cargo run`
+- `npm run build`
+
+Điều này có nghĩa là bộ bảo vệ hữu ích để chặn các lệnh trực tiếp rõ ràng nguy hiểm, nhưng nó **không phải** là sandbox đầy đủ cho các pipeline build chưa được xem xét. Nếu mô hình mối đe dọa của bạn bao gồm mã không đáng tin cậy trong workspace, hãy sử dụng cách ly mạnh hơn như container, VM hoặc quy trình phê duyệt xung quanh các lệnh build và chạy.
+
+### Ví dụ cấu hình
+
+```json
+{
+ "tools": {
+ "exec": {
+ "enable_deny_patterns": true,
+ "custom_deny_patterns": [
+ "\\brm\\s+-r\\b",
+ "\\bkillall\\s+python"
+ ]
+ }
+ }
+}
+```
+
+## Công cụ Cron
+
+Công cụ cron được sử dụng để lên lịch các tác vụ định kỳ.
+
+| Cấu hình | Kiểu | Mặc định | Mô tả |
+|--------------------------|------|----------|-----------------------------------------------------|
+| `exec_timeout_minutes` | int | 5 | Thời gian chờ thực thi tính bằng phút, 0 nghĩa là không giới hạn |
+
+## Công cụ MCP
+
+Công cụ MCP cho phép tích hợp với các máy chủ Model Context Protocol bên ngoài.
+
+### Khám phá công cụ (tải chậm)
+
+Khi kết nối với nhiều máy chủ MCP, việc hiển thị hàng trăm công cụ cùng lúc có thể làm cạn kiệt cửa sổ ngữ cảnh của LLM và tăng chi phí API. Tính năng **Discovery** giải quyết vấn đề này bằng cách giữ các công cụ MCP *ẩn* theo mặc định.
+
+Thay vì tải tất cả các công cụ, LLM được cung cấp một công cụ tìm kiếm nhẹ (sử dụng khớp từ khóa BM25 hoặc Regex). Khi LLM cần một khả năng cụ thể, nó tìm kiếm trong thư viện ẩn. Các công cụ khớp sau đó được tạm thời "mở khóa" và đưa vào ngữ cảnh trong số lượt được cấu hình (`ttl`).
+
+### Cấu hình toàn cục
+
+| Cấu hình | Kiểu | Mặc định | Mô tả |
+|-------------|--------|----------|-----------------------------------------------|
+| `enabled` | bool | false | Bật tích hợp MCP toàn cục |
+| `discovery` | object | `{}` | Cấu hình khám phá công cụ (xem bên dưới) |
+| `servers` | object | `{}` | Ánh xạ tên máy chủ đến cấu hình máy chủ |
+
+### Cấu hình Discovery (`discovery`)
+
+| Cấu hình | Kiểu | Mặc định | Mô tả |
+|----------------------|------|----------|-----------------------------------------------------------------------------------------------------------------------------------|
+| `enabled` | bool | false | Nếu true, các công cụ MCP bị ẩn và được tải theo yêu cầu qua tìm kiếm. Nếu false, tất cả công cụ được tải |
+| `ttl` | int | 5 | Số lượt hội thoại mà một công cụ đã khám phá vẫn được mở khóa |
+| `max_search_results` | int | 5 | Số công cụ tối đa được trả về cho mỗi truy vấn tìm kiếm |
+| `use_bm25` | bool | true | Bật công cụ tìm kiếm ngôn ngữ tự nhiên/từ khóa (`tool_search_tool_bm25`). **Cảnh báo**: tiêu tốn nhiều tài nguyên hơn tìm kiếm regex |
+| `use_regex` | bool | false | Bật công cụ tìm kiếm mẫu regex (`tool_search_tool_regex`) |
+
+> **Lưu ý:** Nếu `discovery.enabled` là `true`, bạn **phải** bật ít nhất một công cụ tìm kiếm (`use_bm25` hoặc `use_regex`),
+> nếu không ứng dụng sẽ không khởi động được.
+
+### Cấu hình từng máy chủ
+
+| Cấu hình | Kiểu | Bắt buộc | Mô tả |
+|------------|--------|----------|--------------------------------------------|
+| `enabled` | bool | có | Bật máy chủ MCP này |
+| `type` | string | không | Loại truyền tải: `stdio`, `sse`, `http` |
+| `command` | string | stdio | Lệnh thực thi cho truyền tải stdio |
+| `args` | array | không | Đối số lệnh cho truyền tải stdio |
+| `env` | object | không | Biến môi trường cho tiến trình stdio |
+| `env_file` | string | không | Đường dẫn đến tệp môi trường cho tiến trình stdio |
+| `url` | string | sse/http | URL endpoint cho truyền tải `sse`/`http` |
+| `headers` | object | không | Header HTTP cho truyền tải `sse`/`http` |
+
+### Hành vi truyền tải
+
+- Nếu bỏ qua `type`, truyền tải được tự động phát hiện:
+ - `url` được đặt → `sse`
+ - `command` được đặt → `stdio`
+- `http` và `sse` đều sử dụng `url` + `headers` tùy chọn.
+- `env` và `env_file` chỉ được áp dụng cho máy chủ `stdio`.
+
+### Ví dụ cấu hình
+
+#### 1) Máy chủ MCP Stdio
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "servers": {
+ "filesystem": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-filesystem",
+ "/tmp"
+ ]
+ }
+ }
+ }
+ }
+}
+```
+
+#### 2) Máy chủ MCP từ xa SSE/HTTP
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "servers": {
+ "remote-mcp": {
+ "enabled": true,
+ "type": "sse",
+ "url": "https://example.com/mcp",
+ "headers": {
+ "Authorization": "Bearer YOUR_TOKEN"
+ }
+ }
+ }
+ }
+ }
+}
+```
+
+#### 3) Thiết lập MCP quy mô lớn với khám phá công cụ được bật
+
+*Trong ví dụ này, LLM chỉ thấy `tool_search_tool_bm25`. Nó sẽ tìm kiếm và mở khóa động các công cụ Github hoặc Postgres chỉ khi được người dùng yêu cầu.*
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "discovery": {
+ "enabled": true,
+ "ttl": 5,
+ "max_search_results": 5,
+ "use_bm25": true,
+ "use_regex": false
+ },
+ "servers": {
+ "github": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-github"
+ ],
+ "env": {
+ "GITHUB_PERSONAL_ACCESS_TOKEN": "YOUR_GITHUB_TOKEN"
+ }
+ },
+ "postgres": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-postgres",
+ "postgresql://user:password@localhost/dbname"
+ ]
+ },
+ "slack": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-slack"
+ ],
+ "env": {
+ "SLACK_BOT_TOKEN": "YOUR_SLACK_BOT_TOKEN",
+ "SLACK_TEAM_ID": "YOUR_SLACK_TEAM_ID"
+ }
+ }
+ }
+ }
+ }
+}
+```
+
+## Công cụ Skills
+
+Công cụ skills cấu hình khám phá và cài đặt kỹ năng thông qua các registry như ClawHub.
+
+### Registry
+
+| Cấu hình | Kiểu | Mặc định | Mô tả |
+|------------------------------------|--------|-----------------------|----------------------------------------------|
+| `registries.clawhub.enabled` | bool | true | Bật registry ClawHub |
+| `registries.clawhub.base_url` | string | `https://clawhub.ai` | URL cơ sở ClawHub |
+| `registries.clawhub.auth_token` | string | `""` | Token Bearer tùy chọn để có giới hạn tốc độ cao hơn |
+| `registries.clawhub.search_path` | string | `/api/v1/search` | Đường dẫn API tìm kiếm |
+| `registries.clawhub.skills_path` | string | `/api/v1/skills` | Đường dẫn API Skills |
+| `registries.clawhub.download_path` | string | `/api/v1/download` | Đường dẫn API tải xuống |
+
+### Ví dụ cấu hình
+
+```json
+{
+ "tools": {
+ "skills": {
+ "registries": {
+ "clawhub": {
+ "enabled": true,
+ "base_url": "https://clawhub.ai",
+ "auth_token": "",
+ "search_path": "/api/v1/search",
+ "skills_path": "/api/v1/skills",
+ "download_path": "/api/v1/download"
+ }
+ }
+ }
+ }
+}
+```
+
+## Biến môi trường
+
+Tất cả các tùy chọn cấu hình có thể được ghi đè qua biến môi trường với định dạng `PICOCLAW_TOOLS__`:
+
+Ví dụ:
+
+- `PICOCLAW_TOOLS_WEB_BRAVE_ENABLED=true`
+- `PICOCLAW_TOOLS_EXEC_ENABLED=false`
+- `PICOCLAW_TOOLS_EXEC_ENABLE_DENY_PATTERNS=false`
+- `PICOCLAW_TOOLS_CRON_EXEC_TIMEOUT_MINUTES=10`
+- `PICOCLAW_TOOLS_MCP_ENABLED=true`
+
+Lưu ý: Cấu hình kiểu map lồng nhau (ví dụ `tools.mcp.servers..*`) được cấu hình trong `config.json` thay vì qua biến môi trường.
diff --git a/docs/vi/troubleshooting.md b/docs/vi/troubleshooting.md
new file mode 100644
index 000000000..961c932aa
--- /dev/null
+++ b/docs/vi/troubleshooting.md
@@ -0,0 +1,45 @@
+# 🐛 Khắc Phục Sự Cố
+
+> Quay lại [README](../../README.vi.md)
+
+## "model ... not found in model_list" hoặc OpenRouter "free is not a valid model ID"
+
+**Triệu chứng:** Bạn thấy một trong các lỗi sau:
+
+- `Error creating provider: model "openrouter/free" not found in model_list`
+- OpenRouter trả về 400: `"free is not a valid model ID"`
+
+**Nguyên nhân:** Trường `model` trong mục `model_list` của bạn là giá trị được gửi đến API. Đối với OpenRouter, bạn phải sử dụng ID mô hình **đầy đủ**, không phải dạng viết tắt.
+
+- **Sai:** `"model": "free"` → OpenRouter nhận được `free` và từ chối.
+- **Đúng:** `"model": "openrouter/free"` → OpenRouter nhận được `openrouter/free` (định tuyến tự động tầng miễn phí).
+
+**Cách sửa:** Trong `~/.picoclaw/config.json` (hoặc đường dẫn cấu hình của bạn):
+
+1. **agents.defaults.model_name** phải khớp với một `model_name` trong `model_list` (ví dụ: `"openrouter-free"`).
+2. **model** của mục đó phải là ID mô hình OpenRouter hợp lệ, ví dụ:
+ - `"openrouter/free"` – tầng miễn phí tự động
+ - `"google/gemini-2.0-flash-exp:free"`
+ - `"meta-llama/llama-3.1-8b-instruct:free"`
+
+Ví dụ:
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "model_name": "openrouter-free"
+ }
+ },
+ "model_list": [
+ {
+ "model_name": "openrouter-free",
+ "model": "openrouter/free",
+ "api_key": "sk-or-v1-YOUR_OPENROUTER_KEY",
+ "api_base": "https://openrouter.ai/api/v1"
+ }
+ ]
+}
+```
+
+Lấy khóa của bạn tại [OpenRouter Keys](https://openrouter.ai/keys).
diff --git a/docs/zh/ANTIGRAVITY_AUTH.md b/docs/zh/ANTIGRAVITY_AUTH.md
new file mode 100644
index 000000000..db7c81dea
--- /dev/null
+++ b/docs/zh/ANTIGRAVITY_AUTH.md
@@ -0,0 +1,809 @@
+> 返回 [README](../../README.zh.md)
+
+# Antigravity 认证与集成指南
+
+## 概述
+
+**Antigravity**(Google Cloud Code Assist)是由 Google 支持的 AI 模型提供商,通过 Google 的云基础设施提供对 Claude Opus 4.6 和 Gemini 等模型的访问。本文档提供了关于认证工作原理、如何获取模型以及如何在 PicoClaw 中实现新提供商的完整指南。
+
+---
+
+## 目录
+
+1. [认证流程](#认证流程)
+2. [OAuth 实现细节](#oauth-实现细节)
+3. [令牌管理](#令牌管理)
+4. [模型列表获取](#模型列表获取)
+5. [用量追踪](#用量追踪)
+6. [提供商插件结构](#提供商插件结构)
+7. [集成要求](#集成要求)
+8. [API 端点](#api-端点)
+9. [配置](#配置)
+10. [在 PicoClaw 中创建新提供商](#在-picoclaw-中创建新提供商)
+
+---
+
+## 认证流程
+
+### 1. 带 PKCE 的 OAuth 2.0
+
+Antigravity 使用 **OAuth 2.0 with PKCE(Proof Key for Code Exchange)** 进行安全认证:
+
+```
+┌─────────────┐ ┌─────────────────┐
+│ Client │ ───(1) Generate PKCE Pair────────> │ │
+│ │ ───(2) Open Auth URL─────────────> │ Google OAuth │
+│ │ │ Server │
+│ │ <──(3) Redirect with Code───────── │ │
+│ │ └─────────────────┘
+│ │ ───(4) Exchange Code for Tokens──> │ Token URL │
+│ │ │ │
+│ │ <──(5) Access + Refresh Tokens──── │ │
+└─────────────┘ └─────────────────┘
+```
+
+### 2. 详细步骤
+
+#### 步骤 1:生成 PKCE 参数
+```typescript
+function generatePkce(): { verifier: string; challenge: string } {
+ const verifier = randomBytes(32).toString("hex");
+ const challenge = createHash("sha256").update(verifier).digest("base64url");
+ return { verifier, challenge };
+}
+```
+
+#### 步骤 2:构建授权 URL
+```typescript
+const AUTH_URL = "https://accounts.google.com/o/oauth2/v2/auth";
+const REDIRECT_URI = "http://localhost:51121/oauth-callback";
+
+function buildAuthUrl(params: { challenge: string; state: string }): string {
+ const url = new URL(AUTH_URL);
+ url.searchParams.set("client_id", CLIENT_ID);
+ url.searchParams.set("response_type", "code");
+ url.searchParams.set("redirect_uri", REDIRECT_URI);
+ url.searchParams.set("scope", SCOPES.join(" "));
+ url.searchParams.set("code_challenge", params.challenge);
+ url.searchParams.set("code_challenge_method", "S256");
+ url.searchParams.set("state", params.state);
+ url.searchParams.set("access_type", "offline");
+ url.searchParams.set("prompt", "consent");
+ return url.toString();
+}
+```
+
+**所需权限范围:**
+```typescript
+const SCOPES = [
+ "https://www.googleapis.com/auth/cloud-platform",
+ "https://www.googleapis.com/auth/userinfo.email",
+ "https://www.googleapis.com/auth/userinfo.profile",
+ "https://www.googleapis.com/auth/cclog",
+ "https://www.googleapis.com/auth/experimentsandconfigs",
+];
+```
+
+#### 步骤 3:处理 OAuth 回调
+
+**自动模式(本地开发):**
+- 在端口 51121 上启动本地 HTTP 服务器
+- 等待来自 Google 的重定向
+- 从查询参数中提取授权码
+
+**手动模式(远程/无头环境):**
+- 向用户显示授权 URL
+- 用户在浏览器中完成认证
+- 用户将完整的重定向 URL 粘贴回终端
+- 从粘贴的 URL 中解析授权码
+
+#### 步骤 4:用授权码交换令牌
+```typescript
+const TOKEN_URL = "https://oauth2.googleapis.com/token";
+
+async function exchangeCode(params: {
+ code: string;
+ verifier: string;
+}): Promise<{ access: string; refresh: string; expires: number }> {
+ const response = await fetch(TOKEN_URL, {
+ method: "POST",
+ headers: { "Content-Type": "application/x-www-form-urlencoded" },
+ body: new URLSearchParams({
+ client_id: CLIENT_ID,
+ client_secret: CLIENT_SECRET,
+ code: params.code,
+ grant_type: "authorization_code",
+ redirect_uri: REDIRECT_URI,
+ code_verifier: params.verifier,
+ }),
+ });
+
+ const data = await response.json();
+
+ return {
+ access: data.access_token,
+ refresh: data.refresh_token,
+ expires: Date.now() + data.expires_in * 1000 - 5 * 60 * 1000, // 5 min buffer
+ };
+}
+```
+
+#### 步骤 5:获取额外的用户数据
+
+**用户邮箱:**
+```typescript
+async function fetchUserEmail(accessToken: string): Promise {
+ const response = await fetch(
+ "https://www.googleapis.com/oauth2/v1/userinfo?alt=json",
+ { headers: { Authorization: `Bearer ${accessToken}` } }
+ );
+ const data = await response.json();
+ return data.email;
+}
+```
+
+**项目 ID(API 调用必需):**
+```typescript
+async function fetchProjectId(accessToken: string): Promise {
+ const headers = {
+ Authorization: `Bearer ${accessToken}`,
+ "Content-Type": "application/json",
+ "User-Agent": "google-api-nodejs-client/9.15.1",
+ "X-Goog-Api-Client": "google-cloud-sdk vscode_cloudshelleditor/0.1",
+ "Client-Metadata": JSON.stringify({
+ ideType: "IDE_UNSPECIFIED",
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+ }),
+ };
+
+ const response = await fetch(
+ "https://cloudcode-pa.googleapis.com/v1internal:loadCodeAssist",
+ {
+ method: "POST",
+ headers,
+ body: JSON.stringify({
+ metadata: {
+ ideType: "IDE_UNSPECIFIED",
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+ },
+ }),
+ }
+ );
+
+ const data = await response.json();
+ return data.cloudaicompanionProject || "rising-fact-p41fc"; // 默认回退值
+}
+```
+
+---
+
+## OAuth 实现细节
+
+### 客户端凭据
+
+**重要:** 这些凭据在源代码中以 base64 编码存储,用于与 pi-ai 同步:
+
+```typescript
+const decode = (s: string) => Buffer.from(s, "base64").toString();
+
+const CLIENT_ID = decode(
+ "MTA3MTAwNjA2MDU5MS10bWhzc2luMmgyMWxjcmUyMzV2dG9sb2poNGc0MDNlcC5hcHBzLmdvb2dsZXVzZXJjb250ZW50LmNvbQ=="
+);
+const CLIENT_SECRET = decode("R09DU1BYLUs1OEZXUjQ4NkxkTEoxbUxCOHNYQzR6NnFEQWY=");
+```
+
+### OAuth 流程模式
+
+1. **自动流程**(有浏览器的本地机器):
+ - 自动打开浏览器
+ - 本地回调服务器捕获重定向
+ - 初始认证后无需用户交互
+
+2. **手动流程**(远程/无头/WSL2 环境):
+ - 显示 URL 供手动复制粘贴
+ - 用户在外部浏览器中完成认证
+ - 用户将完整的重定向 URL 粘贴回来
+
+```typescript
+function shouldUseManualOAuthFlow(isRemote: boolean): boolean {
+ return isRemote || isWSL2Sync();
+}
+```
+
+---
+
+## 令牌管理
+
+### 认证配置文件结构
+
+```typescript
+type OAuthCredential = {
+ type: "oauth";
+ provider: "google-antigravity";
+ access: string; // 访问令牌
+ refresh: string; // 刷新令牌
+ expires: number; // 过期时间戳(毫秒,自 epoch 起)
+ email?: string; // 用户邮箱
+ projectId?: string; // Google Cloud 项目 ID
+};
+```
+
+### 令牌刷新
+
+凭据包含一个刷新令牌,可在当前访问令牌过期时用于获取新的访问令牌。过期时间设置了 5 分钟的缓冲区以防止竞态条件。
+
+---
+
+## 模型列表获取
+
+### 获取可用模型
+
+```typescript
+const BASE_URL = "https://cloudcode-pa.googleapis.com";
+
+async function fetchAvailableModels(
+ accessToken: string,
+ projectId: string
+): Promise {
+ const headers = {
+ Authorization: `Bearer ${accessToken}`,
+ "Content-Type": "application/json",
+ "User-Agent": "antigravity",
+ "X-Goog-Api-Client": "google-cloud-sdk vscode_cloudshelleditor/0.1",
+ };
+
+ const response = await fetch(
+ `${BASE_URL}/v1internal:fetchAvailableModels`,
+ {
+ method: "POST",
+ headers,
+ body: JSON.stringify({ project: projectId }),
+ }
+ );
+
+ const data = await response.json();
+
+ // 返回带有配额信息的模型
+ return Object.entries(data.models).map(([modelId, modelInfo]) => ({
+ id: modelId,
+ displayName: modelInfo.displayName,
+ quotaInfo: {
+ remainingFraction: modelInfo.quotaInfo?.remainingFraction,
+ resetTime: modelInfo.quotaInfo?.resetTime,
+ isExhausted: modelInfo.quotaInfo?.isExhausted,
+ },
+ }));
+}
+```
+
+### 响应格式
+
+```typescript
+type FetchAvailableModelsResponse = {
+ models?: Record;
+};
+```
+
+---
+
+## 用量追踪
+
+### 获取用量数据
+
+```typescript
+export async function fetchAntigravityUsage(
+ token: string,
+ timeoutMs: number
+): Promise {
+ // 1. 获取额度和计划信息
+ const loadCodeAssistRes = await fetch(
+ `${BASE_URL}/v1internal:loadCodeAssist`,
+ {
+ method: "POST",
+ headers: {
+ Authorization: `Bearer ${token}`,
+ "Content-Type": "application/json",
+ },
+ body: JSON.stringify({
+ metadata: {
+ ideType: "ANTIGRAVITY",
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+ },
+ }),
+ }
+ );
+
+ // 提取额度信息
+ const { availablePromptCredits, planInfo, currentTier } = data;
+
+ // 2. 获取模型配额
+ const modelsRes = await fetch(
+ `${BASE_URL}/v1internal:fetchAvailableModels`,
+ {
+ method: "POST",
+ headers: { Authorization: `Bearer ${token}` },
+ body: JSON.stringify({ project: projectId }),
+ }
+ );
+
+ // 构建用量窗口
+ return {
+ provider: "google-antigravity",
+ displayName: "Google Antigravity",
+ windows: [
+ { label: "Credits", usedPercent: calculateUsedPercent(available, monthly) },
+ // 各模型配额...
+ ],
+ plan: currentTier?.name || planType,
+ };
+}
+```
+
+### 用量响应结构
+
+```typescript
+type ProviderUsageSnapshot = {
+ provider: "google-antigravity";
+ displayName: string;
+ windows: UsageWindow[];
+ plan?: string;
+ error?: string;
+};
+
+type UsageWindow = {
+ label: string; // "Credits" 或模型 ID
+ usedPercent: number; // 0-100
+ resetAt?: number; // 配额重置的时间戳
+};
+```
+
+---
+
+## 提供商插件结构
+
+### 插件定义
+
+```typescript
+const antigravityPlugin = {
+ id: "google-antigravity-auth",
+ name: "Google Antigravity Auth",
+ description: "OAuth flow for Google Antigravity (Cloud Code Assist)",
+ configSchema: emptyPluginConfigSchema(),
+
+ register(api: PicoClawPluginApi) {
+ api.registerProvider({
+ id: "google-antigravity",
+ label: "Google Antigravity",
+ docsPath: "/providers/models",
+ aliases: ["antigravity"],
+
+ auth: [
+ {
+ id: "oauth",
+ label: "Google OAuth",
+ hint: "PKCE + localhost callback",
+ kind: "oauth",
+ run: async (ctx: ProviderAuthContext) => {
+ // OAuth 实现在此处
+ },
+ },
+ ],
+ });
+ },
+};
+```
+
+### ProviderAuthContext
+
+```typescript
+type ProviderAuthContext = {
+ config: PicoClawConfig;
+ agentDir?: string;
+ workspaceDir?: string;
+ prompter: WizardPrompter; // UI 提示/通知
+ runtime: RuntimeEnv; // 日志等
+ isRemote: boolean; // 是否在远程运行
+ openUrl: (url: string) => Promise; // 浏览器打开器
+ oauth: {
+ createVpsAwareHandlers: Function;
+ };
+};
+```
+
+### ProviderAuthResult
+
+```typescript
+type ProviderAuthResult = {
+ profiles: Array<{
+ profileId: string;
+ credential: AuthProfileCredential;
+ }>;
+ configPatch?: Partial;
+ defaultModel?: string;
+ notes?: string[];
+};
+```
+
+---
+
+## 集成要求
+
+### 1. 所需环境/依赖
+
+- Go ≥ 1.25
+- PicoClaw 代码库(`pkg/providers/` 和 `pkg/auth/`)
+- `crypto` 和 `net/http` 标准库包
+
+### 2. API 调用所需的请求头
+
+```typescript
+const REQUIRED_HEADERS = {
+ "Authorization": `Bearer ${accessToken}`,
+ "Content-Type": "application/json",
+ "User-Agent": "antigravity", // 或 "google-api-nodejs-client/9.15.1"
+ "X-Goog-Api-Client": "google-cloud-sdk vscode_cloudshelleditor/0.1",
+};
+
+// 对于 loadCodeAssist 调用,还需包含:
+const CLIENT_METADATA = {
+ ideType: "ANTIGRAVITY", // 或 "IDE_UNSPECIFIED"
+ platform: "PLATFORM_UNSPECIFIED",
+ pluginType: "GEMINI",
+};
+```
+
+### 3. 模型 Schema 清理
+
+Antigravity 使用兼容 Gemini 的模型,因此工具 schema 必须进行清理:
+
+```typescript
+const GOOGLE_SCHEMA_UNSUPPORTED_KEYWORDS = new Set([
+ "patternProperties",
+ "additionalProperties",
+ "$schema",
+ "$id",
+ "$ref",
+ "$defs",
+ "definitions",
+ "examples",
+ "minLength",
+ "maxLength",
+ "minimum",
+ "maximum",
+ "multipleOf",
+ "pattern",
+ "format",
+ "minItems",
+ "maxItems",
+ "uniqueItems",
+ "minProperties",
+ "maxProperties",
+]);
+
+// 发送前清理 schema
+function cleanToolSchemaForGemini(schema: Record): unknown {
+ // 移除不支持的关键字
+ // 确保顶层有 type: "object"
+ // 展平 anyOf/oneOf 联合类型
+}
+```
+
+### 4. 思维块处理(Claude 模型)
+
+对于 Antigravity 的 Claude 模型,思维块需要特殊处理:
+
+```typescript
+const ANTIGRAVITY_SIGNATURE_RE = /^[A-Za-z0-9+/]+={0,2}$/;
+
+export function sanitizeAntigravityThinkingBlocks(
+ messages: AgentMessage[]
+): AgentMessage[] {
+ // 验证思维签名
+ // 规范化签名字段
+ // 丢弃未签名的思维块
+}
+```
+
+---
+
+## API 端点
+
+### 认证端点
+
+| 端点 | 方法 | 用途 |
+|------|------|------|
+| `https://accounts.google.com/o/oauth2/v2/auth` | GET | OAuth 授权 |
+| `https://oauth2.googleapis.com/token` | POST | 令牌交换 |
+| `https://www.googleapis.com/oauth2/v1/userinfo` | GET | 用户信息(邮箱) |
+
+### Cloud Code Assist 端点
+
+| 端点 | 方法 | 用途 |
+|------|------|------|
+| `https://cloudcode-pa.googleapis.com/v1internal:loadCodeAssist` | POST | 加载项目信息、额度、计划 |
+| `https://cloudcode-pa.googleapis.com/v1internal:fetchAvailableModels` | POST | 列出可用模型及配额 |
+| `https://cloudcode-pa.googleapis.com/v1internal:streamGenerateContent?alt=sse` | POST | 聊天流式端点 |
+
+**API 请求格式(聊天):**
+`v1internal:streamGenerateContent` 端点期望一个包装标准 Gemini 请求的信封格式:
+
+```json
+{
+ "project": "your-project-id",
+ "model": "model-id",
+ "request": {
+ "contents": [...],
+ "systemInstruction": {...},
+ "generationConfig": {...},
+ "tools": [...]
+ },
+ "requestType": "agent",
+ "userAgent": "antigravity",
+ "requestId": "agent-timestamp-random"
+}
+```
+
+**API 响应格式(SSE):**
+每条 SSE 消息(`data: {...}`)被包装在 `response` 字段中:
+
+```json
+{
+ "response": {
+ "candidates": [...],
+ "usageMetadata": {...},
+ "modelVersion": "...",
+ "responseId": "..."
+ },
+ "traceId": "...",
+ "metadata": {}
+}
+```
+
+---
+
+## 配置
+
+### config.json 配置
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "gemini-flash",
+ "model": "antigravity/gemini-3-flash",
+ "auth_method": "oauth"
+ }
+ ],
+ "agents": {
+ "defaults": {
+ "model_name": "gemini-flash"
+ }
+ }
+}
+```
+
+### 认证配置文件存储
+
+认证配置文件存储在 `~/.picoclaw/auth.json` 中:
+
+```json
+{
+ "credentials": {
+ "google-antigravity": {
+ "access_token": "ya29...",
+ "refresh_token": "1//...",
+ "expires_at": "2026-01-01T00:00:00Z",
+ "provider": "google-antigravity",
+ "auth_method": "oauth",
+ "email": "user@example.com",
+ "project_id": "my-project-id"
+ }
+ }
+}
+```
+
+---
+
+## 在 PicoClaw 中创建新提供商
+
+PicoClaw 提供商以 Go 包的形式实现,位于 `pkg/providers/` 下。要添加新提供商:
+
+### 分步实现
+
+#### 1. 创建提供商文件
+
+在 `pkg/providers/` 中创建新的 Go 文件:
+
+```
+pkg/providers/
+└── your_provider.go
+```
+
+#### 2. 实现 Provider 接口
+
+你的提供商必须实现 `pkg/providers/types.go` 中定义的 `Provider` 接口:
+
+```go
+package providers
+
+type YourProvider struct {
+ apiKey string
+ apiBase string
+}
+
+func NewYourProvider(apiKey, apiBase, proxy string) *YourProvider {
+ if apiBase == "" {
+ apiBase = "https://api.your-provider.com/v1"
+ }
+ return &YourProvider{apiKey: apiKey, apiBase: apiBase}
+}
+
+func (p *YourProvider) Chat(ctx context.Context, messages []Message, tools []Tool, cb StreamCallback) error {
+ // 实现带流式传输的聊天补全
+}
+```
+
+#### 3. 在工厂中注册
+
+将你的提供商添加到 `pkg/providers/factory.go` 中的协议分支:
+
+```go
+case "your-provider":
+ return NewYourProvider(sel.apiKey, sel.apiBase, sel.proxy), nil
+```
+
+#### 4. 添加默认配置(可选)
+
+在 `pkg/config/defaults.go` 中添加默认条目:
+
+```go
+{
+ ModelName: "your-model",
+ Model: "your-provider/model-name",
+ APIKey: "",
+},
+```
+
+#### 5. 添加认证支持(可选)
+
+如果你的提供商需要 OAuth 或特殊认证,在 `cmd/picoclaw/internal/auth/helpers.go` 中添加分支:
+
+```go
+case "your-provider":
+ authLoginYourProvider()
+```
+
+#### 6. 通过 `config.json` 配置
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "your-model",
+ "model": "your-provider/model-name",
+ "api_key": "your-api-key",
+ "api_base": "https://api.your-provider.com/v1"
+ }
+ ]
+}
+```
+
+---
+
+## 测试你的实现
+
+### CLI 命令
+
+```bash
+# 使用提供商进行认证
+picoclaw auth login --provider your-provider
+
+# 列出模型(用于 Antigravity)
+picoclaw auth models
+
+# 启动网关
+picoclaw gateway
+
+# 使用指定模型运行代理
+picoclaw agent -m "Hello" --model your-model
+```
+
+### 测试用环境变量
+
+```bash
+# 覆盖默认模型
+export PICOCLAW_AGENTS_DEFAULTS_MODEL=your-model
+
+# 覆盖提供商设置
+export PICOCLAW_MODEL_LIST='[{"model_name":"your-model","model":"your-provider/model-name","api_key":"..."}]'
+```
+
+---
+
+## 参考资料
+
+- **源文件:**
+ - `pkg/providers/antigravity_provider.go` - Antigravity 提供商实现
+ - `pkg/auth/oauth.go` - OAuth 流程实现
+ - `pkg/auth/store.go` - 认证凭据存储(`~/.picoclaw/auth.json`)
+ - `pkg/providers/factory.go` - 提供商工厂和协议路由
+ - `pkg/providers/types.go` - 提供商接口定义
+ - `cmd/picoclaw/internal/auth/helpers.go` - 认证 CLI 命令
+
+- **文档:**
+ - `docs/ANTIGRAVITY_USAGE.md` - Antigravity 使用指南
+ - `docs/migration/model-list-migration.md` - 迁移指南
+
+---
+
+## 注意事项
+
+1. **Google Cloud 项目:** Antigravity 要求在你的 Google Cloud 项目上启用 Gemini for Google Cloud
+2. **配额:** 使用 Google Cloud 项目配额(非独立计费)
+3. **模型访问:** 可用模型取决于你的 Google Cloud 项目配置
+4. **思维块:** 通过 Antigravity 使用的 Claude 模型需要对带签名的思维块进行特殊处理
+5. **Schema 清理:** 工具 schema 必须清理以移除不支持的 JSON Schema 关键字
+
+---
+
+---
+
+## 常见错误处理
+
+### 1. 速率限制(HTTP 429)
+
+当项目/模型配额耗尽时,Antigravity 会返回 429 错误。错误响应通常在 `details` 字段中包含 `quotaResetDelay`。
+
+**429 错误示例:**
+```json
+{
+ "error": {
+ "code": 429,
+ "message": "You have exhausted your capacity on this model. Your quota will reset after 4h30m28s.",
+ "status": "RESOURCE_EXHAUSTED",
+ "details": [
+ {
+ "@type": "type.googleapis.com/google.rpc.ErrorInfo",
+ "metadata": {
+ "quotaResetDelay": "4h30m28.060903746s"
+ }
+ }
+ ]
+ }
+}
+```
+
+### 2. 空响应(受限模型)
+
+某些模型可能出现在可用模型列表中,但返回空响应(200 OK 但 SSE 流为空)。这通常发生在当前项目没有权限使用的预览版或受限模型上。
+
+**处理方式:** 将空响应视为错误,通知用户该模型可能对其项目受限或无效。
+
+---
+
+## 故障排除
+
+### "Token expired"(令牌已过期)
+- 刷新 OAuth 令牌:`picoclaw auth login --provider antigravity`
+
+### "Gemini for Google Cloud is not enabled"(Gemini for Google Cloud 未启用)
+- 在 Google Cloud Console 中启用该 API
+
+### "Project not found"(项目未找到)
+- 确保你的 Google Cloud 项目已启用必要的 API
+- 检查认证过程中项目 ID 是否正确获取
+
+### 模型未出现在列表中
+- 验证 OAuth 认证是否成功完成
+- 检查认证配置文件存储:`~/.picoclaw/auth.json`
+- 重新运行 `picoclaw auth login --provider antigravity`
diff --git a/docs/zh/ANTIGRAVITY_USAGE.md b/docs/zh/ANTIGRAVITY_USAGE.md
new file mode 100644
index 000000000..2218618a9
--- /dev/null
+++ b/docs/zh/ANTIGRAVITY_USAGE.md
@@ -0,0 +1,72 @@
+> 返回 [README](../../README.zh.md)
+
+# 在 PicoClaw 中使用 Antigravity 提供商
+
+本指南介绍如何在 PicoClaw 中设置和使用 **Antigravity**(Google Cloud Code Assist)提供商。
+
+## 前提条件
+
+1. 一个 Google 账户。
+2. 已启用 Google Cloud Code Assist(通常通过"Gemini for Google Cloud"引导流程获取)。
+
+## 1. 身份验证
+
+要使用 Antigravity 进行身份验证,请运行以下命令:
+
+```bash
+picoclaw auth login --provider antigravity
+```
+
+### 手动验证(无界面/VPS 环境)
+如果你在服务器(Coolify/Docker)上运行且无法访问 `localhost`,请按照以下步骤操作:
+1. 运行上述命令。
+2. 复制提供的 URL 并在本地浏览器中打开。
+3. 完成登录。
+4. 浏览器将重定向到 `localhost:51121` URL(页面将无法加载)。
+5. **从浏览器地址栏复制该最终 URL**。
+6. **将其粘贴回 PicoClaw 正在等待的终端中**。
+
+PicoClaw 将自动提取授权码并完成流程。
+
+## 2. 管理模型
+
+### 列出可用模型
+查看你的项目可以访问哪些模型并检查其配额:
+
+```bash
+picoclaw auth models
+```
+
+### 切换模型
+你可以在 `~/.picoclaw/config.json` 中更改默认模型,或通过 CLI 覆盖:
+
+```bash
+# 为单个命令覆盖
+picoclaw agent -m "Hello" --model claude-opus-4-6-thinking
+```
+
+## 3. 实际使用(Coolify/Docker)
+
+如果你通过 Coolify 或 Docker 部署,请按照以下步骤进行测试:
+
+1. **环境变量**:
+ * `PICOCLAW_AGENTS_DEFAULTS_MODEL=gemini-flash`
+2. **身份验证持久化**:
+ 如果你已在本地登录,可以将凭据复制到服务器:
+ ```bash
+ scp ~/.picoclaw/auth.json user@your-server:~/.picoclaw/
+ ```
+ *或者*,如果你有终端访问权限,可以在服务器上运行一次 `auth login` 命令。
+
+## 4. 故障排除
+
+* **空响应**:如果模型返回空回复,可能是该模型在你的项目中受到限制。请尝试 `gemini-3-flash` 或 `claude-opus-4-6-thinking`。
+* **429 速率限制**:Antigravity 有严格的配额限制。如果触发限制,PicoClaw 将在错误消息中显示"重置时间"。
+* **404 未找到**:确保你使用的是 `picoclaw auth models` 列表中的模型 ID。请使用短 ID(例如 `gemini-3-flash`),而非完整路径。
+
+## 5. 可用模型总结
+
+根据测试,以下模型最为可靠:
+* `gemini-3-flash`(快速,高可用性)
+* `gemini-2.5-flash-lite`(轻量级)
+* `claude-opus-4-6-thinking`(强大,包含推理能力)
diff --git a/docs/zh/chat-apps.md b/docs/zh/chat-apps.md
new file mode 100644
index 000000000..a0206a7d6
--- /dev/null
+++ b/docs/zh/chat-apps.md
@@ -0,0 +1,608 @@
+# 💬 聊天应用配置
+
+> 返回 [README](../../README.zh.md)
+
+## 💬 聊天应用集成 (Chat Apps)
+
+PicoClaw 支持多种聊天平台,使您的 Agent 能够连接到任何地方。
+
+> **注意**: 所有 Webhook 类渠道(LINE、WeCom 等)均挂载在同一个 Gateway HTTP 服务器上(`gateway.host`:`gateway.port`,默认 `127.0.0.1:18790`),无需为每个渠道单独配置端口。注意:飞书(Feishu)使用 WebSocket/SDK 模式,不通过该共享 HTTP webhook 服务器接收消息。
+
+### 核心渠道
+
+| 渠道 | 设置难度 | 特性说明 | 文档链接 |
+| -------------------- | ----------- | ----------------------------------------- | --------------------------------------------------------------------------------------------------------------- |
+| **Telegram** | ⭐ 简单 | 推荐,支持语音转文字,长轮询无需公网 | [查看文档](../channels/telegram/README.zh.md) |
+| **Discord** | ⭐ 简单 | Socket Mode,支持群组/私信,Bot 生态成熟 | [查看文档](../channels/discord/README.zh.md) |
+| **WhatsApp** | ⭐ 简单 | 原生 (QR 扫码) 或 Bridge URL | [查看文档](#whatsapp) |
+| **Slack** | ⭐ 简单 | **Socket Mode** (无需公网 IP),企业级支持 | [查看文档](../channels/slack/README.zh.md) |
+| **Matrix** | ⭐⭐ 中等 | 联邦协议,支持自建 homeserver 与公开服务器 | [查看文档](../channels/matrix/README.zh.md) |
+| **QQ** | ⭐⭐ 中等 | 官方机器人 API,适合国内社群 | [查看文档](../channels/qq/README.zh.md) |
+| **钉钉 (DingTalk)** | ⭐⭐ 中等 | Stream 模式无需公网,企业办公首选 | [查看文档](../channels/dingtalk/README.zh.md) |
+| **LINE** | ⭐⭐⭐ 较难 | 需要 HTTPS Webhook | [查看文档](../channels/line/README.zh.md) |
+| **企业微信 (WeCom)** | ⭐⭐⭐ 较难 | 支持群机器人(Webhook)、自建应用(API)和智能机器人(AI Bot) | [Bot 文档](../channels/wecom/wecom_bot/README.zh.md) / [App 文档](../channels/wecom/wecom_app/README.zh.md) / [AI Bot 文档](../channels/wecom/wecom_aibot/README.zh.md) |
+| **飞书 (Feishu)** | ⭐⭐⭐ 较难 | 企业级协作,功能丰富 | [查看文档](../channels/feishu/README.zh.md) |
+| **IRC** | ⭐⭐ 中等 | 服务器 + TLS 配置 | - |
+| **OneBot** | ⭐⭐ 中等 | 兼容 NapCat/Go-CQHTTP,社区生态丰富 | [查看文档](../channels/onebot/README.zh.md) |
+| **MaixCam** | ⭐ 简单 | 专为 AI 摄像头设计的硬件集成通道 | [查看文档](../channels/maixcam/README.zh.md) |
+| **Pico** | ⭐ 简单 | PicoClaw 原生协议通道 | |
+
+---
+
+
+Telegram(推荐)
+
+**1. 创建 Bot**
+
+* 打开 Telegram,搜索 `@BotFather`
+* 发送 `/newbot`,按提示操作
+* 复制 Token
+
+**2. 配置**
+
+```json
+{
+ "channels": {
+ "telegram": {
+ "enabled": true,
+ "token": "YOUR_BOT_TOKEN",
+ "allow_from": ["YOUR_USER_ID"]
+ }
+ }
+}
+```
+
+> 通过 Telegram 上的 `@userinfobot` 获取你的 User ID。
+
+**3. 运行**
+
+```bash
+picoclaw gateway
+```
+
+**4. Telegram 命令菜单(启动时自动注册)**
+
+PicoClaw 使用统一的命令定义来源。启动时会自动将 Telegram 支持的命令(例如 `/start`、`/help`、`/show`、`/list`)注册到 Bot 命令菜单,确保菜单展示与实际行为一致。
+Telegram 侧保留的是命令菜单注册能力;通用命令的实际执行统一走 Agent Loop 中的 commands executor。
+
+如果注册因网络或 API 短暂异常失败,不会阻塞 channel 启动;系统会在后台自动重试。
+
+
+
+
+Discord
+
+**1. 创建 Bot**
+
+* 前往
+* 创建应用 → Bot → 添加 Bot
+* 复制 Bot Token
+
+**2. 启用 Intents**
+
+* 在 Bot 设置中启用 **MESSAGE CONTENT INTENT**
+* (可选)启用 **SERVER MEMBERS INTENT**(如需基于成员数据的白名单)
+
+**3. 获取 User ID**
+
+* Discord 设置 → 高级 → 启用 **开发者模式**
+* 右键点击头像 → **复制用户 ID**
+
+**4. 配置**
+
+```json
+{
+ "channels": {
+ "discord": {
+ "enabled": true,
+ "token": "YOUR_BOT_TOKEN",
+ "allow_from": ["YOUR_USER_ID"]
+ }
+ }
+}
+```
+
+**5. 邀请 Bot**
+
+* OAuth2 → URL Generator
+* Scopes: `bot`
+* Bot Permissions: `Send Messages`, `Read Message History`
+* 打开生成的邀请链接,将 Bot 添加到服务器
+
+**可选:群组触发模式**
+
+默认情况下 Bot 会回复服务器频道中的所有消息。如需仅在 @提及时回复:
+
+```json
+{
+ "channels": {
+ "discord": {
+ "group_trigger": { "mention_only": true }
+ }
+ }
+}
+```
+
+也可通过关键词前缀触发(如 `!bot`):
+
+```json
+{
+ "channels": {
+ "discord": {
+ "group_trigger": { "prefixes": ["!bot"] }
+ }
+ }
+}
+```
+
+**6. 运行**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+WhatsApp(原生 whatsmeow)
+
+PicoClaw 支持两种 WhatsApp 连接方式:
+
+- **原生(推荐):** 进程内使用 [whatsmeow](https://github.com/tulir/whatsmeow),无需独立 Bridge。设置 `"use_native": true` 并留空 `bridge_url`。首次运行时用 WhatsApp 扫描 QR 码(关联设备)。会话存储在工作区下(如 `workspace/whatsapp/`)。原生渠道为**可选**构建,使用 `-tags whatsapp_native` 编译(如 `make build-whatsapp-native` 或 `go build -tags whatsapp_native ./cmd/...`)。
+- **Bridge:** 连接外部 WebSocket Bridge。设置 `bridge_url`(如 `ws://localhost:3001`),保持 `use_native` 为 false。
+
+**配置(原生)**
+
+```json
+{
+ "channels": {
+ "whatsapp": {
+ "enabled": true,
+ "use_native": true,
+ "session_store_path": "",
+ "allow_from": []
+ }
+ }
+}
+```
+
+如果 `session_store_path` 为空,会话存储在 `/whatsapp/`。运行 `picoclaw gateway`;首次运行时在终端扫描 QR 码(WhatsApp → 关联设备)。
+
+
+
+
+Matrix
+
+**1. 准备 Bot 账号**
+
+* 使用你的 homeserver(如 `https://matrix.org` 或自建)
+* 创建 Bot 用户并获取 access token
+
+**2. 配置**
+
+```json
+{
+ "channels": {
+ "matrix": {
+ "enabled": true,
+ "homeserver": "https://matrix.org",
+ "user_id": "@your-bot:matrix.org",
+ "access_token": "YOUR_MATRIX_ACCESS_TOKEN",
+ "allow_from": []
+ }
+ }
+}
+```
+
+**3. 运行**
+
+```bash
+picoclaw gateway
+```
+
+完整选项(`device_id`、`join_on_invite`、`group_trigger`、`placeholder`、`reasoning_channel_id`)请参考 [Matrix 渠道配置指南](../channels/matrix/README.md)。
+
+
+
+
+QQ
+
+**快速设置(推荐)**
+
+QQ 开放平台提供了一键创建 OpenClaw 兼容机器人的页面:
+
+1. 打开 [QQ 机器人快速创建](https://q.qq.com/qqbot/openclaw/index.html),扫码登录
+2. 机器人自动创建 — 复制 **App ID** 和 **App Secret**
+3. 配置 PicoClaw:
+
+```json
+{
+ "channels": {
+ "qq": {
+ "enabled": true,
+ "app_id": "YOUR_APP_ID",
+ "app_secret": "YOUR_APP_SECRET",
+ "allow_from": []
+ }
+ }
+}
+```
+
+4. 运行 `picoclaw gateway`,打开 QQ 与机器人聊天
+
+> App Secret 仅显示一次,请立即保存 — 再次查看将强制重置。
+>
+> 通过快速创建页面创建的机器人初始仅限创建者使用,不支持群聊。如需启用群聊访问,请在 [QQ 开放平台](https://q.qq.com/) 配置沙箱模式。
+
+**手动设置**
+
+如果你更喜欢手动创建机器人:
+
+* 登录 [QQ 开放平台](https://q.qq.com/) 注册成为开发者
+* 创建 QQ 机器人 — 自定义头像和名称
+* 从机器人设置中复制 **App ID** 和 **App Secret**
+* 按上述方式配置并运行 `picoclaw gateway`
+
+
+
+
+Slack
+
+**1. 创建 Slack App**
+
+* 前往 [Slack API](https://api.slack.com/apps) 创建新应用
+* 在 **OAuth & Permissions** 中添加 Bot 权限范围:`chat:write`、`app_mentions:read`、`im:history`、`im:read`、`im:write`
+* 将应用安装到你的工作区
+* 复制 **Bot Token**(`xoxb-...`)和 **App-Level Token**(`xapp-...`,启用 Socket Mode 后获取)
+
+**2. 配置**
+
+```json
+{
+ "channels": {
+ "slack": {
+ "enabled": true,
+ "bot_token": "xoxb-YOUR-BOT-TOKEN",
+ "app_token": "xapp-YOUR-APP-TOKEN",
+ "allow_from": []
+ }
+ }
+}
+```
+
+**3. 运行**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+IRC
+
+**1. 配置**
+
+```json
+{
+ "channels": {
+ "irc": {
+ "enabled": true,
+ "server": "irc.libera.chat:6697",
+ "tls": true,
+ "nick": "picoclaw-bot",
+ "channels": ["#your-channel"],
+ "password": "",
+ "allow_from": []
+ }
+ }
+}
+```
+
+可选:`nickserv_password` 用于 NickServ 认证,`sasl_user`/`sasl_password` 用于 SASL 认证。
+
+**2. 运行**
+
+```bash
+picoclaw gateway
+```
+
+Bot 将连接到 IRC 服务器并加入指定的频道。
+
+
+
+
+钉钉 (DingTalk)
+
+**1. 创建 Bot**
+
+* 前往 [开放平台](https://open.dingtalk.com/)
+* 创建内部应用
+* 复制 Client ID 和 Client Secret
+
+**2. 配置**
+
+```json
+{
+ "channels": {
+ "dingtalk": {
+ "enabled": true,
+ "client_id": "YOUR_CLIENT_ID",
+ "client_secret": "YOUR_CLIENT_SECRET",
+ "allow_from": []
+ }
+ }
+}
+```
+
+> `allow_from` 留空表示允许所有用户,或指定钉钉用户 ID 限制访问。
+
+**3. 运行**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+LINE
+
+**1. 创建 LINE Official Account**
+
+- 前往 [LINE Developers Console](https://developers.line.biz/)
+- 创建 Provider → 创建 Messaging API Channel
+- 复制 **Channel Secret** 和 **Channel Access Token**
+
+**2. 配置**
+
+```json
+{
+ "channels": {
+ "line": {
+ "enabled": true,
+ "channel_secret": "YOUR_CHANNEL_SECRET",
+ "channel_access_token": "YOUR_CHANNEL_ACCESS_TOKEN",
+ "webhook_path": "/webhook/line",
+ "allow_from": []
+ }
+ }
+}
+```
+
+> LINE Webhook 挂载在共享 Gateway 服务器上(`gateway.host`:`gateway.port`,默认 `127.0.0.1:18790`)。
+
+**3. 设置 Webhook URL**
+
+LINE 要求 HTTPS Webhook。使用反向代理或隧道:
+
+```bash
+# 示例:使用 ngrok(Gateway 默认端口 18790)
+ngrok http 18790
+```
+
+然后在 LINE Developers Console 中将 Webhook URL 设置为 `https://your-domain/webhook/line` 并启用 **Use webhook**。
+
+**4. 运行**
+
+```bash
+picoclaw gateway
+```
+
+> 在群聊中,Bot 仅在被 @提及时回复。回复会引用原始消息。
+
+
+
+
+飞书 (Feishu)
+
+PicoClaw 通过 WebSocket/SDK 模式连接飞书 — 无需公网 Webhook URL 或回调服务器。
+
+**1. 创建应用**
+
+* 前往 [飞书开放平台](https://open.feishu.cn/) 创建应用
+* 在应用设置中启用 **机器人** 能力
+* 创建版本并发布应用(应用必须发布后才能生效)
+* 复制 **App ID**(以 `cli_` 开头)和 **App Secret**
+
+**2. 配置**
+
+```json
+{
+ "channels": {
+ "feishu": {
+ "enabled": true,
+ "app_id": "cli_xxx",
+ "app_secret": "YOUR_APP_SECRET",
+ "allow_from": []
+ }
+ }
+}
+```
+
+可选:`encrypt_key` 和 `verification_token` 用于事件加密(生产环境推荐)。
+
+**3. 运行并聊天**
+
+```bash
+picoclaw gateway
+```
+
+打开飞书,搜索你的机器人名称即可开始聊天。也可以将机器人添加到群组 — 使用 `group_trigger.mention_only: true` 设置为仅在 @提及时回复。
+
+完整选项请参考 [飞书渠道配置指南](../channels/feishu/README.zh.md)。
+
+
+
+
+企业微信 (WeCom)
+
+PicoClaw 支持三种企业微信集成方式:
+
+**方式 1: 群机器人 (Bot)** — 设置简单,支持群聊
+**方式 2: 自建应用 (App)** — 功能更多,支持主动推送,仅私聊
+**方式 3: 智能机器人 (AI Bot)** — 官方 AI Bot,流式回复,支持群聊和私聊
+
+详细设置请参考 [企业微信 AI Bot 配置指南](../channels/wecom/wecom_aibot/README.zh.md)。
+
+**快速设置 — 群机器人:**
+
+**1. 创建 Bot**
+
+* 企业微信管理后台 → 群聊 → 添加群机器人
+* 复制 Webhook URL(格式:`https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=xxx`)
+
+**2. 配置**
+
+```json
+{
+ "channels": {
+ "wecom": {
+ "enabled": true,
+ "token": "YOUR_TOKEN",
+ "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
+ "webhook_url": "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=YOUR_KEY",
+ "webhook_path": "/webhook/wecom",
+ "allow_from": []
+ }
+ }
+}
+```
+
+> WeCom Webhook 挂载在共享 Gateway 服务器上(`gateway.host`:`gateway.port`,默认 `127.0.0.1:18790`)。
+
+**快速设置 — 自建应用:**
+
+**1. 创建应用**
+
+* 企业微信管理后台 → 应用管理 → 创建应用
+* 复制 **AgentId** 和 **Secret**
+* 前往"我的企业"页面,复制 **CorpID**
+
+**2. 配置接收消息**
+
+* 在应用详情中,点击"接收消息" → "设置 API"
+* 设置 URL 为 `http://your-server:18790/webhook/wecom-app`
+* 生成 **Token** 和 **EncodingAESKey**
+
+**3. 配置**
+
+```json
+{
+ "channels": {
+ "wecom_app": {
+ "enabled": true,
+ "corp_id": "wwxxxxxxxxxxxxxxxx",
+ "corp_secret": "YOUR_CORP_SECRET",
+ "agent_id": 1000002,
+ "token": "YOUR_TOKEN",
+ "encoding_aes_key": "YOUR_ENCODING_AES_KEY",
+ "webhook_path": "/webhook/wecom-app",
+ "allow_from": []
+ }
+ }
+}
+```
+
+**4. 运行**
+
+```bash
+picoclaw gateway
+```
+
+> **注意**: WeCom Webhook 回调挂载在 Gateway 端口(默认 18790)。使用反向代理配置 HTTPS。
+
+**快速设置 — 智能机器人 (AI Bot):**
+
+**1. 创建 AI Bot**
+
+* 企业微信管理后台 → 应用管理 → AI Bot
+* 在 AI Bot 设置中配置回调 URL:`http://your-server:18790/webhook/wecom-aibot`
+* 复制 **Token** 并点击"随机生成" **EncodingAESKey**
+
+**2. 配置**
+
+```json
+{
+ "channels": {
+ "wecom_aibot": {
+ "enabled": true,
+ "token": "YOUR_TOKEN",
+ "encoding_aes_key": "YOUR_43_CHAR_ENCODING_AES_KEY",
+ "webhook_path": "/webhook/wecom-aibot",
+ "allow_from": [],
+ "welcome_message": "你好!有什么可以帮你的?",
+ "processing_message": "⏳ Processing, please wait. The results will be sent shortly."
+ }
+ }
+}
+```
+
+**3. 运行**
+
+```bash
+picoclaw gateway
+```
+
+> **注意**: 企业微信 AI Bot 使用流式拉取协议,无回复超时问题。长任务(>30 秒)会自动切换到 `response_url` 推送投递。
+
+
+
+
+OneBot(通过 OneBot 协议连接 QQ)
+
+OneBot 是 QQ 机器人的开放协议。PicoClaw 通过 WebSocket 连接任何 OneBot v11 兼容实现(如 [Lagrange](https://github.com/LagrangeDev/Lagrange.Core)、[NapCat](https://github.com/NapNeko/NapCatQQ))。
+
+**1. 设置 OneBot 实现**
+
+安装并运行 OneBot v11 兼容的 QQ 机器人框架,启用其 WebSocket 服务器。
+
+**2. 配置**
+
+```json
+{
+ "channels": {
+ "onebot": {
+ "enabled": true,
+ "ws_url": "ws://127.0.0.1:8080",
+ "access_token": "",
+ "allow_from": []
+ }
+ }
+}
+```
+
+| 字段 | 说明 |
+|------|------|
+| `ws_url` | OneBot 实现的 WebSocket URL |
+| `access_token` | 认证用的访问令牌(如果在 OneBot 中配置了的话) |
+| `reconnect_interval` | 重连间隔(秒)(默认:5) |
+
+**3. 运行**
+
+```bash
+picoclaw gateway
+```
+
+
+
+
+MaixCam
+
+专为 Sipeed AI 摄像头硬件设计的集成通道。
+
+```json
+{
+ "channels": {
+ "maixcam": {
+ "enabled": true
+ }
+ }
+}
+```
+
+```bash
+picoclaw gateway
+```
+
+
diff --git a/docs/zh/configuration.md b/docs/zh/configuration.md
new file mode 100644
index 000000000..68fb1fd1a
--- /dev/null
+++ b/docs/zh/configuration.md
@@ -0,0 +1,258 @@
+# ⚙️ 配置指南
+
+> 返回 [README](../../README.zh.md)
+
+## ⚙️ 配置详解
+
+配置文件路径: `~/.picoclaw/config.json`
+
+### 环境变量
+
+你可以使用环境变量覆盖默认路径。这对于便携安装、容器化部署或将 picoclaw 作为系统服务运行非常有用。这些变量是独立的,控制不同的路径。
+
+| 变量 | 描述 | 默认路径 |
+|-------------------|-----------------------------------------------------------------------------------------------------------------------------------------|---------------------------|
+| `PICOCLAW_CONFIG` | 覆盖配置文件的路径。这直接告诉 picoclaw 加载哪个 `config.json`,忽略所有其他位置。 | `~/.picoclaw/config.json` |
+| `PICOCLAW_HOME` | 覆盖 picoclaw 数据根目录。这会更改 `workspace` 和其他数据目录的默认位置。 | `~/.picoclaw` |
+
+**示例:**
+
+```bash
+# 使用特定的配置文件运行 picoclaw
+# 工作区路径将从该配置文件中读取
+PICOCLAW_CONFIG=/etc/picoclaw/production.json picoclaw gateway
+
+# 在 /opt/picoclaw 中存储所有数据运行 picoclaw
+# 配置将从默认的 ~/.picoclaw/config.json 加载
+# 工作区将在 /opt/picoclaw/workspace 创建
+PICOCLAW_HOME=/opt/picoclaw picoclaw agent
+
+# 同时使用两者进行完全自定义设置
+PICOCLAW_HOME=/srv/picoclaw PICOCLAW_CONFIG=/srv/picoclaw/main.json picoclaw gateway
+```
+
+### 工作区布局 (Workspace Layout)
+
+PicoClaw 将数据存储在您配置的工作区中(默认:`~/.picoclaw/workspace`):
+
+```
+~/.picoclaw/workspace/
+├── sessions/ # 对话会话和历史
+├── memory/ # 长期记忆 (MEMORY.md)
+├── state/ # 持久化状态 (最后一次频道等)
+├── cron/ # 定时任务数据库
+├── skills/ # 自定义技能
+├── AGENT.md # Agent 行为指南
+├── HEARTBEAT.md # 周期性任务提示词 (每 30 分钟检查一次)
+├── IDENTITY.md # Agent 身份设定
+├── SOUL.md # Agent 灵魂/性格
+└── USER.md # 用户偏好
+```
+
+> **提示:** 对 `AGENT.md`、`SOUL.md`、`USER.md` 和 `memory/MEMORY.md` 的修改会通过文件修改时间(mtime)在运行时自动检测。**无需重启 gateway**,Agent 将在下一次请求时自动加载最新内容。
+
+### 技能来源 (Skill Sources)
+
+默认情况下,技能会按以下顺序加载:
+
+1. `~/.picoclaw/workspace/skills`(工作区)
+2. `~/.picoclaw/skills`(全局)
+3. `<构建时嵌入路径>/skills`(内置)
+
+在高级/测试场景下,可通过以下环境变量覆盖内置技能目录:
+
+```bash
+export PICOCLAW_BUILTIN_SKILLS=/path/to/skills
+```
+
+### 统一命令执行策略
+
+- 通用斜杠命令通过 `pkg/agent/loop.go` 中的 `commands.Executor` 统一执行。
+- Channel 适配器不再在本地消费通用命令;它们只负责把入站文本转发到 bus/agent 路径。Telegram 仍会在启动时自动注册其支持的命令菜单。
+- 未注册的斜杠命令(例如 `/foo`)会透传给 LLM 按普通输入处理。
+- 已注册但当前 channel 不支持的命令(例如 WhatsApp 上的 `/show`)会返回明确的用户可见错误,并停止后续处理。
+
+### 🔒 安全沙箱 (Security Sandbox)
+
+PicoClaw 默认在沙箱环境中运行。Agent 只能访问配置的工作区内的文件和执行命令。
+
+#### 默认配置
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "workspace": "~/.picoclaw/workspace",
+ "restrict_to_workspace": true
+ }
+ }
+}
+```
+
+| 选项 | 默认值 | 描述 |
+| ----------------------- | ----------------------- | ----------------------------- |
+| `workspace` | `~/.picoclaw/workspace` | Agent 的工作目录 |
+| `restrict_to_workspace` | `true` | 限制文件/命令访问在工作区内 |
+
+#### 受保护的工具
+
+当 `restrict_to_workspace: true` 时,以下工具会被沙箱化:
+
+| 工具 | 功能 | 限制 |
+| ------------- | ------------ | ------------------------------ |
+| `read_file` | 读取文件 | 仅限工作区内的文件 |
+| `write_file` | 写入文件 | 仅限工作区内的文件 |
+| `list_dir` | 列出目录 | 仅限工作区内的目录 |
+| `edit_file` | 编辑文件 | 仅限工作区内的文件 |
+| `append_file` | 追加文件 | 仅限工作区内的文件 |
+| `exec` | 执行命令 | 命令路径必须在工作区内 |
+
+#### 额外的 Exec 保护
+
+即使 `restrict_to_workspace: false`,`exec` 工具也会阻止以下危险命令:
+
+* `rm -rf`、`del /f`、`rmdir /s` — 批量删除
+* `format`、`mkfs`、`diskpart` — 磁盘格式化
+* `dd if=` — 磁盘镜像
+* 写入 `/dev/sd[a-z]` — 直接磁盘写入
+* `shutdown`、`reboot`、`poweroff` — 系统关机
+* Fork bomb `:(){ :|:& };:`
+
+### 文件访问控制
+
+| 配置键 | 类型 | 默认值 | 描述 |
+|--------|------|--------|------|
+| `tools.allow_read_paths` | string[] | `[]` | 允许在工作区外读取的额外路径 |
+| `tools.allow_write_paths` | string[] | `[]` | 允许在工作区外写入的额外路径 |
+
+### Exec 安全配置
+
+| 配置键 | 类型 | 默认值 | 描述 |
+|--------|------|--------|------|
+| `tools.exec.allow_remote` | bool | `false` | 允许从远程渠道(Telegram/Discord 等)执行 exec 工具 |
+| `tools.exec.enable_deny_patterns` | bool | `true` | 启用危险命令拦截 |
+| `tools.exec.custom_deny_patterns` | string[] | `[]` | 自定义阻止的正则表达式模式 |
+| `tools.exec.custom_allow_patterns` | string[] | `[]` | 自定义允许的正则表达式模式 |
+
+> **安全提示:** Symlink 保护默认启用——所有文件路径在白名单匹配前都会通过 `filepath.EvalSymlinks` 解析,防止符号链接逃逸攻击。
+
+#### 已知限制:构建工具的子进程
+
+exec 安全守卫仅检查 PicoClaw 直接启动的命令行。它不会递归检查由 `make`、`go run`、`cargo`、`npm run` 或自定义构建脚本等开发工具产生的子进程。
+
+这意味着顶层命令通过初始守卫检查后,仍可以编译或启动其他二进制文件。实际上,应将构建脚本、Makefile、包脚本和生成的二进制文件视为与直接 shell 命令同等级别的可执行代码进行审查。
+
+对于高风险环境:
+
+* 执行前审查构建脚本。
+* 对编译并运行的工作流优先使用审批/手动审查。
+* 如果需要比内置守卫更强的隔离,请在容器或虚拟机中运行 PicoClaw。
+
+#### 错误示例
+
+```
+[ERROR] tool: Tool execution failed
+{tool=exec, error=Command blocked by safety guard (path outside working dir)}
+```
+
+```
+[ERROR] tool: Tool execution failed
+{tool=exec, error=Command blocked by safety guard (dangerous pattern detected)}
+```
+
+#### 禁用限制(安全风险)
+
+如果需要 Agent 访问工作区外的路径:
+
+**方法 1: 配置文件**
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "restrict_to_workspace": false
+ }
+ }
+}
+```
+
+**方法 2: 环境变量**
+
+```bash
+export PICOCLAW_AGENTS_DEFAULTS_RESTRICT_TO_WORKSPACE=false
+```
+
+> ⚠️ **警告**: 禁用此限制将允许 Agent 访问系统上的任何路径。仅在受控环境中谨慎使用。
+
+#### 安全边界一致性
+
+`restrict_to_workspace` 设置在所有执行路径中一致应用:
+
+| 执行路径 | 安全边界 |
+| ---------------- | ---------------------------- |
+| 主 Agent | `restrict_to_workspace` ✅ |
+| 子 Agent / Spawn | 继承相同限制 ✅ |
+| 心跳任务 | 继承相同限制 ✅ |
+
+所有路径共享相同的工作区限制——无法通过子 Agent 或定时任务绕过安全边界。
+
+### 心跳 / 周期性任务 (Heartbeat)
+
+PicoClaw 可以自动执行周期性任务。在工作区创建 `HEARTBEAT.md` 文件:
+
+```markdown
+# Periodic Tasks
+
+- Check my email for important messages
+- Review my calendar for upcoming events
+- Check the weather forecast
+```
+
+Agent 将每隔 30 分钟(可配置)读取此文件,并使用可用工具执行任务。
+
+#### 使用 Spawn 的异步任务
+
+对于耗时较长的任务(网络搜索、API 调用),使用 `spawn` 工具创建一个 **子 Agent (subagent)**:
+
+```markdown
+# Periodic Tasks
+
+## Quick Tasks (respond directly)
+
+- Report current time
+
+## Long Tasks (use spawn for async)
+
+- Search the web for AI news and summarize
+- Check email and report important messages
+```
+
+**关键行为:**
+
+| 特性 | 描述 |
+| ---------------- | ---------------------------------------- |
+| **spawn** | 创建异步子 Agent,不阻塞主心跳进程 |
+| **独立上下文** | 子 Agent 拥有独立上下文,无会话历史 |
+| **message tool** | 子 Agent 通过 message 工具直接与用户通信 |
+| **非阻塞** | spawn 后,心跳继续处理下一个任务 |
+
+**配置:**
+
+```json
+{
+ "heartbeat": {
+ "enabled": true,
+ "interval": 30
+ }
+}
+```
+
+| 选项 | 默认值 | 描述 |
+| ---------- | ------ | ---------------------------- |
+| `enabled` | `true` | 启用/禁用心跳 |
+| `interval` | `30` | 检查间隔,单位分钟 (最小: 5) |
+
+**环境变量:**
+
+- `PICOCLAW_HEARTBEAT_ENABLED=false` 禁用
+- `PICOCLAW_HEARTBEAT_INTERVAL=60` 更改间隔
diff --git a/docs/zh/credential_encryption.md b/docs/zh/credential_encryption.md
new file mode 100644
index 000000000..2105e4307
--- /dev/null
+++ b/docs/zh/credential_encryption.md
@@ -0,0 +1,158 @@
+> 返回 [README](../../README.zh.md)
+
+# 凭据加密
+
+PicoClaw 支持对 `model_list` 配置条目中的 `api_key` 值进行加密。
+加密后的密钥以 `enc://` 字符串形式存储,并在启动时自动解密。
+
+---
+
+## 快速开始
+
+**1. 设置密码短语**
+
+```bash
+export PICOCLAW_KEY_PASSPHRASE="your-passphrase"
+```
+
+**2. 加密 API 密钥**
+
+运行 `picoclaw onboard` — 它会提示你输入密码短语并生成 SSH 密钥,
+然后在下一次 `SaveConfig` 调用时自动重新加密配置中所有明文 `api_key` 条目。生成的 `enc://` 值如下所示:
+
+```
+enc://AAAA...base64...
+```
+
+**3. 将输出粘贴到你的配置中**
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "gpt-4o",
+ "model": "openai/gpt-4o",
+ "api_key": "enc://AAAA...base64...",
+ "api_base": "https://api.openai.com/v1"
+ }
+ ]
+}
+```
+
+---
+
+## 支持的 `api_key` 格式
+
+| 格式 | 示例 | 行为 |
+|------|------|------|
+| 明文 | `sk-abc123` | 直接使用 |
+| 文件引用 | `file://openai.key` | 从配置文件所在目录读取内容 |
+| 加密 | `enc://` | 启动时使用 `PICOCLAW_KEY_PASSPHRASE` 解密 |
+| 空值 | `""` | 原样传递(用于 `auth_method: oauth`) |
+
+---
+
+## 加密设计
+
+### 密钥派生
+
+加密使用 **HKDF-SHA256**,并以 SSH 私钥作为第二因子。
+
+```
+sshHash = SHA256(ssh_private_key_file_bytes)
+ikm = HMAC-SHA256(key=sshHash, message=passphrase)
+aes_key = HKDF-SHA256(ikm, salt, info="picoclaw-credential-v1", 32 bytes)
+```
+
+### 加密
+
+```
+AES-256-GCM(key=aes_key, nonce=random[12], plaintext=api_key)
+```
+
+### 传输格式
+
+```
+enc://
+```
+
+| 字段 | 大小 | 描述 |
+|------|------|------|
+| `salt` | 16 字节 | 每次加密随机生成;输入 HKDF |
+| `nonce` | 12 字节 | 每次加密随机生成;AES-GCM IV |
+| `ciphertext` | 可变 | AES-256-GCM 密文 + 16 字节认证标签 |
+
+GCM 认证标签会自动附加到密文之后。任何篡改都会导致解密失败并报错,而不是返回损坏的明文。
+
+### 性能
+
+| 操作 | 耗时 (ARM Cortex-A) |
+|------|---------------------|
+| 密钥派生 (HKDF) | < 1 ms |
+| AES-256-GCM 解密 | < 1 ms |
+| **启动总开销** | **每个密钥 < 2 ms** |
+
+---
+
+## 使用 SSH 密钥的双因子安全
+
+当提供 SSH 私钥时,破解加密需要**同时具备**:
+
+1. **密码短语** (`PICOCLAW_KEY_PASSPHRASE`)
+2. **SSH 私钥文件**
+
+这意味着仅泄露配置文件不足以恢复 API 密钥,即使密码短语较弱也是如此。SSH 密钥贡献 256 位熵(Ed25519),与密码短语强度无关。
+
+### 威胁模型
+
+| 攻击者拥有 | 能否解密? |
+|------------|-----------|
+| 仅配置文件 | 否 — 需要密码短语 + SSH 密钥 |
+| 仅 SSH 密钥 | 否 — 需要密码短语 |
+| 仅密码短语 | 否 — 需要 SSH 密钥 |
+| 配置文件 + SSH 密钥 + 密码短语 | 是 — 完全泄露 |
+
+---
+
+## 环境变量
+
+| 变量 | 是否必需 | 描述 |
+|------|----------|------|
+| `PICOCLAW_KEY_PASSPHRASE` | 是(用于 `enc://`) | 用于密钥派生的密码短语 |
+| `PICOCLAW_SSH_KEY_PATH` | 否 | SSH 私钥路径。如未设置,自动从 `~/.ssh/picoclaw_ed25519.key` 检测 |
+
+### SSH 密钥自动检测
+
+如果未设置 `PICOCLAW_SSH_KEY_PATH`,PicoClaw 会查找专用密钥:
+
+```
+~/.ssh/picoclaw_ed25519.key
+```
+
+此专用文件避免与用户现有的 SSH 密钥冲突。
+运行 `picoclaw onboard` 可自动生成该密钥。
+
+`os.UserHomeDir()` 用于跨平台主目录解析(在 Windows 上读取 `USERPROFILE`,在 Unix/macOS 上读取 `HOME`)。
+
+> **注意:** SSH 密钥文件是凭据加密的必要条件。如果未找到密钥且未设置 `PICOCLAW_SSH_KEY_PATH`,加密/解密将失败。运行 `picoclaw onboard` 可自动生成密钥。
+
+---
+
+## 迁移
+
+由于唯一的密钥材料是 `PICOCLAW_KEY_PASSPHRASE` 和 SSH 私钥文件,迁移非常简单:
+
+1. 将配置文件复制到新机器。
+2. 将 `PICOCLAW_KEY_PASSPHRASE` 设置为相同的值。
+3. 将 SSH 私钥文件复制到相同路径(或将 `PICOCLAW_SSH_KEY_PATH` 设置为新位置)。
+
+无需重新加密。
+
+---
+
+## 安全注意事项
+
+- **密码短语和 SSH 密钥都是必需的。** SSH 密钥作为第二因子 — 没有它,加密/解密将失败。如果密钥不存在,运行 `picoclaw onboard` 生成。
+- **SSH 密钥在运行时为只读。** PicoClaw 不会写入或修改 SSH 密钥文件。
+- **仍然支持明文密钥。** 不使用 `enc://` 的现有配置不受影响。
+- **`enc://` 格式通过版本控制**,通过 HKDF `info` 字段(`picoclaw-credential-v1`)实现,允许未来升级算法而不破坏现有加密值。
diff --git a/docs/zh/debug.md b/docs/zh/debug.md
new file mode 100644
index 000000000..e7f20d777
--- /dev/null
+++ b/docs/zh/debug.md
@@ -0,0 +1,36 @@
+# 调试 PicoClaw
+
+> 返回 [README](../../README.zh.md)
+
+PicoClaw 在处理每一个请求时,都会在后台执行多个复杂的交互操作——从消息路由和复杂度评估,到工具执行和模型故障适配。能够准确地看到正在发生什么至关重要,这不仅有助于排查潜在问题,也有助于真正理解代理的运作方式。
+
+## 以调试模式启动 PicoClaw
+
+要获取代理运行的详细信息(LLM 请求、工具调用、消息路由),可以使用调试标志启动 PicoClaw 网关:
+
+```bash
+picoclaw gateway --debug
+# or
+picoclaw gateway -d
+```
+
+在此模式下,系统会对日志进行详细格式化,并显示系统提示词和工具执行结果的预览。
+
+## 禁用日志截断(完整日志)
+
+默认情况下,PicoClaw 会在调试日志中截断过长的字符串(例如*系统提示词*或大型 JSON 输出结果),以保持控制台的可读性。
+
+如果你需要检查某个命令的完整输出,或发送给 LLM 模型的确切载荷,可以使用 `--no-truncate` 标志。
+
+**注意:** 此标志*仅*在与 `--debug` 模式组合使用时有效。
+
+```bash
+picoclaw gateway --debug --no-truncate
+
+```
+
+当此标志激活时,全局截断功能将被禁用。这在以下场景中非常有用:
+
+* 验证发送给提供商的消息的确切语法。
+* 读取 `exec`、`web_fetch` 或 `read_file` 等工具的完整输出。
+* 调试保存在内存中的会话历史。
diff --git a/docs/zh/docker.md b/docs/zh/docker.md
new file mode 100644
index 000000000..10bc46544
--- /dev/null
+++ b/docs/zh/docker.md
@@ -0,0 +1,169 @@
+# 🐳 Docker 与快速开始
+
+> 返回 [README](../../README.zh.md)
+
+## 🐳 Docker Compose
+
+您也可以使用 Docker Compose 运行 PicoClaw,无需在本地安装任何环境。
+
+```bash
+# 1. 克隆仓库
+git clone https://github.com/sipeed/picoclaw.git
+cd picoclaw
+
+# 2. 首次运行 — 自动生成 docker/data/config.json 后退出
+# (仅在 config.json 和 workspace/ 都不存在时触发)
+docker compose -f docker/docker-compose.yml --profile gateway up
+# 容器打印 "First-run setup complete." 后自动停止
+
+# 3. 填写 API Key 等配置
+vim docker/data/config.json # 设置 provider API key、Bot Token 等
+
+# 4. 正式启动
+docker compose -f docker/docker-compose.yml --profile gateway up -d
+```
+
+> [!TIP]
+> **Docker 用户**: 默认情况下, Gateway 监听 `127.0.0.1`,该端口不会暴露到容器外。如果需要通过端口映射访问健康检查接口,请在环境变量中设置 `PICOCLAW_GATEWAY_HOST=0.0.0.0` 或修改 `config.json`。
+
+```bash
+# 5. 查看日志
+docker compose -f docker/docker-compose.yml logs -f picoclaw-gateway
+
+# 6. 停止
+docker compose -f docker/docker-compose.yml --profile gateway down
+```
+
+### Launcher 模式 (Web 控制台)
+
+`launcher` 镜像包含所有三个二进制文件(`picoclaw`、`picoclaw-launcher`、`picoclaw-launcher-tui`),默认启动 Web 控制台,提供基于浏览器的配置和聊天界面。
+
+```bash
+docker compose -f docker/docker-compose.yml --profile launcher up -d
+```
+
+在浏览器中打开 http://localhost:18800。Launcher 会自动管理 Gateway 进程。
+
+> [!WARNING]
+> Web 控制台尚不支持身份验证。请勿将其暴露到公网。
+
+### Agent 模式 (一次性运行)
+
+```bash
+# 提问
+docker compose -f docker/docker-compose.yml run --rm picoclaw-agent -m "2+2 等于几?"
+
+# 交互模式
+docker compose -f docker/docker-compose.yml run --rm picoclaw-agent
+```
+
+### 更新镜像
+
+```bash
+docker compose -f docker/docker-compose.yml pull
+docker compose -f docker/docker-compose.yml --profile gateway up -d
+```
+
+---
+
+## 🚀 快速开始
+
+> [!TIP]
+> 在 `~/.picoclaw/config.json` 中设置您的 API Key。获取 API Key: [火山引擎 (CodingPlan)](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) (LLM) · [OpenRouter](https://openrouter.ai/keys) (LLM) · [Zhipu (智谱)](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) (LLM)。网络搜索是 **可选的** — 获取免费的 [Tavily API](https://tavily.com) (每月 1000 次免费查询) 或 [Brave Search API](https://brave.com/search/api) (每月 2000 次免费查询)。
+
+**1. 初始化 (Initialize)**
+
+```bash
+picoclaw onboard
+```
+
+**2. 配置 (Configure)** (`~/.picoclaw/config.json`)
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "workspace": "~/.picoclaw/workspace",
+ "model_name": "gpt-5.4",
+ "max_tokens": 8192,
+ "temperature": 0.7,
+ "max_tool_iterations": 20
+ }
+ },
+ "model_list": [
+ {
+ "model_name": "ark-code-latest",
+ "model": "volcengine/ark-code-latest",
+ "api_key": "sk-your-api-key",
+ "api_base":"https://ark.cn-beijing.volces.com/api/coding/v3"
+ },
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_key": "your-api-key",
+ "request_timeout": 300
+ },
+ {
+ "model_name": "claude-sonnet-4.6",
+ "model": "anthropic/claude-sonnet-4.6",
+ "api_key": "your-anthropic-key"
+ }
+ ],
+ "tools": {
+ "web": {
+ "enabled": true,
+ "fetch_limit_bytes": 10485760,
+ "format": "plaintext",
+ "brave": {
+ "enabled": false,
+ "api_key": "YOUR_BRAVE_API_KEY",
+ "max_results": 5
+ },
+ "tavily": {
+ "enabled": false,
+ "api_key": "YOUR_TAVILY_API_KEY",
+ "max_results": 5
+ },
+ "duckduckgo": {
+ "enabled": true,
+ "max_results": 5
+ },
+ "perplexity": {
+ "enabled": false,
+ "api_key": "YOUR_PERPLEXITY_API_KEY",
+ "max_results": 5
+ },
+ "searxng": {
+ "enabled": false,
+ "base_url": "http://your-searxng-instance:8888",
+ "max_results": 5
+ }
+ }
+ }
+}
+```
+
+> **新功能**: `model_list` 配置格式支持零代码添加 provider。详见[模型配置](providers.md#模型配置-model_list)章节。
+> `request_timeout` 为可选项,单位为秒。若省略或设置为 `<= 0`,PicoClaw 使用默认超时(120 秒)。
+
+**3. 获取 API Key**
+
+* **LLM 提供商**: [OpenRouter](https://openrouter.ai/keys) · [Zhipu](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) · [Anthropic](https://console.anthropic.com) · [OpenAI](https://platform.openai.com) · [Gemini](https://aistudio.google.com/api-keys)
+* **网络搜索** (可选):
+ * [Brave Search](https://brave.com/search/api) - 付费 ($5/1000 次查询,约 $5-6/月)
+ * [Perplexity](https://www.perplexity.ai) - AI 驱动的搜索与聊天界面
+ * [SearXNG](https://github.com/searxng/searxng) - 自建元搜索引擎(免费,无需 API Key)
+ * [Tavily](https://tavily.com) - 专为 AI Agent 优化 (1000 请求/月)
+ * DuckDuckGo - 内置回退(无需 API Key)
+
+> **注意**: 完整的配置模板请参考 `config.example.json`。
+
+**4. 对话 (Chat)**
+
+```bash
+picoclaw agent -m "2+2 等于几?"
+```
+
+就是这样!您在 2 分钟内就拥有了一个可工作的 AI 助手。
+
+---
diff --git a/docs/zh/hardware-compatibility.md b/docs/zh/hardware-compatibility.md
new file mode 100644
index 000000000..66bd08072
--- /dev/null
+++ b/docs/zh/hardware-compatibility.md
@@ -0,0 +1,152 @@
+> 返回 [README](../../README.zh.md)
+
+# 🖥️ PicoClaw 硬件兼容性列表
+
+PicoClaw 几乎可以在任何 Linux 设备上运行。本页面记录了已验证的芯片、产品和开发板。
+
+**你的硬件不在列表中?** 提交 PR 来添加它!欢迎硬件厂商贡献和联合推广。
+
+---
+
+## 1. 已验证的芯片支持
+
+### x86
+
+| 厂商 | 芯片 | 备注 |
+|------|------|------|
+| Intel | Any x86 CPU (i386+) | 所有桌面/服务器/笔记本处理器 |
+| AMD | Any x86 CPU | 所有桌面/服务器/笔记本处理器 |
+
+### ARM
+
+| 子架构 | 典型芯片 | 备注 |
+|--------|----------|------|
+| ARMv6 | [BCM2835](https://www.raspberrypi.com/documentation/computers/processors.html#bcm2835) (Raspberry Pi 1/Zero) | 单核 ARM1176JZF-S |
+| ARMv7 | [Allwinner V3s](https://linux-sunxi.org/V3s) | 单核 Cortex-A7,用于 LicheePi Zero |
+| ARM64 | [Allwinner H618](https://linux-sunxi.org/H618) | 四核 Cortex-A53,用于 Orange Pi Zero 3 |
+| ARM64 | [BCM2711](https://www.raspberrypi.com/documentation/computers/processors.html#bcm2711) (Raspberry Pi 4) | 四核 Cortex-A72 |
+| ARM64 | [BCM2712](https://www.raspberrypi.com/documentation/computers/processors.html#bcm2712) (Raspberry Pi 5) | 四核 Cortex-A76 |
+| ARM64 | [AX630C](https://www.axera-tech.com/) (爱芯元智) | 双核 Cortex-A53 + NPU,用于 NanoKVM-Pro / MaixCAM2 |
+
+### RISC-V (riscv64)
+
+| 厂商 | 芯片 | 核心 | 备注 |
+|------|------|------|------|
+| [SOPHGO (算能)](https://www.sophgo.com/) | SG2002 | C906 @ 1GHz | 256MB DDR3 片上内存,用于 LicheeRV-Nano / NanoKVM / MaixCAM |
+| [Allwinner (全志)](https://www.allwinnertech.com/) | V861 | Dual C907 | 128MB DDR3L 片上内存,1 TOPS NPU,4K AI 摄像头 SiP |
+| [Allwinner (全志)](https://www.allwinnertech.com/) | V881 | C907 | RISC-V AI 摄像头系列 |
+| [Arterytek (匠芯创)](https://www.arterytek.com/) | D213 | RISC-V | 用于 HaaS506-LD1 工业 RTU |
+| [SpacemiT (进迭)](https://www.spacemit.com/) | K1 | 8x X60 @ 1.8GHz | 用于 Milk-V Jupiter, BananaPi BPI-F3 |
+| [SpacemiT (进迭)](https://www.spacemit.com/) | K3 | 8x X100 @ 2.5GHz | 符合 RVA23 规范,1024 位 RVV,FP8 AI 推理 |
+| [Zhihe (知合)](https://www.zhihe-tech.com/) | A210 | High-perf RISC-V | 8 核,16MB L3 缓存,桌面级 |
+| [Canaan (嘉楠)](https://www.canaan-creative.com/) | K230 | Dual C908 @ 1.6GHz | 6 TOPS KPU,用于 CanMV-K230 |
+
+### MIPS
+
+| 厂商 | 芯片 | 备注 |
+|------|------|------|
+| MediaTek | [MT7620](https://www.mediatek.com/products/home-networking/mt7620) | MIPS24KEc @ 580MHz,用于许多 OpenWrt 路由器(如小米路由器 3G) |
+
+### LoongArch (loong64)
+
+| 厂商 | 芯片 | 备注 |
+|------|------|------|
+| [Loongson (龙芯)](https://www.loongson.cn/) | 3A5000 | 四核 LA464 @ 2.5GHz,桌面/工作站 |
+| [Loongson (龙芯)](https://www.loongson.cn/) | 3A6000 | 四核 4C/8T @ 2.5GHz,IPC 可与 Intel 第十代相媲美 |
+| [Loongson (龙芯)](https://www.loongson.cn/) | 2K1000LA | 双核 @ 1GHz,工业/物联网应用 |
+
+---
+
+## 2. 已验证的产品(按发布日期排列)
+
+已通过 PicoClaw 测试的消费产品、路由器和工业设备。
+
+| 年份 | 产品 | 架构 | SoC | 内存 | 类别 |
+|------|------|------|-----|------|------|
+| 2009 | Nokia N900 | ARM (A8) | OMAP3430 | 256MB | 智能手机 |
+| 2012 | Samsung Galaxy Note 10.1 (N8000) | ARM (A9) | Exynos 4412 | 2GB | 平板电脑 |
+| 2016 | Xiaomi Router 3G (小米路由器3G) | MIPS | MT7620 | 256MB | 路由器 (OpenWrt) |
+| 2018 | Phicomm N1 (斐讯N1) | ARM64 (A53) | S905D | 2GB | 电视盒子 / 家庭服务器 |
+| 2019 | Xiaomi AI Speaker (小爱音箱) | ARM64 (A53) | — | 256MB | 智能音箱 |
+| 2024 | [NanoKVM](https://wiki.sipeed.com/hardware/en/kvm/NanoKVM/introduction.html) | RISC-V | SG2002 | 256MB | IP-KVM |
+| 2025 | HaaS506-LD1 | RISC-V | D213 | 128MB | 工业 RTU |
+| 2025 | [NanoKVM-Pro](https://wiki.sipeed.com/hardware/en/kvm/NanoKVM_Pro/introduction.html) | ARM64 (A53) | AX630C | 1GB | 专业 IP-KVM |
+| 2026 | [MaixCAM2](https://wiki.sipeed.com/hardware/en/maixcam/index.html) | ARM64 (A53) | AX630C | 1/4GB | 4K AI 摄像头 |
+
+---
+
+## 3. 已验证的开发板(按发布日期排列)
+
+| 年份 | 开发板 | 架构 | SoC | 内存 | 购买链接 |
+|------|--------|------|-----|------|----------|
+| 2012 | [Raspberry Pi 1 Model B](https://www.raspberrypi.com/products/) | ARMv6 | BCM2835 | 512MB | — |
+| 2015 | [Raspberry Pi 2 Model B](https://www.raspberrypi.com/products/raspberry-pi-2-model-b/) | ARMv7 (A7) | BCM2836 | 1GB | — |
+| 2015 | [Raspberry Pi Zero](https://www.raspberrypi.com/products/raspberry-pi-zero/) | ARMv6 | BCM2835 | 512MB | — |
+| 2016 | [Raspberry Pi 3 Model B](https://www.raspberrypi.com/products/raspberry-pi-3-model-b/) | ARM64 (A53) | BCM2837 | 1GB | — |
+| 2017 | [LicheePi Zero](https://wiki.sipeed.com/hardware/en/lichee/Zero/Zero.html) | ARMv7 (A7) | Allwinner V3s | 64MB | [Sipeed](https://sipeed.com/) |
+| 2019 | [Raspberry Pi 4 Model B](https://www.raspberrypi.com/products/raspberry-pi-4-model-b/) | ARM64 (A72) | BCM2711 | 1~8GB | [RPi](https://www.raspberrypi.com/) |
+| 2023 | [Raspberry Pi 5](https://www.raspberrypi.com/products/raspberry-pi-5/) | ARM64 (A76) | BCM2712 | 2~8GB | [RPi](https://www.raspberrypi.com/) |
+| 2024 | [LicheeRV-Nano](https://wiki.sipeed.com/hardware/en/lichee/RV_Nano/1_intro.html) | RISC-V | SG2002 | 256MB | [AliExpress](https://www.aliexpress.com/item/1005006519668532.html) |
+| 2024 | [MaixCAM-Pro](https://wiki.sipeed.com/hardware/en/maixcam/index.html) | RISC-V | SG2002 | 256MB | [Sipeed](https://sipeed.com/) |
+| 2024 | [Milk-V Duo 64M](https://milkv.io/docs/duo/getting-started/duo) | RISC-V | CV1800B | 64MB | [Milk-V](https://milkv.io/) |
+| 2024 | [CanMV-K230](https://developer.canaan-creative.com/k230_canmv/en/main/) | RISC-V | K230 | 512MB | [Canaan](https://www.canaan-creative.com/) |
+
+---
+
+## 4. 同样适用于
+
+### Android 手机(通过 Termux)
+
+任何 ARM64 Android 手机(2015 年以后),1GB 以上内存。安装 [Termux](https://github.com/termux/termux-app),使用 `proot` 运行 PicoClaw。
+
+> 参见 [README:在旧 Android 手机上运行](../../README.zh.md#-run-on-old-android-phones) 获取设置说明。
+
+### 桌面 / 服务器 / 云
+
+| 平台 | 备注 |
+|------|------|
+| x86_64 Linux | 原生二进制文件,无依赖 |
+| x86_64 Windows | 原生二进制文件 |
+| macOS (Intel / Apple Silicon) | 原生二进制文件 |
+| Docker (any platform) | `docker compose` 一行命令,参见 [Docker 指南](docker.md) |
+| OpenWrt routers | MIPS/ARM 构建,需要 >32MB 可用内存 |
+| FreeBSD / NetBSD | 提供 x86_64 和 arm64 构建 |
+
+---
+
+## 5. 最低要求
+
+| 资源 | 最低要求 | 推荐配置 |
+|------|----------|----------|
+| 内存 | 10MB 可用 | 32MB 以上可用 |
+| 存储 | 20MB(二进制文件) | 50MB 以上(含工作区) |
+| CPU | 任意(单核 0.6GHz 以上) | — |
+| 操作系统 | Linux (kernel 3.x+) | Linux 5.x+ |
+| 网络 | 必需(用于 LLM API 调用) | 以太网或 WiFi |
+
+---
+
+## 6. 如何测试与贡献
+
+```bash
+# 1. 下载适合你架构的版本
+wget https://github.com/sipeed/picoclaw/releases/latest/download/picoclaw_Linux_arm64.tar.gz
+tar xzf picoclaw_Linux_arm64.tar.gz
+
+# 2. 初始化
+./picoclaw onboard
+
+# 3. 测试
+./picoclaw agent -m "Hello, what board am I running on?"
+```
+
+可用构建版本:`linux-amd64`, `linux-arm64`, `linux-arm`, `linux-riscv64`, `linux-loong64`, `linux-mipsle`
+
+### 添加你的硬件
+
+1. Fork 本仓库
+2. 将你的芯片/产品/开发板添加到相应的表格中
+3. 包含:名称、架构、SoC、内存、年份,以及可用的链接
+4. 提交 PR
+
+硬件厂商:想要添加官方支持或联合推广?请提交 issue 或通过 [Discord](https://discord.gg/V4sAZ9XWpN) 联系我们。
diff --git a/docs/zh/providers.md b/docs/zh/providers.md
new file mode 100644
index 000000000..9092e7dfe
--- /dev/null
+++ b/docs/zh/providers.md
@@ -0,0 +1,427 @@
+# 🔌 提供商与模型配置
+
+> 返回 [README](../../README.zh.md)
+
+### 提供商 (Providers)
+
+> [!NOTE]
+> Groq 通过 Whisper 提供免费的语音转录。如果配置了 Groq,任意渠道的音频消息都将在 Agent 层面自动转录为文字。
+
+| 提供商 | 用途 | 获取 API Key |
+| -------------------- | ---------------------------- | -------------------------------------------------------------------- |
+| `gemini` | LLM (Gemini 直连) | [aistudio.google.com](https://aistudio.google.com) |
+| `zhipu` | LLM (智谱直连) | [bigmodel.cn](https://bigmodel.cn) |
+| `volcengine` | LLM (火山引擎直连) | [volcengine.com](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) |
+| `openrouter` | LLM (推荐,可访问所有模型) | [openrouter.ai](https://openrouter.ai) |
+| `anthropic` | LLM (Claude 直连) | [console.anthropic.com](https://console.anthropic.com) |
+| `openai` | LLM (GPT 直连) | [platform.openai.com](https://platform.openai.com) |
+| `deepseek` | LLM (DeepSeek 直连) | [platform.deepseek.com](https://platform.deepseek.com) |
+| `qwen` | LLM (通义千问) | [dashscope.console.aliyun.com](https://dashscope.console.aliyun.com) |
+| `groq` | LLM + **语音转录** (Whisper) | [console.groq.com](https://console.groq.com) |
+| `cerebras` | LLM (Cerebras 直连) | [cerebras.ai](https://cerebras.ai) |
+| `vivgrid` | LLM (Vivgrid 直连) | [vivgrid.com](https://vivgrid.com) |
+| `moonshot` | LLM (Kimi/Moonshot 直连) | [platform.moonshot.cn](https://platform.moonshot.cn) |
+| `minimax` | LLM (Minimax 直连) | [platform.minimaxi.com](https://platform.minimaxi.com) |
+| `avian` | LLM (Avian 直连) | [avian.io](https://avian.io) |
+| `mistral` | LLM (Mistral 直连) | [console.mistral.ai](https://console.mistral.ai) |
+| `longcat` | LLM (Longcat 直连) | [longcat.ai](https://longcat.ai) |
+| `modelscope` | LLM (ModelScope 直连) | [modelscope.cn](https://modelscope.cn) |
+
+### 模型配置 (model_list)
+
+> **新功能!** PicoClaw 现在采用**以模型为中心**的配置方式。只需使用 `厂商/模型` 格式(如 `zhipu/glm-4.7`)即可添加新的 provider——**无需修改任何代码!**
+
+该设计同时支持**多 Agent 场景**,提供灵活的 Provider 选择:
+
+- **不同 Agent 使用不同 Provider**:每个 Agent 可以使用自己的 LLM provider
+- **模型回退(Fallback)**:配置主模型和备用模型,提高可靠性
+- **负载均衡**:在多个 API 端点之间分配请求
+- **集中化配置**:在一个地方管理所有 provider
+
+#### 📋 所有支持的厂商
+
+| 厂商 | `model` 前缀 | 默认 API Base | 协议 | 获取 API Key |
+| ------------------- | ----------------- | --------------------------------------------------- | --------- | ----------------------------------------------------------------- |
+| **OpenAI** | `openai/` | `https://api.openai.com/v1` | OpenAI | [获取密钥](https://platform.openai.com) |
+| **Anthropic** | `anthropic/` | `https://api.anthropic.com/v1` | Anthropic | [获取密钥](https://console.anthropic.com) |
+| **智谱 AI (GLM)** | `zhipu/` | `https://open.bigmodel.cn/api/paas/v4` | OpenAI | [获取密钥](https://open.bigmodel.cn/usercenter/proj-mgmt/apikeys) |
+| **DeepSeek** | `deepseek/` | `https://api.deepseek.com/v1` | OpenAI | [获取密钥](https://platform.deepseek.com) |
+| **Google Gemini** | `gemini/` | `https://generativelanguage.googleapis.com/v1beta` | OpenAI | [获取密钥](https://aistudio.google.com/api-keys) |
+| **Groq** | `groq/` | `https://api.groq.com/openai/v1` | OpenAI | [获取密钥](https://console.groq.com) |
+| **Moonshot** | `moonshot/` | `https://api.moonshot.cn/v1` | OpenAI | [获取密钥](https://platform.moonshot.cn) |
+| **通义千问 (Qwen)** | `qwen/` | `https://dashscope.aliyuncs.com/compatible-mode/v1` | OpenAI | [获取密钥](https://dashscope.console.aliyun.com) |
+| **NVIDIA** | `nvidia/` | `https://integrate.api.nvidia.com/v1` | OpenAI | [获取密钥](https://build.nvidia.com) |
+| **Ollama** | `ollama/` | `http://localhost:11434/v1` | OpenAI | 本地(无需密钥) |
+| **OpenRouter** | `openrouter/` | `https://openrouter.ai/api/v1` | OpenAI | [获取密钥](https://openrouter.ai/keys) |
+| **LiteLLM Proxy** | `litellm/` | `http://localhost:4000/v1` | OpenAI | 你的 LiteLLM 代理密钥 |
+| **VLLM** | `vllm/` | `http://localhost:8000/v1` | OpenAI | 本地 |
+| **Cerebras** | `cerebras/` | `https://api.cerebras.ai/v1` | OpenAI | [获取密钥](https://cerebras.ai) |
+| **火山引擎(Doubao)** | `volcengine/` | `https://ark.cn-beijing.volces.com/api/v3` | OpenAI | [获取密钥](https://www.volcengine.com/activity/codingplan?utm_campaign=PicoClaw&utm_content=PicoClaw&utm_medium=devrel&utm_source=OWO&utm_term=PicoClaw) |
+| **神算云** | `shengsuanyun/` | `https://router.shengsuanyun.com/api/v1` | OpenAI | - |
+| **BytePlus** | `byteplus/` | `https://ark.ap-southeast.bytepluses.com/api/v3` | OpenAI | [获取密钥](https://www.byteplus.com) |
+| **Vivgrid** | `vivgrid/` | `https://api.vivgrid.com/v1` | OpenAI | [获取密钥](https://vivgrid.com) |
+| **LongCat** | `longcat/` | `https://api.longcat.chat/openai` | OpenAI | [获取密钥](https://longcat.chat/platform) |
+| **ModelScope (魔搭)**| `modelscope/` | `https://api-inference.modelscope.cn/v1` | OpenAI | [获取 Token](https://modelscope.cn/my/tokens) |
+| **Antigravity** | `antigravity/` | Google Cloud | 自定义 | 仅 OAuth |
+| **GitHub Copilot** | `github-copilot/` | `localhost:4321` | gRPC | - |
+
+#### 基础配置示例
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "ark-code-latest",
+ "model": "volcengine/ark-code-latest",
+ "api_key": "sk-your-api-key"
+ },
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_key": "sk-your-openai-key"
+ },
+ {
+ "model_name": "claude-sonnet-4.6",
+ "model": "anthropic/claude-sonnet-4.6",
+ "api_key": "sk-ant-your-key"
+ },
+ {
+ "model_name": "glm-4.7",
+ "model": "zhipu/glm-4.7",
+ "api_key": "your-zhipu-key"
+ }
+ ],
+ "agents": {
+ "defaults": {
+ "model_name": "gpt-5.4"
+ }
+ }
+}
+```
+
+#### 各厂商配置示例
+
+**OpenAI**
+
+```json
+{
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_key": "sk-..."
+}
+```
+
+**火山引擎(Doubao)**
+
+```json
+{
+ "model_name": "ark-code-latest",
+ "model": "volcengine/ark-code-latest",
+ "api_key": "sk-..."
+}
+```
+
+**智谱 AI (GLM)**
+
+```json
+{
+ "model_name": "glm-4.7",
+ "model": "zhipu/glm-4.7",
+ "api_key": "your-key"
+}
+```
+
+**DeepSeek**
+
+```json
+{
+ "model_name": "deepseek-chat",
+ "model": "deepseek/deepseek-chat",
+ "api_key": "sk-..."
+}
+```
+
+**Anthropic (使用 OAuth)**
+
+```json
+{
+ "model_name": "claude-sonnet-4.6",
+ "model": "anthropic/claude-sonnet-4.6",
+ "auth_method": "oauth"
+}
+```
+
+> 运行 `picoclaw auth login --provider anthropic` 来设置 OAuth 凭证。
+
+**Anthropic Messages API(原生格式)**
+
+用于直接访问 Anthropic API 或仅支持 Anthropic 原生消息格式的自定义端点:
+
+```json
+{
+ "model_name": "claude-opus-4-6",
+ "model": "anthropic-messages/claude-opus-4-6",
+ "api_key": "sk-ant-your-key",
+ "api_base": "https://api.anthropic.com"
+}
+```
+
+> 使用 `anthropic-messages` 协议的场景:
+> - 使用仅支持 Anthropic 原生 `/v1/messages` 端点的第三方代理(不支持 OpenAI 兼容的 `/v1/chat/completions`)
+> - 连接到 MiniMax、Synthetic 等需要 Anthropic 原生消息格式的服务
+> - 现有的 `anthropic` 协议返回 404 错误(说明端点不支持 OpenAI 兼容格式)
+>
+> **注意:** `anthropic` 协议使用 OpenAI 兼容格式(`/v1/chat/completions`),而 `anthropic-messages` 使用 Anthropic 原生格式(`/v1/messages`)。请根据端点支持的格式选择。
+
+**Ollama (本地)**
+
+```json
+{
+ "model_name": "llama3",
+ "model": "ollama/llama3"
+}
+```
+
+**自定义代理/API**
+
+```json
+{
+ "model_name": "my-custom-model",
+ "model": "openai/custom-model",
+ "api_base": "https://my-proxy.com/v1",
+ "api_key": "sk-...",
+ "request_timeout": 300
+}
+```
+
+**LiteLLM Proxy**
+
+```json
+{
+ "model_name": "lite-gpt4",
+ "model": "litellm/lite-gpt4",
+ "api_base": "http://localhost:4000/v1",
+ "api_key": "sk-..."
+}
+```
+
+PicoClaw 在发送请求前仅去除外层 `litellm/` 前缀,因此 `litellm/lite-gpt4` 会发送 `lite-gpt4`,而 `litellm/openai/gpt-4o` 会发送 `openai/gpt-4o`。
+
+#### 负载均衡
+
+为同一个模型名称配置多个端点——PicoClaw 会自动在它们之间轮询:
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_base": "https://api1.example.com/v1",
+ "api_key": "sk-key1"
+ },
+ {
+ "model_name": "gpt-5.4",
+ "model": "openai/gpt-5.4",
+ "api_base": "https://api2.example.com/v1",
+ "api_key": "sk-key2"
+ }
+ ]
+}
+```
+
+#### 从旧的 `providers` 配置迁移
+
+旧的 `providers` 配置格式**已弃用**,但为向后兼容仍支持。
+
+**旧配置(已弃用):**
+
+```json
+{
+ "providers": {
+ "zhipu": {
+ "api_key": "your-key",
+ "api_base": "https://open.bigmodel.cn/api/paas/v4"
+ }
+ },
+ "agents": {
+ "defaults": {
+ "provider": "zhipu",
+ "model": "glm-4.7"
+ }
+ }
+}
+```
+
+**新配置(推荐):**
+
+```json
+{
+ "model_list": [
+ {
+ "model_name": "glm-4.7",
+ "model": "zhipu/glm-4.7",
+ "api_key": "your-key"
+ }
+ ],
+ "agents": {
+ "defaults": {
+ "model_name": "glm-4.7"
+ }
+ }
+}
+```
+
+详细的迁移指南请参考 [docs/migration/model-list-migration.md](../migration/model-list-migration.md)。
+
+### Provider 架构
+
+PicoClaw 按协议族路由 Provider:
+
+- OpenAI 兼容协议:OpenRouter、OpenAI 兼容网关、Groq、智谱、vLLM 风格端点。
+- Anthropic 协议:Claude 原生 API 行为。
+- Codex/OAuth 路径:OpenAI OAuth/Token 认证路由。
+
+这使得运行时保持轻量,同时让新的 OpenAI 兼容后端基本只需配置操作(`api_base` + `api_key`)。
+
+
+智谱 (Zhipu) 配置示例
+
+**1. 获取 API key 和 base URL**
+
+- 获取 [API key](https://bigmodel.cn/usercenter/proj-mgmt/apikeys)
+
+**2. 配置**
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "workspace": "~/.picoclaw/workspace",
+ "model_name": "glm-4.7",
+ "max_tokens": 8192,
+ "temperature": 0.7,
+ "max_tool_iterations": 20
+ }
+ },
+ "providers": {
+ "zhipu": {
+ "api_key": "Your API Key",
+ "api_base": "https://open.bigmodel.cn/api/paas/v4"
+ }
+ }
+}
+```
+
+**3. 运行**
+
+```bash
+picoclaw agent -m "你好"
+```
+
+
+
+
+完整配置示例
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "model_name": "anthropic/claude-opus-4-5"
+ }
+ },
+ "session": {
+ "dm_scope": "per-channel-peer"
+ },
+ "providers": {
+ "openrouter": {
+ "api_key": "sk-or-v1-xxx"
+ },
+ "groq": {
+ "api_key": "gsk_xxx"
+ }
+ },
+ "channels": {
+ "telegram": {
+ "enabled": true,
+ "token": "123456:ABC...",
+ "allow_from": ["123456789"]
+ },
+ "discord": {
+ "enabled": true,
+ "token": "",
+ "allow_from": [""]
+ },
+ "whatsapp": {
+ "enabled": false,
+ "bridge_url": "ws://localhost:3001",
+ "use_native": false,
+ "session_store_path": "",
+ "allow_from": []
+ },
+ "feishu": {
+ "enabled": false,
+ "app_id": "cli_xxx",
+ "app_secret": "xxx",
+ "encrypt_key": "",
+ "verification_token": "",
+ "allow_from": []
+ },
+ "qq": {
+ "enabled": false,
+ "app_id": "",
+ "app_secret": "",
+ "allow_from": []
+ }
+ },
+ "tools": {
+ "web": {
+ "brave": {
+ "enabled": false,
+ "api_key": "BSA...",
+ "max_results": 5
+ },
+ "duckduckgo": {
+ "enabled": true,
+ "max_results": 5
+ },
+ "perplexity": {
+ "enabled": false,
+ "api_key": "",
+ "max_results": 5
+ },
+ "searxng": {
+ "enabled": false,
+ "base_url": "http://localhost:8888",
+ "max_results": 5
+ }
+ },
+ "cron": {
+ "exec_timeout_minutes": 5
+ }
+ },
+ "heartbeat": {
+ "enabled": true,
+ "interval": 30
+ }
+}
+```
+
+
+
+---
+
+## 📝 API Key 对比
+
+| 服务 | 价格 | 适用场景 |
+| --- | --- | --- |
+| **OpenRouter** | 免费: 200K tokens/月 | 多模型聚合 (Claude, GPT-4 等) |
+| **火山引擎 CodingPlan** | ¥9.9/首月 | 最适合国内用户,多种 SOTA 模型(豆包、DeepSeek 等) |
+| **智谱 (Zhipu)** | 免费: 200K tokens/月 | 适合中国用户 |
+| **Brave Search** | $5/1000 次查询 | 网络搜索功能 |
+| **SearXNG** | 免费(自建) | 隐私优先的元搜索引擎(70+ 搜索引擎) |
+| **Groq** | 免费额度可用 | 极速推理 (Llama, Mixtral) |
+| **Cerebras** | 免费额度可用 | 极速推理 (Llama, Qwen 等) |
+| **LongCat** | 免费: 最多 5M tokens/天 | 极速推理 |
+| **ModelScope (魔搭)** | 免费: 2000 次请求/天 | 推理 (Qwen, GLM, DeepSeek 等) |
diff --git a/docs/zh/spawn-tasks.md b/docs/zh/spawn-tasks.md
new file mode 100644
index 000000000..781462af2
--- /dev/null
+++ b/docs/zh/spawn-tasks.md
@@ -0,0 +1,70 @@
+# 🔄 异步任务与 Spawn
+
+> 返回 [README](../../README.zh.md)
+
+PicoClaw 通过 `spawn` 工具支持**异步任务执行**。主要由 **Heartbeat(心跳)** 系统使用,在不阻塞主 Agent 循环的情况下运行耗时任务。
+
+## Heartbeat
+
+心跳系统会定期检查 `workspace/HEARTBEAT.md` 中的计划任务。首次运行时会自动生成默认模板,你可以自定义它来定义快速任务(内联处理)和长任务(通过 `spawn` 委派)。
+
+**`HEARTBEAT.md` 示例:**
+
+```markdown
+## Quick Tasks (respond directly)
+
+- Report current time
+
+## Long Tasks (use spawn for async)
+
+- Search the web for AI news and summarize
+- Check email and report important messages
+```
+
+**关键行为:**
+
+| 特性 | 描述 |
+| ---------------- | ---------------------------------------- |
+| **spawn** | 创建异步子 Agent,不阻塞主心跳进程 |
+| **独立上下文** | 子 Agent 拥有独立上下文,无会话历史 |
+| **message tool** | 子 Agent 通过 message 工具直接与用户通信 |
+| **非阻塞** | spawn 后,心跳继续处理下一个任务 |
+
+#### 子 Agent 通信原理
+
+```
+心跳触发 (Heartbeat triggers)
+ ↓
+Agent 读取 HEARTBEAT.md
+ ↓
+对于长任务: spawn 子 Agent
+ ↓ ↓
+继续下一个任务 子 Agent 独立工作
+ ↓ ↓
+所有任务完成 子 Agent 使用 "message" 工具
+ ↓ ↓
+响应 HEARTBEAT_OK 用户直接收到结果
+```
+
+子 Agent 可以访问工具(message, web_search 等),并且无需通过主 Agent 即可独立与用户通信。
+
+**配置:**
+
+```json
+{
+ "heartbeat": {
+ "enabled": true,
+ "interval": 30
+ }
+}
+```
+
+| 选项 | 默认值 | 描述 |
+| ---------- | ------ | ---------------------------- |
+| `enabled` | `true` | 启用/禁用心跳 |
+| `interval` | `30` | 检查间隔,单位分钟 (最小: 5) |
+
+**环境变量:**
+
+- `PICOCLAW_HEARTBEAT_ENABLED=false` 禁用
+- `PICOCLAW_HEARTBEAT_INTERVAL=60` 更改间隔
diff --git a/docs/zh/tools_configuration.md b/docs/zh/tools_configuration.md
new file mode 100644
index 000000000..f13448952
--- /dev/null
+++ b/docs/zh/tools_configuration.md
@@ -0,0 +1,421 @@
+# 🔧 工具配置
+
+> 返回 [README](../../README.zh.md)
+
+PicoClaw 的工具配置位于 `config.json` 的 `tools` 字段中。
+
+## 目录结构
+
+```json
+{
+ "tools": {
+ "web": {
+ ...
+ },
+ "mcp": {
+ ...
+ },
+ "exec": {
+ ...
+ },
+ "cron": {
+ ...
+ },
+ "skills": {
+ ...
+ }
+ }
+}
+```
+
+## Web 工具
+
+Web 工具用于网页搜索和抓取。
+
+### Web Fetcher
+用于抓取和处理网页内容的通用设置。
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|---------------------|--------|---------------|----------------------------------------------------------------------------------------|
+| `enabled` | bool | true | 启用网页抓取功能。 |
+| `fetch_limit_bytes` | int | 10485760 | 抓取网页负载的最大大小,单位为字节(默认 10MB)。 |
+| `format` | string | "plaintext" | 抓取内容的输出格式。选项:`plaintext` 或 `markdown`(推荐)。 |
+
+### Brave
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|---------------|----------|--------|------------------------------------------------|
+| `enabled` | bool | false | 启用 Brave 搜索 |
+| `api_key` | string | - | Brave Search API 密钥 |
+| `api_keys` | string[] | - | 多个 API 密钥轮换(优先于 `api_key`) |
+| `max_results` | int | 5 | 最大结果数 |
+
+### DuckDuckGo
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|---------------|------|--------|-----------------------|
+| `enabled` | bool | true | 启用 DuckDuckGo 搜索 |
+| `max_results` | int | 5 | 最大结果数 |
+
+### Perplexity
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|---------------|----------|--------|------------------------------------------------|
+| `enabled` | bool | false | 启用 Perplexity 搜索 |
+| `api_key` | string | - | Perplexity API 密钥 |
+| `api_keys` | string[] | - | 多个 API 密钥轮换(优先于 `api_key`) |
+| `max_results` | int | 5 | 最大结果数 |
+
+### Tavily
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|---------------|--------|--------|-----------------------------------|
+| `enabled` | bool | false | 启用 Tavily 搜索 |
+| `api_key` | string | - | Tavily API 密钥 |
+| `base_url` | string | - | 自定义 Tavily API 基础 URL |
+| `max_results` | int | 0 | 最大结果数(0 = 默认) |
+
+### SearXNG
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|---------------|--------|--------------------------|-----------------------|
+| `enabled` | bool | false | 启用 SearXNG 搜索 |
+| `base_url` | string | `http://localhost:8888` | SearXNG 实例 URL |
+| `max_results` | int | 5 | 最大结果数 |
+
+### GLM Search
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|-----------------|--------|------------------------------------------------------|-----------------------|
+| `enabled` | bool | false | 启用 GLM 搜索 |
+| `api_key` | string | - | GLM API 密钥 |
+| `base_url` | string | `https://open.bigmodel.cn/api/paas/v4/web_search` | GLM Search API URL |
+| `search_engine` | string | `search_std` | 搜索引擎类型 |
+| `max_results` | int | 5 | 最大结果数 |
+
+### 其他 Web 设置
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|--------------------------|----------|--------|-------------------------------------------------|
+| `prefer_native` | bool | true | 优先使用 provider 原生搜索而非配置的搜索引擎 |
+| `private_host_whitelist` | string[] | `[]` | 允许 Web 抓取的私有/内部主机白名单 |
+
+## Exec 工具
+
+Exec 工具用于执行 shell 命令。
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|------------------------|-------|--------|--------------------------------|
+| `enabled` | bool | true | 启用 exec 工具 |
+| `enable_deny_patterns` | bool | true | 启用默认的危险命令拦截 |
+| `custom_deny_patterns` | array | [] | 自定义拒绝模式(正则表达式) |
+
+### 禁用 Exec 工具
+
+要完全禁用 `exec` 工具,请将 `enabled` 设置为 `false`:
+
+**通过配置文件:**
+```json
+{
+ "tools": {
+ "exec": {
+ "enabled": false
+ }
+ }
+}
+```
+
+**通过环境变量:**
+```bash
+PICOCLAW_TOOLS_EXEC_ENABLED=false
+```
+
+> **注意:** 禁用后,代理将无法执行 shell 命令。这也会影响 Cron 工具运行计划 shell 命令的能力。
+
+### 功能说明
+
+- **`enable_deny_patterns`**:设为 `false` 可完全禁用默认的危险命令拦截模式
+- **`custom_deny_patterns`**:添加自定义拒绝正则模式;匹配的命令将被拦截
+
+### 默认拦截的命令模式
+
+默认情况下,PicoClaw 会拦截以下危险命令:
+
+- 删除命令:`rm -rf`、`del /f/q`、`rmdir /s`
+- 磁盘操作:`format`、`mkfs`、`diskpart`、`dd if=`、写入 `/dev/sd*`
+- 系统操作:`shutdown`、`reboot`、`poweroff`
+- 命令替换:`$()`、`${}`、反引号
+- 管道到 shell:`| sh`、`| bash`
+- 权限提升:`sudo`、`chmod`、`chown`
+- 进程控制:`pkill`、`killall`、`kill -9`
+- 远程操作:`curl | sh`、`wget | sh`、`ssh`
+- 包管理:`apt`、`yum`、`dnf`、`npm install -g`、`pip install --user`
+- 容器:`docker run`、`docker exec`
+- Git:`git push`、`git force`
+- 其他:`eval`、`source *.sh`
+
+### 已知架构限制
+
+exec 守卫仅验证发送给 PicoClaw 的顶层命令。它**不会**递归检查该命令启动后由构建工具或脚本生成的子进程。
+
+以下工作流在初始命令被允许后可以绕过直接命令守卫:
+
+- `make run`
+- `go run ./cmd/...`
+- `cargo run`
+- `npm run build`
+
+这意味着守卫对于拦截明显危险的直接命令很有用,但它**不是**未审查构建管道的完整沙箱。如果你的威胁模型包括工作区中的不受信任代码,请使用更强的隔离措施,如容器、虚拟机或围绕构建和运行命令的审批流程。
+
+### 配置示例
+
+```json
+{
+ "tools": {
+ "exec": {
+ "enable_deny_patterns": true,
+ "custom_deny_patterns": [
+ "\\brm\\s+-r\\b",
+ "\\bkillall\\s+python"
+ ]
+ }
+ }
+}
+```
+
+## Cron 工具
+
+Cron 工具用于调度周期性任务。
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|------------------------|------|--------|-------------------------------------|
+| `exec_timeout_minutes` | int | 5 | 执行超时时间(分钟),0 表示无限制 |
+| `allow_command` | bool | false | 允许 cron 任务执行 shell 命令 |
+
+## MCP 工具
+
+MCP 工具支持与外部 Model Context Protocol 服务器集成。
+
+### 工具发现(延迟加载)
+
+当连接多个 MCP 服务器时,同时暴露数百个工具可能会耗尽 LLM 的上下文窗口并增加 API 成本。**Discovery** 功能通过默认*隐藏* MCP 工具来解决此问题。
+
+LLM 不会加载所有工具,而是获得一个轻量级搜索工具(使用 BM25 关键词匹配或正则表达式)。当 LLM 需要特定功能时,它会搜索隐藏的工具库。匹配的工具随后被临时"解锁"并注入上下文中,持续配置的轮数(`ttl`)。
+
+### 全局配置
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|-------------|--------|--------|--------------------------------------|
+| `enabled` | bool | false | 全局启用 MCP 集成 |
+| `discovery` | object | `{}` | 工具发现配置(见下文) |
+| `servers` | object | `{}` | 服务器名称到服务器配置的映射 |
+
+### Discovery 配置(`discovery`)
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|----------------------|------|--------|---------------------------------------------------------------------------------------------------------------|
+| `enabled` | bool | false | 如果为 true,MCP 工具将被隐藏并按需通过搜索加载。如果为 false,所有工具都会被加载 |
+| `ttl` | int | 5 | 已发现工具保持解锁状态的对话轮数 |
+| `max_search_results` | int | 5 | 每次搜索查询返回的最大工具数 |
+| `use_bm25` | bool | true | 启用自然语言/关键词搜索工具(`tool_search_tool_bm25`)。**警告**:比正则搜索消耗更多资源 |
+| `use_regex` | bool | false | 启用正则模式搜索工具(`tool_search_tool_regex`) |
+
+> **注意:** 如果 `discovery.enabled` 为 `true`,你**必须**启用至少一个搜索引擎(`use_bm25` 或 `use_regex`),
+> 否则应用程序将无法启动。
+
+### 单服务器配置
+
+| 配置项 | 类型 | 必需 | 描述 |
+|------------|--------|----------|------------------------------------|
+| `enabled` | bool | 是 | 启用此 MCP 服务器 |
+| `type` | string | 否 | 传输类型:`stdio`、`sse`、`http` |
+| `command` | string | stdio | stdio 传输的可执行命令 |
+| `args` | array | 否 | stdio 传输的命令参数 |
+| `env` | object | 否 | stdio 进程的环境变量 |
+| `env_file` | string | 否 | stdio 进程的环境文件路径 |
+| `url` | string | sse/http | `sse`/`http` 传输的端点 URL |
+| `headers` | object | 否 | `sse`/`http` 传输的 HTTP 头 |
+
+### 传输行为
+
+- 如果省略 `type`,传输方式将自动检测:
+ - 设置了 `url` → `sse`
+ - 设置了 `command` → `stdio`
+- `http` 和 `sse` 都使用 `url` + 可选的 `headers`。
+- `env` 和 `env_file` 仅应用于 `stdio` 服务器。
+
+### 配置示例
+
+#### 1) Stdio MCP 服务器
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "servers": {
+ "filesystem": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-filesystem",
+ "/tmp"
+ ]
+ }
+ }
+ }
+ }
+}
+```
+
+#### 2) 远程 SSE/HTTP MCP 服务器
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "servers": {
+ "remote-mcp": {
+ "enabled": true,
+ "type": "sse",
+ "url": "https://example.com/mcp",
+ "headers": {
+ "Authorization": "Bearer YOUR_TOKEN"
+ }
+ }
+ }
+ }
+ }
+}
+```
+
+#### 3) 启用工具发现的大规模 MCP 设置
+
+*在此示例中,LLM 只会看到 `tool_search_tool_bm25`。它将仅在用户请求时动态搜索并解锁 Github 或 Postgres 工具。*
+
+```json
+{
+ "tools": {
+ "mcp": {
+ "enabled": true,
+ "discovery": {
+ "enabled": true,
+ "ttl": 5,
+ "max_search_results": 5,
+ "use_bm25": true,
+ "use_regex": false
+ },
+ "servers": {
+ "github": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-github"
+ ],
+ "env": {
+ "GITHUB_PERSONAL_ACCESS_TOKEN": "YOUR_GITHUB_TOKEN"
+ }
+ },
+ "postgres": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-postgres",
+ "postgresql://user:password@localhost/dbname"
+ ]
+ },
+ "slack": {
+ "enabled": true,
+ "command": "npx",
+ "args": [
+ "-y",
+ "@modelcontextprotocol/server-slack"
+ ],
+ "env": {
+ "SLACK_BOT_TOKEN": "YOUR_SLACK_BOT_TOKEN",
+ "SLACK_TEAM_ID": "YOUR_SLACK_TEAM_ID"
+ }
+ }
+ }
+ }
+ }
+}
+```
+
+## Skills 工具
+
+Skills 工具配置通过 ClawHub 等注册表进行技能发现和安装。
+
+### 注册表
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|------------------------------------|--------|----------------------|--------------------------------------|
+| `registries.clawhub.enabled` | bool | true | 启用 ClawHub 注册表 |
+| `registries.clawhub.base_url` | string | `https://clawhub.ai` | ClawHub 基础 URL |
+| `registries.clawhub.auth_token` | string | `""` | 可选的 Bearer 令牌,用于更高速率限制 |
+| `registries.clawhub.search_path` | string | `""` | 搜索 API 路径 |
+| `registries.clawhub.skills_path` | string | `""` | Skills API 路径 |
+| `registries.clawhub.download_path` | string | `""` | 下载 API 路径 |
+| `registries.clawhub.timeout` | int | 0 | 请求超时时间(秒),0 = 默认 |
+| `registries.clawhub.max_zip_size` | int | 0 | 技能 zip 最大大小(字节),0 = 默认 |
+| `registries.clawhub.max_response_size` | int | 0 | API 响应最大大小(字节),0 = 默认 |
+
+### GitHub 集成
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|------------------|--------|--------|-------------------------------|
+| `github.proxy` | string | `""` | GitHub API 请求的 HTTP 代理 |
+| `github.token` | string | `""` | GitHub 个人访问令牌 |
+
+### 搜索设置
+
+| 配置项 | 类型 | 默认值 | 描述 |
+|----------------------------|------|--------|--------------------------|
+| `max_concurrent_searches` | int | 2 | 最大并发技能搜索请求数 |
+| `search_cache.max_size` | int | 50 | 最大缓存搜索结果数 |
+| `search_cache.ttl_seconds` | int | 300 | 缓存 TTL(秒) |
+
+### 配置示例
+
+```json
+{
+ "tools": {
+ "skills": {
+ "registries": {
+ "clawhub": {
+ "enabled": true,
+ "base_url": "https://clawhub.ai",
+ "auth_token": ""
+ }
+ },
+ "github": {
+ "proxy": "",
+ "token": ""
+ },
+ "max_concurrent_searches": 2,
+ "search_cache": {
+ "max_size": 50,
+ "ttl_seconds": 300
+ }
+ }
+ }
+}
+```
+
+## 环境变量
+
+所有配置选项都可以通过格式为 `PICOCLAW_TOOLS__` 的环境变量覆盖:
+
+例如:
+
+- `PICOCLAW_TOOLS_WEB_BRAVE_ENABLED=true`
+- `PICOCLAW_TOOLS_EXEC_ENABLED=false`
+- `PICOCLAW_TOOLS_EXEC_ENABLE_DENY_PATTERNS=false`
+- `PICOCLAW_TOOLS_CRON_EXEC_TIMEOUT_MINUTES=10`
+- `PICOCLAW_TOOLS_MCP_ENABLED=true`
+
+注意:嵌套的映射式配置(例如 `tools.mcp.servers..*`)在 `config.json` 中配置,而非通过环境变量。
diff --git a/docs/zh/troubleshooting.md b/docs/zh/troubleshooting.md
new file mode 100644
index 000000000..be4d4f5d7
--- /dev/null
+++ b/docs/zh/troubleshooting.md
@@ -0,0 +1,45 @@
+# 🐛 疑难解答
+
+> 返回 [README](../../README.zh.md)
+
+## "model ... not found in model_list" 或 OpenRouter "free is not a valid model ID"
+
+**症状:** 你看到以下任一错误:
+
+- `Error creating provider: model "openrouter/free" not found in model_list`
+- OpenRouter 返回 400:`"free is not a valid model ID"`
+
+**原因:** `model_list` 条目中的 `model` 字段是发送给 API 的内容。对于 OpenRouter,你必须使用**完整的**模型 ID,而不是简写。
+
+- **错误:** `"model": "free"` → OpenRouter 收到 `free` 并拒绝。
+- **正确:** `"model": "openrouter/free"` → OpenRouter 收到 `openrouter/free`(自动免费层路由)。
+
+**修复方法:** 在 `~/.picoclaw/config.json`(或你的配置路径)中:
+
+1. **agents.defaults.model_name** 必须匹配 `model_list` 中的某个 `model_name`(例如 `"openrouter-free"`)。
+2. 该条目的 **model** 必须是有效的 OpenRouter 模型 ID,例如:
+ - `"openrouter/free"` – 自动免费层
+ - `"google/gemini-2.0-flash-exp:free"`
+ - `"meta-llama/llama-3.1-8b-instruct:free"`
+
+示例片段:
+
+```json
+{
+ "agents": {
+ "defaults": {
+ "model_name": "openrouter-free"
+ }
+ },
+ "model_list": [
+ {
+ "model_name": "openrouter-free",
+ "model": "openrouter/free",
+ "api_key": "sk-or-v1-YOUR_OPENROUTER_KEY",
+ "api_base": "https://openrouter.ai/api/v1"
+ }
+ ]
+}
+```
+
+在 [OpenRouter Keys](https://openrouter.ai/keys) 获取你的密钥。
diff --git a/examples/pico-echo-server/README.md b/examples/pico-echo-server/README.md
new file mode 100644
index 000000000..f6b5d8020
--- /dev/null
+++ b/examples/pico-echo-server/README.md
@@ -0,0 +1,47 @@
+# pico-echo-server
+
+Minimal Pico Protocol WebSocket server for testing the `pico_client` channel.
+
+## Usage
+
+```bash
+go run ./examples/pico-echo-server -addr :9090 -token secret
+```
+
+### Flags
+
+| Flag | Default | Description |
+|----------|---------|------------------------------------|
+| `-addr` | `:9090` | Listen address |
+| `-token` | (none) | Auth token; empty disables auth |
+
+## How it works
+
+- Listens for WebSocket connections at `/ws`
+- Authenticates via `Authorization: Bearer ` header or `?token=` query param
+- Prints received `message.send` content to stdout
+- Responds to `ping` with `pong`
+- Lines typed into stdin are broadcast as `message.create` to all connected clients
+
+## Testing with pico_client
+
+1. Start the server:
+ ```bash
+ go run ./examples/pico-echo-server -token mytoken
+ ```
+
+2. Configure `pico_client` in your `config.json`:
+ ```json
+ {
+ "channels": {
+ "pico_client": {
+ "enabled": true,
+ "url": "ws://localhost:9090/ws",
+ "token": "mytoken",
+ "session_id": "test-session"
+ }
+ }
+ }
+ ```
+
+3. Start picoclaw — the client connects and you can exchange messages interactively via stdin/stdout.
diff --git a/examples/pico-echo-server/main.go b/examples/pico-echo-server/main.go
new file mode 100644
index 000000000..46970fb34
--- /dev/null
+++ b/examples/pico-echo-server/main.go
@@ -0,0 +1,160 @@
+// pico-echo-server is a minimal Pico Protocol WebSocket server for testing
+// the pico_client channel. It accepts connections, prints received messages
+// to stdout, and forwards stdin lines as message.create to all connected clients.
+//
+// Usage:
+//
+// go run ./examples/pico-echo-server -addr :9090 -token secret
+//
+// Then configure pico_client with url=ws://localhost:9090/ws&token=secret.
+package main
+
+import (
+ "bufio"
+ "encoding/json"
+ "flag"
+ "fmt"
+ "log"
+ "net/http"
+ "os"
+ "strings"
+ "sync"
+ "time"
+
+ "github.com/gorilla/websocket"
+)
+
+type picoMessage struct {
+ Type string `json:"type"`
+ ID string `json:"id,omitempty"`
+ SessionID string `json:"session_id,omitempty"`
+ Timestamp int64 `json:"timestamp,omitempty"`
+ Payload map[string]any `json:"payload,omitempty"`
+}
+
+var upgrader = websocket.Upgrader{CheckOrigin: func(*http.Request) bool { return true }}
+
+type server struct {
+ token string
+ mu sync.Mutex
+ conns map[*websocket.Conn]string // conn → sessionID
+}
+
+func (s *server) handleWS(w http.ResponseWriter, r *http.Request) {
+ if s.token != "" {
+ auth := r.Header.Get("Authorization")
+ if auth != "Bearer "+s.token {
+ http.Error(w, "unauthorized", http.StatusUnauthorized)
+ return
+ }
+ }
+
+ conn, err := upgrader.Upgrade(w, r, nil)
+ if err != nil {
+ log.Printf("upgrade: %v", err)
+ return
+ }
+
+ sessionID := r.URL.Query().Get("session_id")
+ if sessionID == "" {
+ sessionID = fmt.Sprintf("sess-%d", time.Now().UnixMilli())
+ }
+
+ s.mu.Lock()
+ s.conns[conn] = sessionID
+ s.mu.Unlock()
+
+ log.Printf("[+] client connected (session=%s)", sessionID)
+
+ defer func() {
+ s.mu.Lock()
+ delete(s.conns, conn)
+ s.mu.Unlock()
+ conn.Close()
+ log.Printf("[-] client disconnected (session=%s)", sessionID)
+ }()
+
+ for {
+ _, raw, err := conn.ReadMessage()
+ if err != nil {
+ if websocket.IsUnexpectedCloseError(err, websocket.CloseGoingAway, websocket.CloseNormalClosure) {
+ log.Printf("read error: %v", err)
+ }
+ return
+ }
+
+ var msg picoMessage
+ if err := json.Unmarshal(raw, &msg); err != nil {
+ log.Printf("bad json: %v", err)
+ continue
+ }
+
+ switch msg.Type {
+ case "ping":
+ pong := picoMessage{Type: "pong", ID: msg.ID, Timestamp: time.Now().UnixMilli()}
+ conn.WriteJSON(pong)
+
+ case "message.send":
+ content, _ := msg.Payload["content"].(string)
+ fmt.Printf("[%s] %s\n", sessionID, content)
+
+ case "typing.start":
+ log.Printf("[%s] typing...", sessionID)
+
+ case "typing.stop":
+ log.Printf("[%s] stopped typing", sessionID)
+
+ default:
+ log.Printf("[%s] unknown type: %s", sessionID, msg.Type)
+ }
+ }
+}
+
+func (s *server) broadcast(content string) {
+ msg := picoMessage{
+ Type: "message.create",
+ Timestamp: time.Now().UnixMilli(),
+ Payload: map[string]any{"content": content},
+ }
+
+ s.mu.Lock()
+ defer s.mu.Unlock()
+
+ for conn, sid := range s.conns {
+ msg.SessionID = sid
+ if err := conn.WriteJSON(msg); err != nil {
+ log.Printf("write to %s failed: %v", sid, err)
+ }
+ }
+}
+
+func main() {
+ addr := flag.String("addr", ":9090", "listen address")
+ token := flag.String("token", "", "auth token (empty = no auth)")
+ flag.Parse()
+
+ s := &server{
+ token: *token,
+ conns: make(map[*websocket.Conn]string),
+ }
+
+ http.HandleFunc("/ws", s.handleWS)
+
+ log.Printf("listening on %s", *addr)
+ log.Printf("connect with: ws://localhost%s/ws", *addr)
+ fmt.Println("Type messages to send to connected clients (Ctrl+C to quit):")
+
+ go func() {
+ scanner := bufio.NewScanner(os.Stdin)
+ for scanner.Scan() {
+ line := strings.TrimSpace(scanner.Text())
+ if line == "" {
+ continue
+ }
+ s.broadcast(line)
+ log.Printf("[server] sent: %s", line)
+ }
+ }()
+
+ log.Fatal(http.ListenAndServe(*addr, nil))
+}
diff --git a/go.mod b/go.mod
index f29ef7207..cfc930d37 100644
--- a/go.mod
+++ b/go.mod
@@ -1,13 +1,15 @@
module github.com/sipeed/picoclaw
-go 1.25.7
+go 1.25.8
require (
+ github.com/BurntSushi/toml v1.6.0
+ fyne.io/systray v1.12.0
github.com/adhocore/gronx v1.19.6
- github.com/anthropics/anthropic-sdk-go v1.22.1
+ github.com/anthropics/anthropic-sdk-go v1.26.0
github.com/bwmarrin/discordgo v0.29.0
- github.com/caarlos0/env/v11 v11.3.1
- github.com/ergochat/irc-go v0.5.0
+ github.com/caarlos0/env/v11 v11.4.0
+ github.com/ergochat/irc-go v0.6.0
github.com/ergochat/readline v0.1.3
github.com/gdamore/tcell/v2 v2.13.8
github.com/gomarkdown/markdown v0.0.0-20260217112301-37c66b85d6ab
@@ -16,8 +18,8 @@ require (
github.com/h2non/filetype v1.1.3
github.com/larksuite/oapi-sdk-go/v3 v3.5.3
github.com/mdp/qrterminal/v3 v3.2.1
- github.com/modelcontextprotocol/go-sdk v1.3.1
- github.com/mymmrac/telego v1.6.0
+ github.com/modelcontextprotocol/go-sdk v1.4.1
+ github.com/mymmrac/telego v1.7.0
github.com/open-dingtalk/dingtalk-stream-sdk-go v0.9.1
github.com/openai/openai-go/v3 v3.22.0
github.com/rivo/tview v0.42.0
@@ -27,40 +29,41 @@ require (
github.com/stretchr/testify v1.11.1
github.com/tencent-connect/botgo v0.2.1
go.mau.fi/whatsmeow v0.0.0-20260219150138-7ae702b1eed4
- golang.org/x/oauth2 v0.35.0
+ golang.org/x/oauth2 v0.36.0
+ golang.org/x/term v0.41.0
golang.org/x/time v0.14.0
google.golang.org/protobuf v1.36.11
gopkg.in/yaml.v3 v3.0.1
- maunium.net/go/mautrix v0.26.3
+ maunium.net/go/mautrix v0.26.4
modernc.org/sqlite v1.46.1
)
require (
- filippo.io/edwards25519 v1.1.1 // indirect
+ filippo.io/edwards25519 v1.2.0 // indirect
github.com/beeper/argo-go v1.1.2 // indirect
github.com/coder/websocket v1.8.14 // indirect
github.com/davecgh/go-spew v1.1.1 // indirect
github.com/dustin/go-humanize v1.0.1 // indirect
github.com/elliotchance/orderedmap/v3 v3.1.0 // indirect
github.com/gdamore/encoding v1.0.1 // indirect
+ github.com/godbus/dbus/v5 v5.1.0 // indirect
github.com/inconshreveable/mousetrap v1.1.0 // indirect
github.com/lucasb-eyer/go-colorful v1.3.0 // indirect
github.com/mattn/go-colorable v0.1.14 // indirect
github.com/mattn/go-isatty v0.0.20 // indirect
github.com/ncruces/go-strftime v1.0.0 // indirect
- github.com/petermattis/goid v0.0.0-20260113132338-7c7de50cc741 // indirect
+ github.com/petermattis/goid v0.0.0-20260226131333-17d1149c6ac6 // indirect
github.com/pmezard/go-difflib v1.0.0 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/rivo/uniseg v0.4.7 // indirect
github.com/segmentio/asm v1.1.3 // indirect
- github.com/segmentio/encoding v0.5.3 // indirect
+ github.com/segmentio/encoding v0.5.4 // indirect
github.com/spf13/pflag v1.0.10 // indirect
github.com/vektah/gqlparser/v2 v2.5.27 // indirect
go.mau.fi/libsignal v0.2.1 // indirect
- go.mau.fi/util v0.9.6 // indirect
- golang.org/x/exp v0.0.0-20260212183809-81e46e3db34a // indirect
- golang.org/x/term v0.40.0 // indirect
- golang.org/x/text v0.34.0 // indirect
+ go.mau.fi/util v0.9.7 // indirect
+ golang.org/x/exp v0.0.0-20260312153236-7ab1446f8b90 // indirect
+ golang.org/x/text v0.35.0 // indirect
modernc.org/libc v1.67.6 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
@@ -73,7 +76,7 @@ require (
github.com/bytedance/sonic v1.15.0 // indirect
github.com/bytedance/sonic/loader v0.5.0 // indirect
github.com/cloudwego/base64x v0.1.6 // indirect
- github.com/github/copilot-sdk/go v0.1.23
+ github.com/github/copilot-sdk/go v0.1.32
github.com/go-resty/resty/v2 v2.17.1 // indirect
github.com/gogo/protobuf v1.3.2 // indirect
github.com/google/jsonschema-go v0.4.2 // indirect
@@ -87,11 +90,11 @@ require (
github.com/twitchyliquid64/golang-asm v0.15.1 // indirect
github.com/valyala/bytebufferpool v1.0.0 // indirect
github.com/valyala/fasthttp v1.69.0 // indirect
- github.com/valyala/fastjson v1.6.7 // indirect
+ github.com/valyala/fastjson v1.6.10 // indirect
github.com/yosida95/uritemplate/v3 v3.0.2 // indirect
golang.org/x/arch v0.24.0 // indirect
- golang.org/x/crypto v0.48.0 // indirect
- golang.org/x/net v0.51.0 // indirect
- golang.org/x/sync v0.19.0 // indirect
- golang.org/x/sys v0.41.0 // indirect
+ golang.org/x/crypto v0.49.0
+ golang.org/x/net v0.52.0
+ golang.org/x/sync v0.20.0 // indirect
+ golang.org/x/sys v0.42.0 // indirect
)
diff --git a/go.sum b/go.sum
index addbab56c..f24b997d4 100644
--- a/go.sum
+++ b/go.sum
@@ -1,6 +1,10 @@
cloud.google.com/go/compute/metadata v0.3.0/go.mod h1:zFmK7XCadkQkj6TtorcaGlCW1hT1fIilQDwofLpJ20k=
-filippo.io/edwards25519 v1.1.1 h1:YpjwWWlNmGIDyXOn8zLzqiD+9TyIlPhGFG96P39uBpw=
-filippo.io/edwards25519 v1.1.1/go.mod h1:BxyFTGdWcka3PhytdK4V28tE5sGfRvvvRV7EaN4VDT4=
+filippo.io/edwards25519 v1.2.0 h1:crnVqOiS4jqYleHd9vaKZ+HKtHfllngJIiOpNpoJsjo=
+filippo.io/edwards25519 v1.2.0/go.mod h1:xzAOLCNug/yB62zG1bQ8uziwrIqIuxhctzJT18Q77mc=
+fyne.io/systray v1.12.0 h1:CA1Kk0e2zwFlxtc02L3QFSiIbxJ/P0n582YrZHT7aTM=
+fyne.io/systray v1.12.0/go.mod h1:RVwqP9nYMo7h5zViCBHri2FgjXF7H2cub7MAq4NSoLs=
+github.com/BurntSushi/toml v1.6.0 h1:dRaEfpa2VI55EwlIW72hMRHdWouJeRF7TPYhI+AUQjk=
+github.com/BurntSushi/toml v1.6.0/go.mod h1:ukJfTF/6rtPPRCnwkur4qwRxa8vTRFBF0uk2lLoLwho=
github.com/DATA-DOG/go-sqlmock v1.5.2 h1:OcvFkGmslmlZibjAjaHm3L//6LiuBgolP7OputlJIzU=
github.com/DATA-DOG/go-sqlmock v1.5.2/go.mod h1:88MAG/4G7SMwSE3CeA0ZKzrT5CiOU3OJ+JlNzwDqpNU=
github.com/adhocore/gronx v1.19.6 h1:5KNVcoR9ACgL9HhEqCm5QXsab/gI4QDIybTAWcXDKDc=
@@ -11,8 +15,8 @@ github.com/andreyvit/diff v0.0.0-20170406064948-c7f18ee00883 h1:bvNMNQO63//z+xNg
github.com/andreyvit/diff v0.0.0-20170406064948-c7f18ee00883/go.mod h1:rCTlJbsFo29Kk6CurOXKm700vrz8f0KW0JNfpkRJY/8=
github.com/andybalholm/brotli v1.2.0 h1:ukwgCxwYrmACq68yiUqwIWnGY0cTPox/M94sVwToPjQ=
github.com/andybalholm/brotli v1.2.0/go.mod h1:rzTDkvFWvIrjDXZHkuS16NPggd91W3kUSvPlQ1pLaKY=
-github.com/anthropics/anthropic-sdk-go v1.22.1 h1:xbsc3vJKCX/ELDZSpTNfz9wCgrFsamwFewPb1iI0Xh0=
-github.com/anthropics/anthropic-sdk-go v1.22.1/go.mod h1:WTz31rIUHUHqai2UslPpw5CwXrQP3geYBioRV4WOLvE=
+github.com/anthropics/anthropic-sdk-go v1.26.0 h1:oUTzFaUpAevfuELAP1sjL6CQJ9HHAfT7CoSYSac11PY=
+github.com/anthropics/anthropic-sdk-go v1.26.0/go.mod h1:qUKmaW+uuPB64iy1l+4kOSvaLqPXnHTTBKH6RVZ7q5Q=
github.com/beeper/argo-go v1.1.2 h1:UQI2G8F+NLfGTOmTUI0254pGKx/HUU/etbUGTJv91Fs=
github.com/beeper/argo-go v1.1.2/go.mod h1:M+LJAnyowKVQ6Rdj6XYGEn+qcVFkb3R/MUpqkGR0hM4=
github.com/bwmarrin/discordgo v0.29.0 h1:FmWeXFaKUwrcL3Cx65c20bTRW+vOb6k8AnaP+EgjDno=
@@ -23,8 +27,8 @@ github.com/bytedance/sonic v1.15.0 h1:/PXeWFaR5ElNcVE84U0dOHjiMHQOwNIx3K4ymzh/uS
github.com/bytedance/sonic v1.15.0/go.mod h1:tFkWrPz0/CUCLEF4ri4UkHekCIcdnkqXw9VduqpJh0k=
github.com/bytedance/sonic/loader v0.5.0 h1:gXH3KVnatgY7loH5/TkeVyXPfESoqSBSBEiDd5VjlgE=
github.com/bytedance/sonic/loader v0.5.0/go.mod h1:AR4NYCk5DdzZizZ5djGqQ92eEhCCcdf5x77udYiSJRo=
-github.com/caarlos0/env/v11 v11.3.1 h1:cArPWC15hWmEt+gWk7YBi7lEXTXCvpaSdCiZE2X5mCA=
-github.com/caarlos0/env/v11 v11.3.1/go.mod h1:qupehSf/Y0TUTsxKywqRt/vJjN5nz6vauiYEUUr8P4U=
+github.com/caarlos0/env/v11 v11.4.0 h1:Kcb6t5kIIr4XkoQC9AF2j+8E1Jsrl3Wz/hhm1LtoGAc=
+github.com/caarlos0/env/v11 v11.4.0/go.mod h1:qupehSf/Y0TUTsxKywqRt/vJjN5nz6vauiYEUUr8P4U=
github.com/cespare/xxhash/v2 v2.1.2/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/cespare/xxhash/v2 v2.2.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/cloudwego/base64x v0.1.6 h1:t11wG9AECkCDk5fMSoxmufanudBtJ+/HemLstXDLI2M=
@@ -38,12 +42,14 @@ github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSs
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/dgryski/go-rendezvous v0.0.0-20200823014737-9f7001d12a5f/go.mod h1:cuUVRXasLTGF7a8hSLbxyZXjz+1KgoB3wDUb6vlszIc=
+github.com/dnaeon/go-vcr v1.2.0 h1:zHCHvJYTMh1N7xnV7zf1m1GPBF9Ad0Jk/whtQ1663qI=
+github.com/dnaeon/go-vcr v1.2.0/go.mod h1:R4UdLID7HZT3taECzJs4YgbbH6PIGXB6W/sc5OLb6RQ=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/elliotchance/orderedmap/v3 v3.1.0 h1:j4DJ5ObEmMBt/lcwIecKcoRxIQUEnw0L804lXYDt/pg=
github.com/elliotchance/orderedmap/v3 v3.1.0/go.mod h1:G+Hc2RwaZvJMcS4JpGCOyViCnGeKf0bTYCGTO4uhjSo=
-github.com/ergochat/irc-go v0.5.0 h1:woQ1RS9YbfgqPgSpPBBQeczXGIGzR0aC7dEgk469fTw=
-github.com/ergochat/irc-go v0.5.0/go.mod h1:2vi7KNpIPWnReB5hmLpl92eMywQvuIeIIGdt/FQCph0=
+github.com/ergochat/irc-go v0.6.0 h1:Y0AGV76aeihJfCtLaQh+OyJKFiKGrYC0VTkeMZ6XW28=
+github.com/ergochat/irc-go v0.6.0/go.mod h1:2vi7KNpIPWnReB5hmLpl92eMywQvuIeIIGdt/FQCph0=
github.com/ergochat/readline v0.1.3 h1:/DytGTmwdUJcLAe3k3VJgowh5vNnsdifYT6uVaf4pSo=
github.com/ergochat/readline v0.1.3/go.mod h1:o3ux9QLHLm77bq7hDB21UTm6HlV2++IPDMfIfKDuOgY=
github.com/fsnotify/fsnotify v1.4.7/go.mod h1:jwhsz4b93w/PPRr/qN1Yymfu8t87LnFCMoQvtojpjFo=
@@ -52,8 +58,8 @@ github.com/gdamore/encoding v1.0.1 h1:YzKZckdBL6jVt2Gc+5p82qhrGiqMdG/eNs6Wy0u3Uh
github.com/gdamore/encoding v1.0.1/go.mod h1:0Z0cMFinngz9kS1QfMjCP8TY7em3bZYeeklsSDPivEo=
github.com/gdamore/tcell/v2 v2.13.8 h1:Mys/Kl5wfC/GcC5Cx4C2BIQH9dbnhnkPgS9/wF3RlfU=
github.com/gdamore/tcell/v2 v2.13.8/go.mod h1:+Wfe208WDdB7INEtCsNrAN6O2m+wsTPk1RAovjaILlo=
-github.com/github/copilot-sdk/go v0.1.23 h1:uExtO/inZQndCZMiSAA1hvXINiz9tqo/MZgQzFzurxw=
-github.com/github/copilot-sdk/go v0.1.23/go.mod h1:GdwwBfMbm9AABLEM3x5IZKw4ZfwCYxZ1BgyytmZenQ0=
+github.com/github/copilot-sdk/go v0.1.32 h1:wc9SFWwxXhJts6vyzzboPLJqcEJGnHE8rMCAY1RrUgo=
+github.com/github/copilot-sdk/go v0.1.32/go.mod h1:qc2iEF7hdO8kzSvbyGvrcGhuk2fzdW4xTtT0+1EH2ts=
github.com/go-redis/redis/v8 v8.11.4/go.mod h1:2Z2wHZXdQpCDXEGzqMockDpNyYvi2l4Pxt6RJr792+w=
github.com/go-resty/resty/v2 v2.6.0/go.mod h1:PwvJS6hvaPkjtjNg9ph+VrSD92bi5Zq73w/BIH7cC3Q=
github.com/go-resty/resty/v2 v2.17.1 h1:x3aMpHK1YM9e4va/TMDRlusDDoZiQ+ViDu/WpA6xTM4=
@@ -62,10 +68,12 @@ github.com/go-task/slim-sprig v0.0.0-20210107165309-348f09dbbbc0/go.mod h1:fyg78
github.com/go-test/deep v1.1.1 h1:0r/53hagsehfO4bzD2Pgr/+RgHqhmf+k1Bpse2cTu1U=
github.com/go-test/deep v1.1.1/go.mod h1:5C2ZWiW0ErCdrYzpqxLbTX7MG14M9iiw8DgHncVwcsE=
github.com/godbus/dbus/v5 v5.0.4/go.mod h1:xhWf0FNVPg57R7Z0UbKHbJfkEywrmjJnf7w5xrFpKfA=
+github.com/godbus/dbus/v5 v5.1.0 h1:4KLkAxT3aOY8Li4FRJe/KvhoNFFxo0m6fNuFUO8QJUk=
+github.com/godbus/dbus/v5 v5.1.0/go.mod h1:xhWf0FNVPg57R7Z0UbKHbJfkEywrmjJnf7w5xrFpKfA=
github.com/gogo/protobuf v1.3.2 h1:Ov1cvc58UF3b5XjBnZv7+opcTcQFZebYjWzi34vdm4Q=
github.com/gogo/protobuf v1.3.2/go.mod h1:P1XiOD3dCwIKUDQYPy72D8LYyHL2YPYrpS2s69NZV8Q=
-github.com/golang-jwt/jwt/v5 v5.2.2 h1:Rl4B7itRWVtYIHFrSNd7vhTiz9UpLdi6gZhZ3wEeDy8=
-github.com/golang-jwt/jwt/v5 v5.2.2/go.mod h1:pqrtFR0X4osieyHYxtmOUWsAWrfe1Q5UVIyoH402zdk=
+github.com/golang-jwt/jwt/v5 v5.3.0 h1:pv4AsKCKKZuqlgs5sUmn4x8UlGa0kEVt/puTpKx9vvo=
+github.com/golang-jwt/jwt/v5 v5.3.0/go.mod h1:fxCRLWMO43lRc8nhHWY6LGqRcf+1gQWArsqaEUEa5bE=
github.com/golang/protobuf v1.2.0/go.mod h1:6lQm79b+lXiMfvg/cZm0SGofjICqVBUtrP5yJMmIC1U=
github.com/golang/protobuf v1.4.0-rc.1/go.mod h1:ceaxUfeHdC40wWswd/P6IGgMaK3YpKi5j83Wpe3EHw8=
github.com/golang/protobuf v1.4.0-rc.1.0.20200221234624-67d41d38c208/go.mod h1:xKAWHe0F5eneWXFV3EuXVDTCmh+JuBKY0li0aMyXATA=
@@ -134,10 +142,10 @@ github.com/mattn/go-sqlite3 v1.14.34 h1:3NtcvcUnFBPsuRcno8pUtupspG/GM+9nZ88zgJcp
github.com/mattn/go-sqlite3 v1.14.34/go.mod h1:Uh1q+B4BYcTPb+yiD3kU8Ct7aC0hY9fxUwlHK0RXw+Y=
github.com/mdp/qrterminal/v3 v3.2.1 h1:6+yQjiiOsSuXT5n9/m60E54vdgFsw0zhADHhHLrFet4=
github.com/mdp/qrterminal/v3 v3.2.1/go.mod h1:jOTmXvnBsMy5xqLniO0R++Jmjs2sTm9dFSuQ5kpz/SU=
-github.com/modelcontextprotocol/go-sdk v1.3.1 h1:TfqtNKOIWN4Z1oqmPAiWDC2Jq7K9OdJaooe0teoXASI=
-github.com/modelcontextprotocol/go-sdk v1.3.1/go.mod h1:DgVX498dMD8UJlseK1S5i1T4tFz2fkBk4xogC3D15nw=
-github.com/mymmrac/telego v1.6.0 h1:Zc8rgyHozvd/7ZgyrigyHdAF9koHYMfilYfyB6wlFC0=
-github.com/mymmrac/telego v1.6.0/go.mod h1:xt6ZWA8zi8KmuzryE1ImEdl9JSwjHNpM4yhC7D8hU4Y=
+github.com/modelcontextprotocol/go-sdk v1.4.1 h1:M4x9GyIPj+HoIlHNGpK2hq5o3BFhC+78PkEaldQRphc=
+github.com/modelcontextprotocol/go-sdk v1.4.1/go.mod h1:Bo/mS87hPQqHSRkMv4dQq1XCu6zv4INdXnFZabkNU6s=
+github.com/mymmrac/telego v1.7.0 h1:yRO/l00tFGG4nY66ufUKb4ARqv7qx9+LsjQv/b0NEyo=
+github.com/mymmrac/telego v1.7.0/go.mod h1:pdLV346EgVuq7Xrh3kMggeBiazeHhsdEoK0RTEOPXRM=
github.com/ncruces/go-strftime v1.0.0 h1:HMFp8mLCTPp341M/ZnA4qaf7ZlsbTc+miZjCLOFAw7w=
github.com/ncruces/go-strftime v1.0.0/go.mod h1:Fwc5htZGVVkseilnfgOVb9mKy6w1naJmn9CehxcKcls=
github.com/nxadm/tail v1.4.4/go.mod h1:kenIhsEOeOJmVchQTgglprH7qJGnHDVpk1VPCcaMI8A=
@@ -152,8 +160,8 @@ github.com/open-dingtalk/dingtalk-stream-sdk-go v0.9.1 h1:Lb/Uzkiw2Ugt2Xf03J5wmv
github.com/open-dingtalk/dingtalk-stream-sdk-go v0.9.1/go.mod h1:ln3IqPYYocZbYvl9TAOrG/cxGR9xcn4pnZRLdCTEGEU=
github.com/openai/openai-go/v3 v3.22.0 h1:6MEoNoV8sbjOVmXdvhmuX3BjVbVdcExbVyGixiyJ8ys=
github.com/openai/openai-go/v3 v3.22.0/go.mod h1:cdufnVK14cWcT9qA1rRtrXx4FTRsgbDPW7Ia7SS5cZo=
-github.com/petermattis/goid v0.0.0-20260113132338-7c7de50cc741 h1:KPpdlQLZcHfTMQRi6bFQ7ogNO0ltFT4PmtwTLW4W+14=
-github.com/petermattis/goid v0.0.0-20260113132338-7c7de50cc741/go.mod h1:pxMtw7cyUw6B2bRH0ZBANSPg+AoSud1I1iyJHI69jH4=
+github.com/petermattis/goid v0.0.0-20260226131333-17d1149c6ac6 h1:rh2lKw/P/EqHa724vYH2+VVQ1YnW4u6EOXl0PMAovZE=
+github.com/petermattis/goid v0.0.0-20260226131333-17d1149c6ac6/go.mod h1:pxMtw7cyUw6B2bRH0ZBANSPg+AoSud1I1iyJHI69jH4=
github.com/pkg/diff v0.0.0-20210226163009-20ebb0f2a09e/go.mod h1:pJLUxLENpZxwdsKMEsNbx1VGcRFpLqf3715MtcvvzbA=
github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
@@ -173,8 +181,8 @@ github.com/rs/zerolog v1.34.0/go.mod h1:bJsvje4Z08ROH4Nhs5iH600c3IkWhwp44iRc54W6
github.com/russross/blackfriday/v2 v2.1.0/go.mod h1:+Rmxgy9KzJVeS9/2gXHxylqXiyQDYRxCVz55jmeOWTM=
github.com/segmentio/asm v1.1.3 h1:WM03sfUOENvvKexOLp+pCqgb/WDjsi7EK8gIsICtzhc=
github.com/segmentio/asm v1.1.3/go.mod h1:Ld3L4ZXGNcSLRg4JBsZ3//1+f/TjYl0Mzen/DQy1EJg=
-github.com/segmentio/encoding v0.5.3 h1:OjMgICtcSFuNvQCdwqMCv9Tg7lEOXGwm1J5RPQccx6w=
-github.com/segmentio/encoding v0.5.3/go.mod h1:HS1ZKa3kSN32ZHVZ7ZLPLXWvOVIiZtyJnO1gPH1sKt0=
+github.com/segmentio/encoding v0.5.4 h1:OW1VRern8Nw6ITAtwSZ7Idrl3MXCFwXHPgqESYfvNt0=
+github.com/segmentio/encoding v0.5.4/go.mod h1:HS1ZKa3kSN32ZHVZ7ZLPLXWvOVIiZtyJnO1gPH1sKt0=
github.com/sergi/go-diff v1.3.1 h1:xkr+Oxo4BOQKmkn/B9eMK0g5Kg/983T9DqqPHwYqD+8=
github.com/sergi/go-diff v1.3.1/go.mod h1:aMJSSKb2lpPvRNec0+w3fl7LP9IOFzdc9Pa4NFbPK1I=
github.com/slack-go/slack v0.17.3 h1:zV5qO3Q+WJAQ/XwbGfNFrRMaJ5T/naqaonyPV/1TP4g=
@@ -216,8 +224,8 @@ github.com/valyala/bytebufferpool v1.0.0 h1:GqA5TC/0021Y/b9FG4Oi9Mr3q7XYx6Kllzaw
github.com/valyala/bytebufferpool v1.0.0/go.mod h1:6bBcMArwyJ5K/AmCkWv1jt77kVWyCJ6HpOuEn7z0Csc=
github.com/valyala/fasthttp v1.69.0 h1:fNLLESD2SooWeh2cidsuFtOcrEi4uB4m1mPrkJMZyVI=
github.com/valyala/fasthttp v1.69.0/go.mod h1:4wA4PfAraPlAsJ5jMSqCE2ug5tqUPwKXxVj8oNECGcw=
-github.com/valyala/fastjson v1.6.7 h1:ZE4tRy0CIkh+qDc5McjatheGX2czdn8slQjomexVpBM=
-github.com/valyala/fastjson v1.6.7/go.mod h1:CLCAqky6SMuOcxStkYQvblddUtoRxhYMGLrsQns1aXY=
+github.com/valyala/fastjson v1.6.10 h1:/yjJg8jaVQdYR3arGxPE2X5z89xrlhS0eGXdv+ADTh4=
+github.com/valyala/fastjson v1.6.10/go.mod h1:e6FubmQouUNP73jtMLmcbxS6ydWIpOfhz34TSfO3JaE=
github.com/vektah/gqlparser/v2 v2.5.27 h1:RHPD3JOplpk5mP5JGX8RKZkt2/Vwj/PZv0HxTdwFp0s=
github.com/vektah/gqlparser/v2 v2.5.27/go.mod h1:D1/VCZtV3LPnQrcPBeR/q5jkSQIPti0uYCP/RI0gIeo=
github.com/xyproto/randomstring v1.0.5 h1:YtlWPoRdgMu3NZtP45drfy1GKoojuR7hmRcnhZqKjWU=
@@ -229,8 +237,8 @@ github.com/yuin/goldmark v1.2.1/go.mod h1:3hX8gzYuyVAZsxl0MRgGTJEmQBFcNTphYh9dec
github.com/yuin/goldmark v1.4.13/go.mod h1:6yULJ656Px+3vBD8DxQVa3kxgyrAnzto9xy5taEt/CY=
go.mau.fi/libsignal v0.2.1 h1:vRZG4EzTn70XY6Oh/pVKrQGuMHBkAWlGRC22/85m9L0=
go.mau.fi/libsignal v0.2.1/go.mod h1:iVvjrHyfQqWajOUaMEsIfo3IqgVMrhWcPiiEzk7NgoU=
-go.mau.fi/util v0.9.6 h1:2nsvxm49KhI3wrFltr0+wSUBlnQ4CMtykuELjpIU+ts=
-go.mau.fi/util v0.9.6/go.mod h1:sIJpRH7Iy5Ad1SBuxQoatxtIeErgzxCtjd/2hCMkYMI=
+go.mau.fi/util v0.9.7 h1:AWGNbJfz1zRcQOKeOEYhKUG2fT+/26Gy6kyqcH8tnBg=
+go.mau.fi/util v0.9.7/go.mod h1:5T2f3ZWZFAGgmFwg3dGw7YK6kIsb9lryDzvynoR98pE=
go.mau.fi/whatsmeow v0.0.0-20260219150138-7ae702b1eed4 h1:hsmlwsM+VqfF70cpdZEeIUKer2XWCQmQPK0u0tHy3ZQ=
go.mau.fi/whatsmeow v0.0.0-20260219150138-7ae702b1eed4/go.mod h1:mXCRFyPEPn4jqWz6Afirn8vY7DpHCPnlKq6I2cWwFHM=
go.uber.org/mock v0.6.0 h1:hyF9dfmbgIX5EfOdasqLsWD6xqpNZlXblLB/Dbnwv3Y=
@@ -244,16 +252,16 @@ golang.org/x/crypto v0.0.0-20200622213623-75b288015ac9/go.mod h1:LzIPMQfyMNhhGPh
golang.org/x/crypto v0.0.0-20210421170649-83a5a9bb288b/go.mod h1:T9bdIzuCu7OtxOm1hfPfRQxPLYneinmdGuTeoZ9dtd4=
golang.org/x/crypto v0.0.0-20210921155107-089bfa567519/go.mod h1:GvvjBRRGRdwPK5ydBHafDWAxML/pGHZbMvKqRZ5+Abc=
golang.org/x/crypto v0.16.0/go.mod h1:gCAAfMLgwOJRpTjQ2zCCt2OcSfYMTeZVSRtQlPC7Nq4=
-golang.org/x/crypto v0.48.0 h1:/VRzVqiRSggnhY7gNRxPauEQ5Drw9haKdM0jqfcCFts=
-golang.org/x/crypto v0.48.0/go.mod h1:r0kV5h3qnFPlQnBSrULhlsRfryS2pmewsg+XfMgkVos=
-golang.org/x/exp v0.0.0-20260212183809-81e46e3db34a h1:ovFr6Z0MNmU7nH8VaX5xqw+05ST2uO1exVfZPVqRC5o=
-golang.org/x/exp v0.0.0-20260212183809-81e46e3db34a/go.mod h1:K79w1Vqn7PoiZn+TkNpx3BUWUQksGO3JcVX6qIjytmA=
+golang.org/x/crypto v0.49.0 h1:+Ng2ULVvLHnJ/ZFEq4KdcDd/cfjrrjjNSXNzxg0Y4U4=
+golang.org/x/crypto v0.49.0/go.mod h1:ErX4dUh2UM+CFYiXZRTcMpEcN8b/1gxEuv3nODoYtCA=
+golang.org/x/exp v0.0.0-20260312153236-7ab1446f8b90 h1:jiDhWWeC7jfWqR9c/uplMOqJ0sbNlNWv0UkzE0vX1MA=
+golang.org/x/exp v0.0.0-20260312153236-7ab1446f8b90/go.mod h1:xE1HEv6b+1SCZ5/uscMRjUBKtIxworgEcEi+/n9NQDQ=
golang.org/x/mod v0.2.0/go.mod h1:s0Qsj1ACt9ePp/hMypM3fl4fZqREWJwdYDEqhRiZZUA=
golang.org/x/mod v0.3.0/go.mod h1:s0Qsj1ACt9ePp/hMypM3fl4fZqREWJwdYDEqhRiZZUA=
golang.org/x/mod v0.6.0-dev.0.20220419223038-86c51ed26bb4/go.mod h1:jJ57K6gSWd91VN4djpZkiMVwK6gcyfeH4XE8wZrZaV4=
golang.org/x/mod v0.8.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
-golang.org/x/mod v0.33.0 h1:tHFzIWbBifEmbwtGz65eaWyGiGZatSrT9prnU8DbVL8=
-golang.org/x/mod v0.33.0/go.mod h1:swjeQEj+6r7fODbD2cqrnje9PnziFuw4bmLbBZFrQ5w=
+golang.org/x/mod v0.34.0 h1:xIHgNUUnW6sYkcM5Jleh05DvLOtwc6RitGHbDk4akRI=
+golang.org/x/mod v0.34.0/go.mod h1:ykgH52iCZe79kzLLMhyCUzhMci+nQj+0XkbXpNYtVjY=
golang.org/x/net v0.0.0-20180906233101-161cd47e91fd/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20190404232315-eb5bcb51f2a3/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg=
golang.org/x/net v0.0.0-20190620200207-3b0461eec859/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s=
@@ -267,19 +275,19 @@ golang.org/x/net v0.0.0-20220722155237-a158d28d115b/go.mod h1:XRhObCWvk6IyKnWLug
golang.org/x/net v0.6.0/go.mod h1:2Tu9+aMcznHK/AK1HMvgo6xiTLG5rD5rZLDS+rp2Bjs=
golang.org/x/net v0.10.0/go.mod h1:0qNGK6F8kojg2nk9dLZ2mShWaEBan6FAoqfSigmmuDg=
golang.org/x/net v0.19.0/go.mod h1:CfAk/cbD4CthTvqiEl8NpboMuiuOYsAr/7NOjZJtv1U=
-golang.org/x/net v0.51.0 h1:94R/GTO7mt3/4wIKpcR5gkGmRLOuE/2hNGeWq/GBIFo=
-golang.org/x/net v0.51.0/go.mod h1:aamm+2QF5ogm02fjy5Bb7CQ0WMt1/WVM7FtyaTLlA9Y=
+golang.org/x/net v0.52.0 h1:He/TN1l0e4mmR3QqHMT2Xab3Aj3L9qjbhRm78/6jrW0=
+golang.org/x/net v0.52.0/go.mod h1:R1MAz7uMZxVMualyPXb+VaqGSa3LIaUqk0eEt3w36Sw=
golang.org/x/oauth2 v0.23.0/go.mod h1:XYTD2NtWslqkgxebSiOHnXEap4TF09sJSc7H1sXbhtI=
-golang.org/x/oauth2 v0.35.0 h1:Mv2mzuHuZuY2+bkyWXIHMfhNdJAdwW3FuWeCPYN5GVQ=
-golang.org/x/oauth2 v0.35.0/go.mod h1:lzm5WQJQwKZ3nwavOZ3IS5Aulzxi68dUSgRHujetwEA=
+golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs=
+golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q=
golang.org/x/sync v0.0.0-20180314180146-1d60e4601c6f/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20190423024810-112230192c58/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20190911185100-cd5d95a43a6e/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20201020160332-67f06af15bc9/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20220722155255-886fb9371eb4/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.1.0/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
-golang.org/x/sync v0.19.0 h1:vV+1eWNmZ5geRlYjzm2adRgW2/mcpevXNg50YZtPCE4=
-golang.org/x/sync v0.19.0/go.mod h1:9KTHXmSnoGruLpwFjVSX0lNNA75CykiMECbovNTZqGI=
+golang.org/x/sync v0.20.0 h1:e0PTpb7pjO8GAtTs2dQ6jYa5BWYlMuX047Dco/pItO4=
+golang.org/x/sync v0.20.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.0.0-20180909124046-d0be0721c37e/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20190215142949-d0b11bdaac8a/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20190412213103-97732733099d/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
@@ -301,15 +309,15 @@ golang.org/x/sys v0.6.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.8.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.12.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.15.0/go.mod h1:/VUhepiaJMQUp4+oa/7Zr1D23ma6VTLIYjOOTFZPUcA=
-golang.org/x/sys v0.41.0 h1:Ivj+2Cp/ylzLiEU89QhWblYnOE9zerudt9Ftecq2C6k=
-golang.org/x/sys v0.41.0/go.mod h1:OgkHotnGiDImocRcuBABYBEXf8A9a87e/uXjp9XT3ks=
+golang.org/x/sys v0.42.0 h1:omrd2nAlyT5ESRdCLYdm3+fMfNFE/+Rf4bDIQImRJeo=
+golang.org/x/sys v0.42.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210927222741-03fcf44c2211/go.mod h1:jbD1KX2456YbFQfuXm/mYQcufACuNUgVhRMnK/tPxf8=
golang.org/x/term v0.5.0/go.mod h1:jMB1sMXY+tzblOD4FWmEbocvup2/aLOaQEp7JmGp78k=
golang.org/x/term v0.8.0/go.mod h1:xPskH00ivmX89bAKVGSKKtLOWNx2+17Eiy94tnKShWo=
golang.org/x/term v0.15.0/go.mod h1:BDl952bC7+uMoWR75FIrCDx79TPU9oHkTZ9yRbYOrX0=
-golang.org/x/term v0.40.0 h1:36e4zGLqU4yhjlmxEaagx2KuYbJq3EwY8K943ZsHcvg=
-golang.org/x/term v0.40.0/go.mod h1:w2P8uVp06p2iyKKuvXIm7N/y0UCRt3UfJTfZ7oOpglM=
+golang.org/x/term v0.41.0 h1:QCgPso/Q3RTJx2Th4bDLqML4W6iJiaXFq2/ftQF13YU=
+golang.org/x/term v0.41.0/go.mod h1:3pfBgksrReYfZ5lvYM0kSO0LIkAl4Yl2bXOkKP7Ec2A=
golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.3/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
golang.org/x/text v0.3.6/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
@@ -317,8 +325,8 @@ golang.org/x/text v0.3.7/go.mod h1:u+2+/6zg+i71rQMx5EYifcz6MCKuco9NR6JIITiCfzQ=
golang.org/x/text v0.7.0/go.mod h1:mrYo+phRRbMaCq/xk9113O4dZlRixOauAjOtrjsXDZ8=
golang.org/x/text v0.9.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.14.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
-golang.org/x/text v0.34.0 h1:oL/Qq0Kdaqxa1KbNeMKwQq0reLCCaFtqu2eNuSeNHbk=
-golang.org/x/text v0.34.0/go.mod h1:homfLqTYRFyVYemLBFl5GgL/DWEiH5wcsQ5gSh1yziA=
+golang.org/x/text v0.35.0 h1:JOVx6vVDFokkpaq1AEptVzLTpDe9KGpj5tR4/X+ybL8=
+golang.org/x/text v0.35.0/go.mod h1:khi/HExzZJ2pGnjenulevKNX1W67CUy0AsXcNubPGCA=
golang.org/x/time v0.14.0 h1:MRx4UaLrDotUKUdCIqzPC48t1Y9hANFKIRpNx+Te8PI=
golang.org/x/time v0.14.0/go.mod h1:eL/Oa2bBBK0TkX57Fyni+NgnyQQN4LitPmob2Hjnqw4=
golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
@@ -328,8 +336,8 @@ golang.org/x/tools v0.0.0-20201224043029-2b0845dc783e/go.mod h1:emZCQorbCU4vsT4f
golang.org/x/tools v0.0.0-20210106214847-113979e3529a/go.mod h1:emZCQorbCU4vsT4fOWvOPXz4eW1wZW4PmDk9uLelYpA=
golang.org/x/tools v0.1.12/go.mod h1:hNGJHUnrk76NpqgfD5Aqm5Crs+Hm0VOH/i9J2+nxYbc=
golang.org/x/tools v0.6.0/go.mod h1:Xwgl3UAJ/d3gWutnCtw505GrjyAbvKui8lOU390QaIU=
-golang.org/x/tools v0.42.0 h1:uNgphsn75Tdz5Ji2q36v/nsFSfR/9BRFvqhGBaJGd5k=
-golang.org/x/tools v0.42.0/go.mod h1:Ma6lCIwGZvHK6XtgbswSoWroEkhugApmsXyrUmBhfr0=
+golang.org/x/tools v0.43.0 h1:12BdW9CeB3Z+J/I/wj34VMl8X+fEXBxVR90JeMX5E7s=
+golang.org/x/tools v0.43.0/go.mod h1:uHkMso649BX2cZK6+RpuIPXS3ho2hZo4FVwfoy1vIk0=
golang.org/x/xerrors v0.0.0-20190717185122-a985d3407aa7/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20191011141410-1b5146add898/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20191204190536-9bdfabe68543/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
@@ -354,12 +362,13 @@ gopkg.in/tomb.v1 v1.0.0-20141024135613-dd632973f1e7/go.mod h1:dt/ZhP58zS4L8KSrWD
gopkg.in/yaml.v2 v2.2.2/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI=
gopkg.in/yaml.v2 v2.2.4/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI=
gopkg.in/yaml.v2 v2.3.0/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI=
+gopkg.in/yaml.v2 v2.4.0 h1:D8xgwECY7CYvx+Y2n4sBz93Jn9JRvxdiyyo8CTfuKaY=
gopkg.in/yaml.v2 v2.4.0/go.mod h1:RDklbk79AGWmwhnvt/jBztapEOGDOx6ZbXqjP6csGnQ=
gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
-maunium.net/go/mautrix v0.26.3 h1:tWZih6Vjw0qGTWuPmg9JUrQPzViTNDPGQLVc5UXC4nk=
-maunium.net/go/mautrix v0.26.3/go.mod h1:v5ZdDoCwUpNqEj5OrhEoUa3L1kEddKPaAya9TgGXN38=
+maunium.net/go/mautrix v0.26.4 h1:enHSnkf0L2V9+VnfJfNhKSReSW6pBKS/x3Su+v+Vovs=
+maunium.net/go/mautrix v0.26.4/go.mod h1:YWw8NWTszsbyFAznboicBObwHPgTSLcuTbVX2kY7U2M=
modernc.org/cc/v4 v4.27.1 h1:9W30zRlYrefrDV2JE2O8VDtJ1yPGownxciz5rrbQZis=
modernc.org/cc/v4 v4.27.1/go.mod h1:uVtb5OGqUKpoLWhqwNQo/8LwvoiEBLvZXIQ/SmO6mL0=
modernc.org/ccgo/v4 v4.30.1 h1:4r4U1J6Fhj98NKfSjnPUN7Ze2c6MnAdL0hWw6+LrJpc=
diff --git a/pkg/agent/context.go b/pkg/agent/context.go
index cb566f02b..022230d41 100644
--- a/pkg/agent/context.go
+++ b/pkg/agent/context.go
@@ -52,7 +52,7 @@ func (cb *ContextBuilder) WithToolDiscovery(useBM25, useRegex bool) *ContextBuil
}
func getGlobalConfigDir() string {
- if home := os.Getenv("PICOCLAW_HOME"); home != "" {
+ if home := os.Getenv(config.EnvHome); home != "" {
return home
}
home, err := os.UserHomeDir()
@@ -65,7 +65,7 @@ func getGlobalConfigDir() string {
func NewContextBuilder(workspace string) *ContextBuilder {
// builtin skills: skills directory in current project
// Use the skills/ directory under the current working directory
- builtinSkillsDir := strings.TrimSpace(os.Getenv("PICOCLAW_BUILTIN_SKILLS"))
+ builtinSkillsDir := strings.TrimSpace(os.Getenv(config.EnvBuiltinSkills))
if builtinSkillsDir == "" {
wd, _ := os.Getwd()
builtinSkillsDir = filepath.Join(wd, "skills")
@@ -469,7 +469,23 @@ func (cb *ContextBuilder) LoadBootstrapFiles() string {
//
// See: https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching
// See: https://platform.openai.com/docs/guides/prompt-caching
-func (cb *ContextBuilder) buildDynamicContext(channel, chatID string) string {
+func formatCurrentSenderLine(senderID, senderDisplayName string) string {
+ senderID = strings.TrimSpace(senderID)
+ senderDisplayName = strings.TrimSpace(senderDisplayName)
+
+ switch {
+ case senderDisplayName != "" && senderID != "":
+ return fmt.Sprintf("Current sender: %s (ID: %s)", senderDisplayName, senderID)
+ case senderDisplayName != "":
+ return fmt.Sprintf("Current sender: %s", senderDisplayName)
+ case senderID != "":
+ return fmt.Sprintf("Current sender: %s", senderID)
+ default:
+ return ""
+ }
+}
+
+func (cb *ContextBuilder) buildDynamicContext(channel, chatID, senderID, senderDisplayName string) string {
now := time.Now().Format("2006-01-02 15:04 (Monday)")
rt := fmt.Sprintf("%s %s, Go %s", runtime.GOOS, runtime.GOARCH, runtime.Version())
@@ -479,6 +495,9 @@ func (cb *ContextBuilder) buildDynamicContext(channel, chatID string) string {
if channel != "" && chatID != "" {
fmt.Fprintf(&sb, "\n\n## Current Session\nChannel: %s\nChat ID: %s", channel, chatID)
}
+ if senderLine := formatCurrentSenderLine(senderID, senderDisplayName); senderLine != "" {
+ fmt.Fprintf(&sb, "\n\n## Current Sender\n%s", senderLine)
+ }
return sb.String()
}
@@ -488,7 +507,7 @@ func (cb *ContextBuilder) BuildMessages(
summary string,
currentMessage string,
media []string,
- channel, chatID string,
+ channel, chatID, senderID, senderDisplayName string,
) []providers.Message {
messages := []providers.Message{}
@@ -504,7 +523,7 @@ func (cb *ContextBuilder) BuildMessages(
staticPrompt := cb.BuildSystemPromptWithCache()
// Build short dynamic context (time, runtime, session) — changes per request
- dynamicCtx := cb.buildDynamicContext(channel, chatID)
+ dynamicCtx := cb.buildDynamicContext(channel, chatID, senderID, senderDisplayName)
// Compose a single system message: static (cached) + dynamic + optional summary.
// Keeping all system content in one message ensures every provider adapter can
diff --git a/pkg/agent/context_cache_test.go b/pkg/agent/context_cache_test.go
index 1f9423a3a..81a1534b9 100644
--- a/pkg/agent/context_cache_test.go
+++ b/pkg/agent/context_cache_test.go
@@ -82,7 +82,7 @@ func TestSingleSystemMessage(t *testing.T) {
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
- msgs := cb.BuildMessages(tt.history, tt.summary, tt.message, nil, "test", "chat1")
+ msgs := cb.BuildMessages(tt.history, tt.summary, tt.message, nil, "test", "chat1", "", "")
systemCount := 0
for _, m := range msgs {
@@ -126,6 +126,68 @@ func TestSingleSystemMessage(t *testing.T) {
}
}
+func TestBuildMessages_CurrentSenderDynamicContext(t *testing.T) {
+ tmpDir := setupWorkspace(t, map[string]string{
+ "IDENTITY.md": "# Identity\nTest agent.",
+ })
+ defer os.RemoveAll(tmpDir)
+
+ cb := NewContextBuilder(tmpDir)
+
+ tests := []struct {
+ name string
+ senderID string
+ senderDisplayName string
+ wantLine string
+ wantSection bool
+ }{
+ {
+ name: "both id and display name",
+ senderID: "feishu:ou_xxx",
+ senderDisplayName: "Zhang San",
+ wantLine: "Current sender: Zhang San (ID: feishu:ou_xxx)",
+ wantSection: true,
+ },
+ {
+ name: "display name only",
+ senderDisplayName: "Alice",
+ wantLine: "Current sender: Alice",
+ wantSection: true,
+ },
+ {
+ name: "id only",
+ senderID: "discord:123",
+ wantLine: "Current sender: discord:123",
+ wantSection: true,
+ },
+ {
+ name: "no sender info",
+ wantSection: false,
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ msgs := cb.BuildMessages(nil, "", "hello", nil, "discord", "chat1", tt.senderID, tt.senderDisplayName)
+ sys := msgs[0].Content
+
+ if tt.wantSection {
+ if !strings.Contains(sys, "## Current Sender") {
+ t.Fatalf("system prompt missing Current Sender section:\n%s", sys)
+ }
+ if !strings.Contains(sys, tt.wantLine) {
+ t.Fatalf("system prompt missing sender line %q:\n%s", tt.wantLine, sys)
+ }
+ return
+ }
+
+ if strings.Contains(sys, "## Current Sender") {
+ t.Fatalf("system prompt should omit Current Sender section:\n%s", sys)
+ }
+ })
+ }
+}
+
// TestMtimeAutoInvalidation verifies that the cache detects source file changes
// via mtime without requiring explicit InvalidateCache().
// Fix: original implementation had no auto-invalidation — edits to bootstrap files,
@@ -576,7 +638,7 @@ func TestConcurrentBuildSystemPromptWithCache(t *testing.T) {
}
// Also exercise BuildMessages concurrently
- msgs := cb.BuildMessages(nil, "", "hello", nil, "test", "chat")
+ msgs := cb.BuildMessages(nil, "", "hello", nil, "test", "chat", "", "")
if len(msgs) < 2 {
errs <- "BuildMessages returned fewer than 2 messages"
return
@@ -664,6 +726,6 @@ func BenchmarkBuildMessagesWithCache(b *testing.B) {
b.ResetTimer()
for i := 0; i < b.N; i++ {
- _ = cb.BuildMessages(history, "summary", "new message", nil, "cli", "test")
+ _ = cb.BuildMessages(history, "summary", "new message", nil, "cli", "test", "", "")
}
}
diff --git a/pkg/agent/events.go b/pkg/agent/events.go
index 95e4c90d0..f4562b360 100644
--- a/pkg/agent/events.go
+++ b/pkg/agent/events.go
@@ -43,6 +43,8 @@ const (
EventKindSubTurnEnd
// EventKindSubTurnResultDelivered is emitted when a sub-turn result is delivered.
EventKindSubTurnResultDelivered
+ // EventKindSubTurnOrphan is emitted when a sub-turn result cannot be delivered.
+ EventKindSubTurnOrphan
// EventKindError is emitted when a turn encounters an execution error.
EventKindError
@@ -67,6 +69,7 @@ var eventKindNames = [...]string{
"subturn_spawn",
"subturn_end",
"subturn_result_delivered",
+ "subturn_orphan",
"error",
}
@@ -236,8 +239,9 @@ type InterruptReceivedPayload struct {
// SubTurnSpawnPayload describes the creation of a child turn.
type SubTurnSpawnPayload struct {
- AgentID string
- Label string
+ AgentID string
+ Label string
+ ParentTurnID string
}
// SubTurnEndPayload describes the completion of a child turn.
@@ -253,6 +257,13 @@ type SubTurnResultDeliveredPayload struct {
ContentLen int
}
+// SubTurnOrphanPayload describes a sub-turn result that could not be delivered.
+type SubTurnOrphanPayload struct {
+ ParentTurnID string
+ ChildTurnID string
+ Reason string
+}
+
// ErrorPayload describes an execution error inside the agent loop.
type ErrorPayload struct {
Stage string
diff --git a/pkg/agent/instance.go b/pkg/agent/instance.go
index c34f9b4a4..34d401186 100644
--- a/pkg/agent/instance.go
+++ b/pkg/agent/instance.go
@@ -3,13 +3,14 @@ package agent
import (
"context"
"fmt"
- "log"
"os"
"path/filepath"
"regexp"
"strings"
"github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/logger"
+ "github.com/sipeed/picoclaw/pkg/media"
"github.com/sipeed/picoclaw/pkg/memory"
"github.com/sipeed/picoclaw/pkg/providers"
"github.com/sipeed/picoclaw/pkg/routing"
@@ -66,7 +67,7 @@ func NewAgentInstance(
readRestrict := restrict && !defaults.AllowReadOutsideWorkspace
// Compile path whitelist patterns from config.
- allowReadPaths := compilePatterns(cfg.Tools.AllowReadPaths)
+ allowReadPaths := buildAllowReadPatterns(cfg)
allowWritePaths := compilePatterns(cfg.Tools.AllowWritePaths)
toolsRegistry := tools.NewToolRegistry()
@@ -82,11 +83,13 @@ func NewAgentInstance(
toolsRegistry.Register(tools.NewListDirTool(workspace, readRestrict, allowReadPaths))
}
if cfg.Tools.IsToolEnabled("exec") {
- execTool, err := tools.NewExecToolWithConfig(workspace, restrict, cfg)
+ execTool, err := tools.NewExecToolWithConfig(workspace, restrict, cfg, allowReadPaths)
if err != nil {
- log.Fatalf("Critical error: unable to initialize exec tool: %v", err)
+ logger.ErrorCF("agent", "Failed to initialize exec tool; continuing without exec",
+ map[string]any{"error": err.Error()})
+ } else {
+ toolsRegistry.Register(execTool)
}
- toolsRegistry.Register(execTool)
}
if cfg.Tools.IsToolEnabled("edit_file") {
@@ -160,59 +163,14 @@ func NewAgentInstance(
}
// Resolve fallback candidates
- modelCfg := providers.ModelConfig{
- Primary: model,
- Fallbacks: fallbacks,
- }
- resolveFromModelList := func(raw string) (string, bool) {
- ensureProtocol := func(model string) string {
- model = strings.TrimSpace(model)
- if model == "" {
- return ""
- }
- if strings.Contains(model, "/") {
- return model
- }
- return "openai/" + model
- }
-
- raw = strings.TrimSpace(raw)
- if raw == "" {
- return "", false
- }
-
- if cfg != nil {
- if mc, err := cfg.GetModelConfig(raw); err == nil && mc != nil && strings.TrimSpace(mc.Model) != "" {
- return ensureProtocol(mc.Model), true
- }
-
- for i := range cfg.ModelList {
- fullModel := strings.TrimSpace(cfg.ModelList[i].Model)
- if fullModel == "" {
- continue
- }
- if fullModel == raw {
- return ensureProtocol(fullModel), true
- }
- _, modelID := providers.ExtractProtocol(fullModel)
- if modelID == raw {
- return ensureProtocol(fullModel), true
- }
- }
- }
-
- return "", false
- }
-
- candidates := providers.ResolveCandidatesWithLookup(modelCfg, defaults.Provider, resolveFromModelList)
+ candidates := resolveModelCandidates(cfg, defaults.Provider, model, fallbacks)
// Model routing setup: pre-resolve light model candidates at creation time
// to avoid repeated model_list lookups on every incoming message.
var router *routing.Router
var lightCandidates []providers.FallbackCandidate
if rc := defaults.Routing; rc != nil && rc.Enabled && rc.LightModel != "" {
- lightModelCfg := providers.ModelConfig{Primary: rc.LightModel}
- resolved := providers.ResolveCandidatesWithLookup(lightModelCfg, defaults.Provider, resolveFromModelList)
+ resolved := resolveModelCandidates(cfg, defaults.Provider, rc.LightModel, nil)
if len(resolved) > 0 {
router = routing.New(routing.RouterConfig{
LightModel: rc.LightModel,
@@ -220,8 +178,8 @@ func NewAgentInstance(
})
lightCandidates = resolved
} else {
- log.Printf("routing: light_model %q not found in model_list — routing disabled for agent %q",
- rc.LightModel, agentID)
+ logger.WarnCF("agent", "Routing light model not found; routing disabled",
+ map[string]any{"light_model": rc.LightModel, "agent_id": agentID})
}
}
@@ -293,6 +251,28 @@ func compilePatterns(patterns []string) []*regexp.Regexp {
return compiled
}
+func buildAllowReadPatterns(cfg *config.Config) []*regexp.Regexp {
+ var configured []string
+ if cfg != nil {
+ configured = cfg.Tools.AllowReadPaths
+ }
+
+ compiled := compilePatterns(configured)
+ mediaDirPattern := regexp.MustCompile(mediaTempDirPattern())
+ for _, pattern := range compiled {
+ if pattern.String() == mediaDirPattern.String() {
+ return compiled
+ }
+ }
+
+ return append(compiled, mediaDirPattern)
+}
+
+func mediaTempDirPattern() string {
+ sep := regexp.QuoteMeta(string(os.PathSeparator))
+ return "^" + regexp.QuoteMeta(filepath.Clean(media.TempDir())) + "(?:" + sep + "|$)"
+}
+
// Close releases resources held by the agent's session store.
func (a *AgentInstance) Close() error {
if a.Sessions != nil {
@@ -308,7 +288,8 @@ func (a *AgentInstance) Close() error {
func initSessionStore(dir string) session.SessionStore {
store, err := memory.NewJSONLStore(dir)
if err != nil {
- log.Printf("memory: init store: %v; using json sessions", err)
+ logger.WarnCF("agent", "Memory JSONL store init failed; falling back to json sessions",
+ map[string]any{"error": err.Error()})
return session.NewSessionManager(dir)
}
@@ -316,11 +297,12 @@ func initSessionStore(dir string) session.SessionStore {
// Migration failure means the store could not write data.
// Fall back to SessionManager to avoid a split state where
// some sessions are in JSONL and others remain in JSON.
- log.Printf("memory: migration failed: %v; falling back to json sessions", merr)
+ logger.WarnCF("agent", "Memory migration failed; falling back to json sessions",
+ map[string]any{"error": merr.Error()})
store.Close()
return session.NewSessionManager(dir)
} else if n > 0 {
- log.Printf("memory: migrated %d session(s) to jsonl", n)
+ logger.InfoCF("agent", "Memory migrated to JSONL", map[string]any{"sessions_migrated": n})
}
return session.NewJSONLBackend(store)
diff --git a/pkg/agent/instance_test.go b/pkg/agent/instance_test.go
index 4f41ecd1c..b3318ad1f 100644
--- a/pkg/agent/instance_test.go
+++ b/pkg/agent/instance_test.go
@@ -1,10 +1,14 @@
package agent
import (
+ "context"
"os"
+ "path/filepath"
+ "strings"
"testing"
"github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/media"
)
func TestNewAgentInstance_UsesDefaultsTemperatureAndMaxTokens(t *testing.T) {
@@ -160,3 +164,119 @@ func TestNewAgentInstance_ResolveCandidatesFromModelListAlias(t *testing.T) {
})
}
}
+
+func TestNewAgentInstance_AllowsMediaTempDirForReadListAndExec(t *testing.T) {
+ workspace := t.TempDir()
+ mediaDir := media.TempDir()
+ if err := os.MkdirAll(mediaDir, 0o700); err != nil {
+ t.Fatalf("MkdirAll(mediaDir) error = %v", err)
+ }
+
+ mediaFile, err := os.CreateTemp(mediaDir, "instance-tool-*.txt")
+ if err != nil {
+ t.Fatalf("CreateTemp(mediaDir) error = %v", err)
+ }
+ mediaPath := mediaFile.Name()
+ if _, err := mediaFile.WriteString("attachment content"); err != nil {
+ mediaFile.Close()
+ t.Fatalf("WriteString(mediaFile) error = %v", err)
+ }
+ if err := mediaFile.Close(); err != nil {
+ t.Fatalf("Close(mediaFile) error = %v", err)
+ }
+ t.Cleanup(func() { _ = os.Remove(mediaPath) })
+
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Workspace: workspace,
+ ModelName: "test-model",
+ RestrictToWorkspace: true,
+ },
+ },
+ Tools: config.ToolsConfig{
+ ReadFile: config.ReadFileToolConfig{Enabled: true},
+ ListDir: config.ToolConfig{Enabled: true},
+ Exec: config.ExecConfig{
+ ToolConfig: config.ToolConfig{Enabled: true},
+ EnableDenyPatterns: true,
+ AllowRemote: true,
+ },
+ },
+ }
+
+ agent := NewAgentInstance(nil, &cfg.Agents.Defaults, cfg, &mockProvider{})
+
+ readTool, ok := agent.Tools.Get("read_file")
+ if !ok {
+ t.Fatal("read_file tool not registered")
+ }
+ readResult := readTool.Execute(context.Background(), map[string]any{"path": mediaPath})
+ if readResult.IsError {
+ t.Fatalf("read_file should allow media temp dir, got: %s", readResult.ForLLM)
+ }
+ if !strings.Contains(readResult.ForLLM, "attachment content") {
+ t.Fatalf("read_file output missing media content: %s", readResult.ForLLM)
+ }
+
+ listTool, ok := agent.Tools.Get("list_dir")
+ if !ok {
+ t.Fatal("list_dir tool not registered")
+ }
+ listResult := listTool.Execute(context.Background(), map[string]any{"path": mediaDir})
+ if listResult.IsError {
+ t.Fatalf("list_dir should allow media temp dir, got: %s", listResult.ForLLM)
+ }
+ if !strings.Contains(listResult.ForLLM, filepath.Base(mediaPath)) {
+ t.Fatalf("list_dir output missing media file: %s", listResult.ForLLM)
+ }
+
+ execTool, ok := agent.Tools.Get("exec")
+ if !ok {
+ t.Fatal("exec tool not registered")
+ }
+ execResult := execTool.Execute(context.Background(), map[string]any{
+ "command": "cat " + filepath.Base(mediaPath),
+ "working_dir": mediaDir,
+ })
+ if execResult.IsError {
+ t.Fatalf("exec should allow media temp dir, got: %s", execResult.ForLLM)
+ }
+ if !strings.Contains(execResult.ForLLM, "attachment content") {
+ t.Fatalf("exec output missing media content: %s", execResult.ForLLM)
+ }
+}
+
+func TestNewAgentInstance_InvalidExecConfigDoesNotExit(t *testing.T) {
+ workspace := t.TempDir()
+
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Workspace: workspace,
+ ModelName: "test-model",
+ },
+ },
+ Tools: config.ToolsConfig{
+ ReadFile: config.ReadFileToolConfig{Enabled: true},
+ Exec: config.ExecConfig{
+ ToolConfig: config.ToolConfig{Enabled: true},
+ EnableDenyPatterns: true,
+ CustomDenyPatterns: []string{"[invalid-regex"},
+ },
+ },
+ }
+
+ agent := NewAgentInstance(nil, &cfg.Agents.Defaults, cfg, &mockProvider{})
+ if agent == nil {
+ t.Fatal("expected agent instance, got nil")
+ }
+
+ if _, ok := agent.Tools.Get("exec"); ok {
+ t.Fatal("exec tool should not be registered when exec config is invalid")
+ }
+
+ if _, ok := agent.Tools.Get("read_file"); !ok {
+ t.Fatal("read_file tool should still be registered")
+ }
+}
diff --git a/pkg/agent/loop.go b/pkg/agent/loop.go
index fb9224298..840aa8fa1 100644
--- a/pkg/agent/loop.go
+++ b/pkg/agent/loop.go
@@ -35,12 +35,17 @@ import (
)
type AgentLoop struct {
- bus *bus.MessageBus
- cfg *config.Config
- registry *AgentRegistry
- state *state.Manager
- eventBus *EventBus
- hooks *HookManager
+ // Core dependencies
+ bus *bus.MessageBus
+ cfg *config.Config
+ registry *AgentRegistry
+ state *state.Manager
+
+ // Event system (from Incoming)
+ eventBus *EventBus
+ hooks *HookManager
+
+ // Runtime state
running atomic.Bool
summarizing sync.Map
fallback *providers.FallbackChain
@@ -52,26 +57,34 @@ type AgentLoop struct {
hookRuntime hookRuntime
steering *steeringQueue
mu sync.RWMutex
- activeTurnMu sync.RWMutex
- activeTurn *turnState
+
+ // Concurrent turn management (from HEAD)
+ activeTurnStates sync.Map // key: sessionKey (string), value: *turnState
+ subTurnCounter atomic.Int64 // Counter for generating unique SubTurn IDs
+
+ // Turn tracking (from Incoming)
turnSeq atomic.Uint64
- // Track active requests for safe provider cleanup
activeRequests sync.WaitGroup
+
+ reloadFunc func() error
}
// processOptions configures how a message is processed
type processOptions struct {
- SessionKey string // Session identifier for history/context
- Channel string // Target channel for tool execution
- ChatID string // Target chat ID for tool execution
- UserMessage string // User message content (may include prefix)
- Media []string // media:// refs from inbound message
- InitialSteeringMessages []providers.Message
- DefaultResponse string // Response when LLM returns empty
- EnableSummary bool // Whether to trigger summarization
- SendResponse bool // Whether to send response via bus
- NoHistory bool // If true, don't load session history (for heartbeat)
- SkipInitialSteeringPoll bool // If true, skip the steering poll at loop start (used by Continue)
+ SessionKey string // Session identifier for history/context
+ Channel string // Target channel for tool execution
+ ChatID string // Target chat ID for tool execution
+ SenderID string // Current sender ID for dynamic context
+ SenderDisplayName string // Current sender display name for dynamic context
+ UserMessage string // User message content (may include prefix)
+ SystemPromptOverride string // Override the default system prompt (Used by SubTurns)
+ Media []string // media:// refs from inbound message
+ InitialSteeringMessages []providers.Message // Steering messages from refactor/agent
+ DefaultResponse string // Response when LLM returns empty
+ EnableSummary bool // Whether to trigger summarization
+ SendResponse bool // Whether to send response via bus
+ NoHistory bool // If true, don't load session history (for heartbeat)
+ SkipInitialSteeringPoll bool // If true, skip the steering poll at loop start (used by Continue)
}
type continuationTarget struct {
@@ -81,7 +94,8 @@ type continuationTarget struct {
}
const (
- defaultResponse = "I've completed processing but have no response to give. Increase `max_tool_iterations` in config.json."
+ defaultResponse = "The model returned an empty response. This may indicate a provider error or token limit."
+ toolLimitResponse = "I've reached `max_tool_iterations` without a final response. Increase `max_tool_iterations` in config.json if this task needs more tool steps."
sessionKeyAgentPrefix = "agent:"
metadataKeyAccountID = "account_id"
metadataKeyGuildID = "guild_id"
@@ -97,9 +111,6 @@ func NewAgentLoop(
) *AgentLoop {
registry := NewAgentRegistry(cfg, provider)
- // Register shared tools to all agents
- registerSharedTools(cfg, msgBus, registry, provider)
-
// Set up shared fallback chain
cooldown := providers.NewCooldownTracker()
fallbackChain := providers.NewFallbackChain(cooldown)
@@ -126,16 +137,22 @@ func NewAgentLoop(
al.hooks = NewHookManager(eventBus)
configureHookManagerFromConfig(al.hooks, cfg)
+ // Register shared tools to all agents (now that al is created)
+ registerSharedTools(al, cfg, msgBus, registry, provider)
+
return al
}
// registerSharedTools registers tools that are shared across all agents (web, message, spawn).
func registerSharedTools(
+ al *AgentLoop,
cfg *config.Config,
msgBus *bus.MessageBus,
registry *AgentRegistry,
provider providers.LLMProvider,
) {
+ allowReadPaths := buildAllowReadPatterns(cfg)
+
for _, agentID := range registry.ListAgentIDs() {
agent, ok := registry.GetAgent(agentID)
if !ok {
@@ -176,7 +193,12 @@ func registerSharedTools(
}
}
if cfg.Tools.IsToolEnabled("web_fetch") {
- fetchTool, err := tools.NewWebFetchToolWithProxy(50000, cfg.Tools.Web.Proxy, cfg.Tools.Web.FetchLimitBytes)
+ fetchTool, err := tools.NewWebFetchToolWithProxy(
+ 50000,
+ cfg.Tools.Web.Proxy,
+ cfg.Tools.Web.Format,
+ cfg.Tools.Web.FetchLimitBytes,
+ cfg.Tools.Web.PrivateHostWhitelist)
if err != nil {
logger.ErrorCF("agent", "Failed to create web fetch tool", map[string]any{"error": err.Error()})
} else {
@@ -214,6 +236,7 @@ func registerSharedTools(
cfg.Agents.Defaults.RestrictToWorkspace,
cfg.Agents.Defaults.GetMaxMediaSize(),
nil,
+ allowReadPaths,
)
agent.Tools.Register(sendFileTool)
}
@@ -241,20 +264,99 @@ func registerSharedTools(
}
}
- // Spawn tool with allowlist checker
- if cfg.Tools.IsToolEnabled("spawn") {
- if cfg.Tools.IsToolEnabled("subagent") {
- subagentManager := tools.NewSubagentManager(provider, agent.Model, agent.Workspace)
- subagentManager.SetLLMOptions(agent.MaxTokens, agent.Temperature)
+ // Spawn and spawn_status tools share a SubagentManager.
+ // Construct it when either tool is enabled (both require subagent).
+ spawnEnabled := cfg.Tools.IsToolEnabled("spawn")
+ spawnStatusEnabled := cfg.Tools.IsToolEnabled("spawn_status")
+ if (spawnEnabled || spawnStatusEnabled) && cfg.Tools.IsToolEnabled("subagent") {
+ subagentManager := tools.NewSubagentManager(provider, agent.Model, agent.Workspace)
+ subagentManager.SetLLMOptions(agent.MaxTokens, agent.Temperature)
+
+ // Set the spawner that links into AgentLoop's turnState
+ subagentManager.SetSpawner(func(
+ ctx context.Context,
+ task, label, targetAgentID string,
+ tls *tools.ToolRegistry,
+ maxTokens int,
+ temperature float64,
+ hasMaxTokens, hasTemperature bool,
+ ) (*tools.ToolResult, error) {
+ // 1. Recover parent Turn State from Context
+ parentTS := turnStateFromContext(ctx)
+ if parentTS == nil {
+ // Fallback: If no turnState exists in context, create an isolated ad-hoc root turn state
+ // so that the tool can still function outside of an agent loop (e.g. tests, raw invocations).
+ parentTS = &turnState{
+ ctx: ctx,
+ turnID: "adhoc-root",
+ depth: 0,
+ session: nil, // Ephemeral session not needed for adhoc spawn
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, 5),
+ }
+ }
+
+ // 2. Build Tools slice from registry
+ var tlSlice []tools.Tool
+ for _, name := range tls.List() {
+ if t, ok := tls.Get(name); ok {
+ tlSlice = append(tlSlice, t)
+ }
+ }
+
+ // 3. System Prompt
+ systemPrompt := "You are a subagent. Complete the given task independently and report the result.\n" +
+ "You have access to tools - use them as needed to complete your task.\n" +
+ "After completing the task, provide a clear summary of what was done.\n\n" +
+ "Task: " + task
+
+ // 4. Resolve Model
+ modelToUse := agent.Model
+ if targetAgentID != "" {
+ if targetAgent, ok := al.GetRegistry().GetAgent(targetAgentID); ok {
+ modelToUse = targetAgent.Model
+ }
+ }
+
+ // 5. Build SubTurnConfig
+ cfg := SubTurnConfig{
+ Model: modelToUse,
+ Tools: tlSlice,
+ SystemPrompt: systemPrompt,
+ }
+ if hasMaxTokens {
+ cfg.MaxTokens = maxTokens
+ }
+
+ // 6. Spawn SubTurn
+ return spawnSubTurn(ctx, al, parentTS, cfg)
+ })
+
+ // Clone the parent's tool registry so subagents can use all
+ // tools registered so far (file, web, etc.) but NOT spawn/
+ // spawn_status which are added below — preventing recursive
+ // subagent spawning.
+ subagentManager.SetTools(agent.Tools.Clone())
+ if spawnEnabled {
spawnTool := tools.NewSpawnTool(subagentManager)
+ spawnTool.SetSpawner(NewSubTurnSpawner(al))
currentAgentID := agentID
spawnTool.SetAllowlistChecker(func(targetAgentID string) bool {
return registry.CanSpawnSubagent(currentAgentID, targetAgentID)
})
+
agent.Tools.Register(spawnTool)
- } else {
- logger.WarnCF("agent", "spawn tool requires subagent to be enabled", nil)
+
+ // Also register the synchronous subagent tool
+ subagentTool := tools.NewSubagentTool(subagentManager)
+ subagentTool.SetSpawner(NewSubTurnSpawner(al))
+ agent.Tools.Register(subagentTool)
}
+ if spawnStatusEnabled {
+ agent.Tools.Register(tools.NewSpawnStatusTool(subagentManager))
+ }
+ } else if (spawnEnabled || spawnStatusEnabled) && !cfg.Tools.IsToolEnabled("subagent") {
+ logger.WarnCF("agent", "spawn/spawn_status tools require subagent to be enabled", nil)
}
}
}
@@ -273,10 +375,9 @@ func (al *AgentLoop) Run(ctx context.Context) error {
select {
case <-ctx.Done():
return nil
- default:
- msg, ok := al.bus.ConsumeInbound(ctx)
+ case msg, ok := <-al.bus.InboundChan():
if !ok {
- continue
+ return nil
}
// Start a goroutine that drains the bus while processMessage is
@@ -291,6 +392,11 @@ func (al *AgentLoop) Run(ctx context.Context) error {
// Process message
func() {
+ defer func() {
+ if al.channelManager != nil {
+ al.channelManager.InvokeTypingStop(msg.Channel, msg.ChatID)
+ }
+ }()
// TODO: Re-enable media cleanup after inbound media is properly consumed by the agent.
// Currently disabled because files are deleted before the LLM can access their content.
// defer func() {
@@ -395,21 +501,47 @@ func (al *AgentLoop) Run(ctx context.Context) error {
al.publishResponseIfNeeded(ctx, target.Channel, target.ChatID, finalResponse)
}
}()
+ default:
+ time.Sleep(time.Microsecond * 200)
}
}
return nil
}
-// drainBusToSteering continuously consumes inbound messages and redirects
-// messages from the active scope into the steering queue. Messages from other
-// scopes are requeued so they can be processed normally after the active turn.
+// drainBusToSteering consumes inbound messages and redirects messages from the
+// active scope into the steering queue. Messages from other scopes are requeued
+// so they can be processed normally after the active turn. It drains all
+// immediately available messages, blocking for the first one until ctx is done.
func (al *AgentLoop) drainBusToSteering(ctx context.Context, activeScope, activeAgentID string) {
+ blocking := true
for {
- msg, ok := al.bus.ConsumeInbound(ctx)
- if !ok {
- return
+ var msg bus.InboundMessage
+
+ if blocking {
+ // Block waiting for the first available message or ctx cancellation.
+ select {
+ case <-ctx.Done():
+ return
+ case m, ok := <-al.bus.InboundChan():
+ if !ok {
+ return
+ }
+ msg = m
+ }
+ } else {
+ // Non-blocking: drain any remaining queued messages, return when empty.
+ select {
+ case m, ok := <-al.bus.InboundChan():
+ if !ok {
+ return
+ }
+ msg = m
+ default:
+ return
+ }
}
+ blocking = false
msgScope, _, scopeOK := al.resolveSteeringTarget(msg)
if !scopeOK || msgScope != activeScope {
@@ -420,7 +552,7 @@ func (al *AgentLoop) drainBusToSteering(ctx context.Context, activeScope, active
"sender_id": msg.SenderID,
})
}
- return
+ continue
}
// Transcribe audio if needed before steering, so the agent sees text.
@@ -603,11 +735,12 @@ func (al *AgentLoop) emitEvent(kind EventKind, meta EventMeta, payload any) {
Payload: payload,
}
- al.logEvent(evt)
-
if al == nil || al.eventBus == nil {
return
}
+
+ al.logEvent(evt)
+
al.eventBus.Emit(evt)
}
@@ -814,7 +947,7 @@ func (al *AgentLoop) ReloadProviderAndConfig(
}
// Ensure shared tools are re-registered on the new registry
- registerSharedTools(cfg, al.bus, registry, provider)
+ registerSharedTools(al, cfg, al.bus, registry, provider)
// Atomically swap the config and registry under write lock
// This ensures readers see a consistent pair
@@ -891,6 +1024,11 @@ func (al *AgentLoop) SetTranscriber(t voice.Transcriber) {
al.transcriber = t
}
+// SetReloadFunc sets the callback function for triggering config reload.
+func (al *AgentLoop) SetReloadFunc(fn func() error) {
+ al.reloadFunc = fn
+}
+
var audioAnnotationRe = regexp.MustCompile(`\[(voice|audio)(?::[^\]]*)?\]`)
// transcribeAudioInMessage resolves audio media refs, transcribes them, and
@@ -1154,14 +1292,16 @@ func (al *AgentLoop) processMessage(ctx context.Context, msg bus.InboundMessage)
})
opts := processOptions{
- SessionKey: sessionKey,
- Channel: msg.Channel,
- ChatID: msg.ChatID,
- UserMessage: msg.Content,
- Media: msg.Media,
- DefaultResponse: defaultResponse,
- EnableSummary: true,
- SendResponse: false,
+ SessionKey: sessionKey,
+ Channel: msg.Channel,
+ ChatID: msg.ChatID,
+ SenderID: msg.SenderID,
+ SenderDisplayName: msg.Sender.DisplayName,
+ UserMessage: msg.Content,
+ Media: msg.Media,
+ DefaultResponse: defaultResponse,
+ EnableSummary: true,
+ SendResponse: false,
}
// context-dependent commands check their own Runtime fields and report
@@ -1216,9 +1356,16 @@ func (al *AgentLoop) resolveSteeringTarget(msg bus.InboundMessage) (string, stri
}
func (al *AgentLoop) requeueInboundMessage(msg bus.InboundMessage) error {
+ if al.bus == nil {
+ return nil
+ }
pubCtx, cancel := context.WithTimeout(context.Background(), time.Second)
defer cancel()
- return al.bus.PublishInbound(pubCtx, msg)
+ return al.bus.PublishOutbound(pubCtx, bus.OutboundMessage{
+ Channel: msg.Channel,
+ ChatID: msg.ChatID,
+ Content: msg.Content,
+ })
}
func (al *AgentLoop) processSystemMessage(
@@ -1293,6 +1440,7 @@ func (al *AgentLoop) runAgentLoop(
agent *AgentInstance,
opts processOptions,
) (string, error) {
+ // Record last channel for heartbeat notifications (skip internal channels and cli)
if opts.Channel != "" && opts.ChatID != "" && !constants.IsInternalChannel(opts.Channel) {
channelKey := fmt.Sprintf("%s:%s", opts.Channel, opts.ChatID)
if err := al.RecordLastChannel(channelKey); err != nil {
@@ -1406,6 +1554,10 @@ func (al *AgentLoop) runTurn(ctx context.Context, ts *turnState) (turnResult, er
defer turnCancel()
ts.setTurnCancel(turnCancel)
+ // Inject turnState and AgentLoop into context so tools (e.g. spawn) can retrieve them.
+ turnCtx = withTurnState(turnCtx, ts)
+ turnCtx = WithAgentLoop(turnCtx, al)
+
al.registerActiveTurn(ts)
defer al.clearActiveTurn(ts)
@@ -1449,6 +1601,8 @@ func (al *AgentLoop) runTurn(ctx context.Context, ts *turnState) (turnResult, er
ts.media,
ts.channel,
ts.chatID,
+ ts.opts.SenderID,
+ ts.opts.SenderDisplayName,
)
cfg := al.GetConfig()
@@ -1477,11 +1631,13 @@ func (al *AgentLoop) runTurn(ctx context.Context, ts *turnState) (turnResult, er
messages = ts.agent.ContextBuilder.BuildMessages(
newHistory, newSummary, ts.userMessage,
ts.media, ts.channel, ts.chatID,
+ ts.opts.SenderID, ts.opts.SenderDisplayName,
)
messages = resolveMediaRefs(messages, al.mediaStore, maxMediaSize)
}
}
+ // Save user message to session (from Incoming)
if !ts.opts.NoHistory && (strings.TrimSpace(ts.userMessage) != "" || len(ts.media) > 0) {
rootMsg := providers.Message{
Role: "user",
@@ -1524,6 +1680,37 @@ turnLoop:
}
}
+ // Check if parent turn has ended (SubTurn support from HEAD)
+ if ts.parentTurnState != nil && ts.IsParentEnded() {
+ if !ts.critical {
+ logger.InfoCF("agent", "Parent turn ended, non-critical SubTurn exiting gracefully", map[string]any{
+ "agent_id": ts.agentID,
+ "iteration": iteration,
+ "turn_id": ts.turnID,
+ })
+ break
+ }
+ logger.InfoCF("agent", "Parent turn ended, critical SubTurn continues running", map[string]any{
+ "agent_id": ts.agentID,
+ "iteration": iteration,
+ "turn_id": ts.turnID,
+ })
+ }
+
+ // Poll for pending SubTurn results (from HEAD)
+ if ts.pendingResults != nil {
+ select {
+ case result, ok := <-ts.pendingResults:
+ if ok && result != nil && result.ForLLM != "" {
+ msg := providers.Message{Role: "user", Content: fmt.Sprintf("[SubTurn Result] %s", result.ForLLM)}
+ pendingMessages = append(pendingMessages, msg)
+ }
+ default:
+ // No results available
+ }
+ }
+
+ // Inject pending steering messages
if len(pendingMessages) > 0 {
resolvedPending := resolveMediaRefs(pendingMessages, al.mediaStore, maxMediaSize)
totalContentLen := 0
@@ -1562,6 +1749,30 @@ turnLoop:
gracefulTerminal, _ := ts.gracefulInterruptRequested()
providerToolDefs := ts.agent.Tools.ToProviderDefs()
+
+ // Native web search support (from HEAD)
+ _, hasWebSearch := ts.agent.Tools.Get("web_search")
+ useNativeSearch := al.cfg.Tools.Web.PreferNative &&
+ hasWebSearch &&
+ func() bool {
+ // Check if provider supports native search
+ if ns, ok := ts.agent.Provider.(interface{ SupportsNativeSearch() bool }); ok {
+ return ns.SupportsNativeSearch()
+ }
+ return false
+ }()
+
+ if useNativeSearch {
+ // Filter out client-side web_search tool
+ filtered := make([]providers.ToolDefinition, 0, len(providerToolDefs))
+ for _, td := range providerToolDefs {
+ if td.Function.Name != "web_search" {
+ filtered = append(filtered, td)
+ }
+ }
+ providerToolDefs = filtered
+ }
+
callMessages := messages
if gracefulTerminal {
callMessages = append(append([]providers.Message(nil), messages...), ts.interruptHintMessage())
@@ -1574,6 +1785,9 @@ turnLoop:
"temperature": ts.agent.Temperature,
"prompt_cache_key": ts.agent.ID,
}
+ if useNativeSearch {
+ llmOpts["native_search"] = true
+ }
if ts.agent.ThinkingLevel != ThinkingOff {
if tc, ok := ts.agent.Provider.(providers.ThinkingCapable); ok && tc.SupportsThinking() {
llmOpts["thinking_level"] = string(ts.agent.ThinkingLevel)
@@ -1783,6 +1997,7 @@ turnLoop:
messages = ts.agent.ContextBuilder.BuildMessages(
newHistory, newSummary, "",
nil, ts.channel, ts.chatID,
+ "", "", // Empty SenderID and SenderDisplayName for retry
)
callMessages = messages
if gracefulTerminal {
@@ -1836,6 +2051,15 @@ turnLoop:
}
}
+ // Save finishReason to turnState for SubTurn truncation detection
+ if innerTS := turnStateFromContext(ctx); innerTS != nil {
+ innerTS.SetLastFinishReason(response.FinishReason)
+ // Save usage for token budget tracking
+ if response.Usage != nil {
+ innerTS.SetLastUsage(response.Usage)
+ }
+ }
+
go al.handleReasoning(
turnCtx,
response.Reasoning,
@@ -2040,10 +2264,28 @@ turnLoop:
},
)
+ // Send tool feedback to chat channel if enabled (from HEAD)
+ if al.cfg.Agents.Defaults.IsToolFeedbackEnabled() && ts.channel != "" {
+ feedbackPreview := utils.Truncate(
+ string(argsJSON),
+ al.cfg.Agents.Defaults.GetToolFeedbackMaxArgsLength(),
+ )
+ feedbackMsg := fmt.Sprintf("\U0001f527 `%s`\n```\n%s\n```", tc.Name, feedbackPreview)
+ fbCtx, fbCancel := context.WithTimeout(turnCtx, 3*time.Second)
+ _ = al.bus.PublishOutbound(fbCtx, bus.OutboundMessage{
+ Channel: ts.channel,
+ ChatID: ts.chatID,
+ Content: feedbackMsg,
+ })
+ fbCancel()
+ }
+
toolCallID := tc.ID
toolIteration := iteration
asyncToolName := toolName
asyncCallback := func(_ context.Context, result *tools.ToolResult) {
+ // Send ForUser content directly to the user (immediate feedback),
+ // mirroring the synchronous tool execution path.
if !result.Silent && result.ForUser != "" {
outCtx, outCancel := context.WithTimeout(context.Background(), 5*time.Second)
defer outCancel()
@@ -2054,6 +2296,7 @@ turnLoop:
})
}
+ // Determine content for the agent loop (ForLLM or error).
content := result.ForLLM
if content == "" && result.Err != nil {
content = result.Err.Error()
@@ -2248,6 +2491,20 @@ turnLoop:
}
break
}
+
+ // Also poll for any SubTurn results that arrived during tool execution.
+ if ts.pendingResults != nil {
+ select {
+ case result, ok := <-ts.pendingResults:
+ if ok && result != nil && result.ForLLM != "" {
+ msg := providers.Message{Role: "user", Content: fmt.Sprintf("[SubTurn Result] %s", result.ForLLM)}
+ messages = append(messages, msg)
+ ts.agent.Sessions.AddFullMessage(ts.sessionKey, msg)
+ }
+ default:
+ // No results available
+ }
+ }
}
ts.agent.Tools.TickTTL()
@@ -2274,7 +2531,11 @@ turnLoop:
}
if finalContent == "" {
- finalContent = ts.opts.DefaultResponse
+ if ts.currentIteration() >= ts.agent.MaxIterations && ts.agent.MaxIterations > 0 {
+ finalContent = toolLimitResponse
+ } else {
+ finalContent = ts.opts.DefaultResponse
+ }
}
ts.setPhase(TurnPhaseFinalizing)
@@ -2353,7 +2614,7 @@ func (al *AgentLoop) selectCandidates(
history []providers.Message,
) (candidates []providers.FallbackCandidate, model string) {
if agent.Router == nil || len(agent.LightCandidates) == 0 {
- return agent.Candidates, agent.Model
+ return agent.Candidates, resolvedCandidateModel(agent.Candidates, agent.Model)
}
_, usedLight, score := agent.Router.SelectModel(userMsg, history, agent.Model)
@@ -2364,7 +2625,7 @@ func (al *AgentLoop) selectCandidates(
"score": score,
"threshold": agent.Router.Threshold(),
})
- return agent.Candidates, agent.Model
+ return agent.Candidates, resolvedCandidateModel(agent.Candidates, agent.Model)
}
logger.InfoCF("agent", "Model routing: light model selected",
@@ -2374,7 +2635,7 @@ func (al *AgentLoop) selectCandidates(
"score": score,
"threshold": agent.Router.Threshold(),
})
- return agent.LightCandidates, agent.Router.LightModel()
+ return agent.LightCandidates, resolvedCandidateModel(agent.LightCandidates, agent.Router.LightModel())
}
// maybeSummarize triggers summarization if the session history exceeds thresholds.
@@ -2862,6 +3123,13 @@ func (al *AgentLoop) buildCommandsRuntime(agent *AgentInstance, opts *processOpt
}
return al.channelManager.GetEnabledChannels()
},
+ GetActiveTurn: func() any {
+ info := al.GetActiveTurn()
+ if info == nil {
+ return nil
+ }
+ return info
+ },
SwitchChannel: func(value string) error {
if al.channelManager == nil {
return fmt.Errorf("channel manager not initialized")
@@ -2872,13 +3140,45 @@ func (al *AgentLoop) buildCommandsRuntime(agent *AgentInstance, opts *processOpt
return nil
},
}
+ rt.ReloadConfig = func() error {
+ if al.reloadFunc == nil {
+ return fmt.Errorf("reload not configured")
+ }
+ return al.reloadFunc()
+ }
if agent != nil {
rt.GetModelInfo = func() (string, string) {
- return agent.Model, cfg.Agents.Defaults.Provider
+ return agent.Model, resolvedCandidateProvider(agent.Candidates, cfg.Agents.Defaults.Provider)
}
rt.SwitchModel = func(value string) (string, error) {
+ value = strings.TrimSpace(value)
+ modelCfg, err := resolvedModelConfig(cfg, value, agent.Workspace)
+ if err != nil {
+ return "", err
+ }
+
+ nextProvider, _, err := providers.CreateProviderFromConfig(modelCfg)
+ if err != nil {
+ return "", fmt.Errorf("failed to initialize model %q: %w", value, err)
+ }
+
+ nextCandidates := resolveModelCandidates(cfg, cfg.Agents.Defaults.Provider, modelCfg.Model, agent.Fallbacks)
+ if len(nextCandidates) == 0 {
+ return "", fmt.Errorf("model %q did not resolve to any provider candidates", value)
+ }
+
oldModel := agent.Model
+ oldProvider := agent.Provider
agent.Model = value
+ agent.Provider = nextProvider
+ agent.Candidates = nextCandidates
+ agent.ThinkingLevel = parseThinkingLevel(modelCfg.ThinkingLevel)
+
+ if oldProvider != nil && oldProvider != nextProvider {
+ if stateful, ok := oldProvider.(providers.StatefulProvider); ok {
+ stateful.Close()
+ }
+ }
return oldModel, nil
}
@@ -2939,6 +3239,28 @@ func extractParentPeer(msg bus.InboundMessage) *routing.RoutePeer {
return &routing.RoutePeer{Kind: parentKind, ID: parentID}
}
+// isNativeSearchProvider reports whether the given LLM provider implements
+// NativeSearchCapable and returns true for SupportsNativeSearch.
+func isNativeSearchProvider(p providers.LLMProvider) bool {
+ if ns, ok := p.(providers.NativeSearchCapable); ok {
+ return ns.SupportsNativeSearch()
+ }
+ return false
+}
+
+// filterClientWebSearch returns a copy of tools with the client-side
+// web_search tool removed. Used when native provider search is preferred.
+func filterClientWebSearch(tools []providers.ToolDefinition) []providers.ToolDefinition {
+ result := make([]providers.ToolDefinition, 0, len(tools))
+ for _, t := range tools {
+ if strings.EqualFold(t.Function.Name, "web_search") {
+ continue
+ }
+ result = append(result, t)
+ }
+ return result
+}
+
// Helper to extract provider from registry for cleanup
func extractProvider(registry *AgentRegistry) (providers.LLMProvider, bool) {
if registry == nil {
diff --git a/pkg/agent/loop_mcp.go b/pkg/agent/loop_mcp.go
index 962789a06..97debbc33 100644
--- a/pkg/agent/loop_mcp.go
+++ b/pkg/agent/loop_mcp.go
@@ -11,6 +11,7 @@ import (
"fmt"
"sync"
+ "github.com/sipeed/picoclaw/pkg/config"
"github.com/sipeed/picoclaw/pkg/logger"
"github.com/sipeed/picoclaw/pkg/mcp"
"github.com/sipeed/picoclaw/pkg/tools"
@@ -111,6 +112,12 @@ func (al *AgentLoop) ensureMCPInitialized(ctx context.Context) error {
for serverName, conn := range servers {
uniqueTools += len(conn.Tools)
+
+ // Determine whether this server's tools should be deferred (hidden).
+ // Per-server "deferred" field takes precedence over the global Discovery.Enabled.
+ serverCfg := al.cfg.Tools.MCP.Servers[serverName]
+ registerAsHidden := serverIsDeferred(al.cfg.Tools.MCP.Discovery.Enabled, serverCfg)
+
for _, tool := range conn.Tools {
for _, agentID := range agentIDs {
agent, ok := al.registry.GetAgent(agentID)
@@ -120,7 +127,7 @@ func (al *AgentLoop) ensureMCPInitialized(ctx context.Context) error {
mcpTool := tools.NewMCPTool(mcpManager, serverName, tool)
- if al.cfg.Tools.MCP.Discovery.Enabled {
+ if registerAsHidden {
agent.Tools.RegisterHidden(mcpTool)
} else {
agent.Tools.Register(mcpTool)
@@ -133,6 +140,7 @@ func (al *AgentLoop) ensureMCPInitialized(ctx context.Context) error {
"server": serverName,
"tool": tool.Name,
"name": mcpTool.Name(),
+ "deferred": registerAsHidden,
})
}
}
@@ -198,3 +206,18 @@ func (al *AgentLoop) ensureMCPInitialized(ctx context.Context) error {
return al.mcp.getInitErr()
}
+
+// serverIsDeferred reports whether an MCP server's tools should be registered
+// as hidden (deferred/discovery mode).
+//
+// The per-server Deferred field takes precedence over the global discoveryEnabled
+// default. When Deferred is nil, discoveryEnabled is used as the fallback.
+func serverIsDeferred(discoveryEnabled bool, serverCfg config.MCPServerConfig) bool {
+ if !discoveryEnabled {
+ return false
+ }
+ if serverCfg.Deferred != nil {
+ return *serverCfg.Deferred
+ }
+ return true
+}
diff --git a/pkg/agent/loop_mcp_test.go b/pkg/agent/loop_mcp_test.go
new file mode 100644
index 000000000..35c3e49c8
--- /dev/null
+++ b/pkg/agent/loop_mcp_test.go
@@ -0,0 +1,75 @@
+// PicoClaw - Ultra-lightweight personal AI agent
+// Inspired by and based on nanobot: https://github.com/HKUDS/nanobot
+// License: MIT
+//
+// Copyright (c) 2026 PicoClaw contributors
+
+package agent
+
+import (
+ "testing"
+
+ "github.com/sipeed/picoclaw/pkg/config"
+)
+
+func boolPtr(b bool) *bool { return &b }
+
+func TestServerIsDeferred(t *testing.T) {
+ tests := []struct {
+ name string
+ discoveryEnabled bool
+ serverDeferred *bool
+ want bool
+ }{
+ // --- global false always wins: per-server deferred is ignored ---
+ {
+ name: "global false: per-server deferred=true is ignored",
+ discoveryEnabled: false,
+ serverDeferred: boolPtr(true),
+ want: false,
+ },
+ {
+ name: "global false: per-server deferred=false stays false",
+ discoveryEnabled: false,
+ serverDeferred: boolPtr(false),
+ want: false,
+ },
+ // --- global true: per-server override applies ---
+ {
+ name: "global true: per-server deferred=false opts out",
+ discoveryEnabled: true,
+ serverDeferred: boolPtr(false),
+ want: false,
+ },
+ {
+ name: "global true: per-server deferred=true stays true",
+ discoveryEnabled: true,
+ serverDeferred: boolPtr(true),
+ want: true,
+ },
+ // --- no per-server override: fall back to global ---
+ {
+ name: "no per-server field, global discovery enabled",
+ discoveryEnabled: true,
+ serverDeferred: nil,
+ want: true,
+ },
+ {
+ name: "no per-server field, global discovery disabled",
+ discoveryEnabled: false,
+ serverDeferred: nil,
+ want: false,
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ serverCfg := config.MCPServerConfig{Deferred: tt.serverDeferred}
+ got := serverIsDeferred(tt.discoveryEnabled, serverCfg)
+ if got != tt.want {
+ t.Errorf("serverIsDeferred(discoveryEnabled=%v, deferred=%v) = %v, want %v",
+ tt.discoveryEnabled, tt.serverDeferred, got, tt.want)
+ }
+ })
+ }
+}
diff --git a/pkg/agent/loop_test.go b/pkg/agent/loop_test.go
index b65c0e21c..71f2d15e4 100644
--- a/pkg/agent/loop_test.go
+++ b/pkg/agent/loop_test.go
@@ -2,7 +2,10 @@ package agent
import (
"context"
+ "encoding/json"
"fmt"
+ "net/http"
+ "net/http/httptest"
"os"
"path/filepath"
"slices"
@@ -30,6 +33,28 @@ func (f *fakeChannel) IsAllowed(string) bool {
func (f *fakeChannel) IsAllowedSender(sender bus.SenderInfo) bool { return true }
func (f *fakeChannel) ReasoningChannelID() string { return f.id }
+type recordingProvider struct {
+ lastMessages []providers.Message
+}
+
+func (r *recordingProvider) Chat(
+ ctx context.Context,
+ messages []providers.Message,
+ tools []providers.ToolDefinition,
+ model string,
+ opts map[string]any,
+) (*providers.LLMResponse, error) {
+ r.lastMessages = append([]providers.Message(nil), messages...)
+ return &providers.LLMResponse{
+ Content: "Mock response",
+ ToolCalls: []providers.ToolCall{},
+ }, nil
+}
+
+func (r *recordingProvider) GetDefaultModel() string {
+ return "mock-model"
+}
+
func newTestAgentLoop(
t *testing.T,
) (al *AgentLoop, cfg *config.Config, msgBus *bus.MessageBus, provider *mockProvider, cleanup func()) {
@@ -54,6 +79,59 @@ func newTestAgentLoop(
return al, cfg, msgBus, provider, func() { os.RemoveAll(tmpDir) }
}
+func TestProcessMessage_IncludesCurrentSenderInDynamicContext(t *testing.T) {
+ tmpDir, err := os.MkdirTemp("", "agent-test-*")
+ if err != nil {
+ t.Fatalf("Failed to create temp dir: %v", err)
+ }
+ defer os.RemoveAll(tmpDir)
+
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Workspace: tmpDir,
+ Model: "test-model",
+ MaxTokens: 4096,
+ MaxToolIterations: 10,
+ },
+ },
+ }
+
+ msgBus := bus.NewMessageBus()
+ provider := &recordingProvider{}
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ response, err := al.processMessage(context.Background(), bus.InboundMessage{
+ Channel: "discord",
+ SenderID: "discord:123",
+ Sender: bus.SenderInfo{
+ DisplayName: "Alice",
+ },
+ ChatID: "group-1",
+ Content: "hello",
+ })
+ if err != nil {
+ t.Fatalf("processMessage() error = %v", err)
+ }
+ if response != "Mock response" {
+ t.Fatalf("processMessage() response = %q, want %q", response, "Mock response")
+ }
+ if len(provider.lastMessages) == 0 {
+ t.Fatal("provider did not receive any messages")
+ }
+
+ systemPrompt := provider.lastMessages[0].Content
+ wantSender := "## Current Sender\nCurrent sender: Alice (ID: discord:123)"
+ if !strings.Contains(systemPrompt, wantSender) {
+ t.Fatalf("system prompt missing sender context %q:\n%s", wantSender, systemPrompt)
+ }
+
+ lastMessage := provider.lastMessages[len(provider.lastMessages)-1]
+ if lastMessage.Role != "user" || lastMessage.Content != "hello" {
+ t.Fatalf("last provider message = %+v, want unchanged user message", lastMessage)
+ }
+}
+
func TestRecordLastChannel(t *testing.T) {
al, cfg, msgBus, provider, cleanup := newTestAgentLoop(t)
defer cleanup()
@@ -342,6 +420,29 @@ func (m *countingMockProvider) GetDefaultModel() string {
return "counting-mock-model"
}
+type toolLimitOnlyProvider struct{}
+
+func (m *toolLimitOnlyProvider) Chat(
+ ctx context.Context,
+ messages []providers.Message,
+ tools []providers.ToolDefinition,
+ model string,
+ opts map[string]any,
+) (*providers.LLMResponse, error) {
+ return &providers.LLMResponse{
+ ToolCalls: []providers.ToolCall{{
+ ID: "call_tool_limit_test",
+ Type: "function",
+ Name: "tool_limit_test_tool",
+ Arguments: map[string]any{"value": "x"},
+ }},
+ }, nil
+}
+
+func (m *toolLimitOnlyProvider) GetDefaultModel() string {
+ return "tool-limit-only-model"
+}
+
// mockCustomTool is a simple mock tool for registration testing
type mockCustomTool struct{}
@@ -364,11 +465,74 @@ func (m *mockCustomTool) Execute(ctx context.Context, args map[string]any) *tool
return tools.SilentResult("Custom tool executed")
}
+type toolLimitTestTool struct{}
+
+func (m *toolLimitTestTool) Name() string {
+ return "tool_limit_test_tool"
+}
+
+func (m *toolLimitTestTool) Description() string {
+ return "Tool used to exhaust the iteration budget in tests"
+}
+
+func (m *toolLimitTestTool) Parameters() map[string]any {
+ return map[string]any{
+ "type": "object",
+ "properties": map[string]any{
+ "value": map[string]any{"type": "string"},
+ },
+ }
+}
+
+func (m *toolLimitTestTool) Execute(ctx context.Context, args map[string]any) *tools.ToolResult {
+ return tools.SilentResult("tool limit test result")
+}
+
// testHelper executes a message and returns the response
type testHelper struct {
al *AgentLoop
}
+func newChatCompletionTestServer(
+ t *testing.T,
+ label string,
+ response string,
+ calls *int,
+ model *string,
+) *httptest.Server {
+ t.Helper()
+
+ return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if r.URL.Path != "/chat/completions" {
+ t.Fatalf("%s server path = %q, want /chat/completions", label, r.URL.Path)
+ }
+ *calls = *calls + 1
+ defer r.Body.Close()
+
+ var req struct {
+ Model string `json:"model"`
+ }
+ decodeErr := json.NewDecoder(r.Body).Decode(&req)
+ if decodeErr != nil {
+ t.Fatalf("decode %s request: %v", label, decodeErr)
+ }
+ *model = req.Model
+
+ w.Header().Set("Content-Type", "application/json")
+ encodeErr := json.NewEncoder(w).Encode(map[string]any{
+ "choices": []map[string]any{
+ {
+ "message": map[string]any{"content": response},
+ "finish_reason": "stop",
+ },
+ },
+ })
+ if encodeErr != nil {
+ t.Fatalf("encode %s response: %v", label, encodeErr)
+ }
+ }))
+}
+
func (h testHelper) executeAndGetResponse(tb testing.TB, ctx context.Context, msg bus.InboundMessage) string {
// Use a short timeout to avoid hanging
timeoutCtx, cancel := context.WithTimeout(ctx, responseTimeout)
@@ -530,11 +694,25 @@ func TestProcessMessage_SwitchModelShowModelConsistency(t *testing.T) {
Defaults: config.AgentDefaults{
Workspace: tmpDir,
Provider: "openai",
- Model: "before-switch",
+ Model: "local",
MaxTokens: 4096,
MaxToolIterations: 10,
},
},
+ ModelList: []config.ModelConfig{
+ {
+ ModelName: "local",
+ Model: "openai/local-model",
+ APIKey: "test-key",
+ APIBase: "https://local.example.invalid/v1",
+ },
+ {
+ ModelName: "deepseek",
+ Model: "openrouter/deepseek/deepseek-v3.2",
+ APIKey: "test-key",
+ APIBase: "https://openrouter.ai/api/v1",
+ },
+ },
}
msgBus := bus.NewMessageBus()
@@ -546,13 +724,13 @@ func TestProcessMessage_SwitchModelShowModelConsistency(t *testing.T) {
Channel: "telegram",
SenderID: "user1",
ChatID: "chat1",
- Content: "/switch model to after-switch",
+ Content: "/switch model to deepseek",
Peer: bus.Peer{
Kind: "direct",
ID: "user1",
},
})
- if !strings.Contains(switchResp, "Switched model from before-switch to after-switch") {
+ if !strings.Contains(switchResp, "Switched model from local to deepseek") {
t.Fatalf("unexpected /switch reply: %q", switchResp)
}
@@ -566,7 +744,7 @@ func TestProcessMessage_SwitchModelShowModelConsistency(t *testing.T) {
ID: "user1",
},
})
- if !strings.Contains(showResp, "Current Model: after-switch (Provider: openai)") {
+ if !strings.Contains(showResp, "Current Model: deepseek (Provider: openrouter)") {
t.Fatalf("unexpected /show model reply after switch: %q", showResp)
}
@@ -575,6 +753,187 @@ func TestProcessMessage_SwitchModelShowModelConsistency(t *testing.T) {
}
}
+func TestProcessMessage_SwitchModelRejectsUnknownAlias(t *testing.T) {
+ tmpDir, err := os.MkdirTemp("", "agent-test-*")
+ if err != nil {
+ t.Fatalf("Failed to create temp dir: %v", err)
+ }
+ defer os.RemoveAll(tmpDir)
+
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Workspace: tmpDir,
+ Provider: "openai",
+ Model: "local",
+ MaxTokens: 4096,
+ MaxToolIterations: 10,
+ },
+ },
+ ModelList: []config.ModelConfig{
+ {
+ ModelName: "local",
+ Model: "openai/local-model",
+ APIKey: "test-key",
+ APIBase: "https://local.example.invalid/v1",
+ },
+ },
+ }
+
+ msgBus := bus.NewMessageBus()
+ provider := &countingMockProvider{response: "LLM reply"}
+ al := NewAgentLoop(cfg, msgBus, provider)
+ helper := testHelper{al: al}
+
+ switchResp := helper.executeAndGetResponse(t, context.Background(), bus.InboundMessage{
+ Channel: "telegram",
+ SenderID: "user1",
+ ChatID: "chat1",
+ Content: "/switch model to missing",
+ Peer: bus.Peer{
+ Kind: "direct",
+ ID: "user1",
+ },
+ })
+ if switchResp != `model "missing" not found in model_list or providers` {
+ t.Fatalf("unexpected /switch error reply: %q", switchResp)
+ }
+
+ showResp := helper.executeAndGetResponse(t, context.Background(), bus.InboundMessage{
+ Channel: "telegram",
+ SenderID: "user1",
+ ChatID: "chat1",
+ Content: "/show model",
+ Peer: bus.Peer{
+ Kind: "direct",
+ ID: "user1",
+ },
+ })
+ if !strings.Contains(showResp, "Current Model: local (Provider: openai)") {
+ t.Fatalf("unexpected /show model reply after rejected switch: %q", showResp)
+ }
+
+ if provider.calls != 0 {
+ t.Fatalf("LLM should not be called for rejected /switch and /show, calls=%d", provider.calls)
+ }
+}
+
+func TestProcessMessage_SwitchModelRoutesSubsequentRequestsToSelectedProvider(t *testing.T) {
+ tmpDir, err := os.MkdirTemp("", "agent-test-*")
+ if err != nil {
+ t.Fatalf("Failed to create temp dir: %v", err)
+ }
+ defer os.RemoveAll(tmpDir)
+
+ localCalls := 0
+ localModel := ""
+ localServer := newChatCompletionTestServer(t, "local", "local reply", &localCalls, &localModel)
+ defer localServer.Close()
+
+ remoteCalls := 0
+ remoteModel := ""
+ remoteServer := newChatCompletionTestServer(t, "remote", "remote reply", &remoteCalls, &remoteModel)
+ defer remoteServer.Close()
+
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Workspace: tmpDir,
+ Provider: "openai",
+ Model: "local",
+ MaxTokens: 4096,
+ MaxToolIterations: 10,
+ },
+ },
+ ModelList: []config.ModelConfig{
+ {
+ ModelName: "local",
+ Model: "openai/Qwen3.5-35B-A3B",
+ APIKey: "local-key",
+ APIBase: localServer.URL,
+ },
+ {
+ ModelName: "deepseek",
+ Model: "openrouter/deepseek/deepseek-v3.2",
+ APIKey: "remote-key",
+ APIBase: remoteServer.URL,
+ },
+ },
+ }
+
+ msgBus := bus.NewMessageBus()
+ provider, _, err := providers.CreateProvider(cfg)
+ if err != nil {
+ t.Fatalf("CreateProvider() error = %v", err)
+ }
+ al := NewAgentLoop(cfg, msgBus, provider)
+ helper := testHelper{al: al}
+
+ firstResp := helper.executeAndGetResponse(t, context.Background(), bus.InboundMessage{
+ Channel: "telegram",
+ SenderID: "user1",
+ ChatID: "chat1",
+ Content: "hello before switch",
+ Peer: bus.Peer{
+ Kind: "direct",
+ ID: "user1",
+ },
+ })
+ if firstResp != "local reply" {
+ t.Fatalf("unexpected response before switch: %q", firstResp)
+ }
+ if localCalls != 1 {
+ t.Fatalf("local calls before switch = %d, want 1", localCalls)
+ }
+ if remoteCalls != 0 {
+ t.Fatalf("remote calls before switch = %d, want 0", remoteCalls)
+ }
+ if localModel != "Qwen3.5-35B-A3B" {
+ t.Fatalf("local model before switch = %q, want %q", localModel, "Qwen3.5-35B-A3B")
+ }
+
+ switchResp := helper.executeAndGetResponse(t, context.Background(), bus.InboundMessage{
+ Channel: "telegram",
+ SenderID: "user1",
+ ChatID: "chat1",
+ Content: "/switch model to deepseek",
+ Peer: bus.Peer{
+ Kind: "direct",
+ ID: "user1",
+ },
+ })
+ if !strings.Contains(switchResp, "Switched model from local to deepseek") {
+ t.Fatalf("unexpected /switch reply: %q", switchResp)
+ }
+
+ secondResp := helper.executeAndGetResponse(t, context.Background(), bus.InboundMessage{
+ Channel: "telegram",
+ SenderID: "user1",
+ ChatID: "chat1",
+ Content: "hello after switch",
+ Peer: bus.Peer{
+ Kind: "direct",
+ ID: "user1",
+ },
+ })
+ if secondResp != "remote reply" {
+ t.Fatalf("unexpected response after switch: %q", secondResp)
+ }
+ if localCalls != 1 {
+ t.Fatalf("local calls after switch = %d, want 1", localCalls)
+ }
+ if remoteCalls != 1 {
+ t.Fatalf("remote calls after switch = %d, want 1", remoteCalls)
+ }
+ if remoteModel != "deepseek-v3.2" {
+ t.Fatalf(
+ "remote model after switch = %q, want %q",
+ remoteModel,
+ "deepseek-v3.2",
+ )
+ }
+}
+
// TestToolResult_SilentToolDoesNotSendUserMessage verifies silent tools don't trigger outbound
func TestToolResult_SilentToolDoesNotSendUserMessage(t *testing.T) {
tmpDir, err := os.MkdirTemp("", "agent-test-*")
@@ -769,6 +1128,89 @@ func TestAgentLoop_ContextExhaustionRetry(t *testing.T) {
}
}
+func TestAgentLoop_EmptyModelResponseUsesAccurateFallback(t *testing.T) {
+ tmpDir, err := os.MkdirTemp("", "agent-test-*")
+ if err != nil {
+ t.Fatalf("Failed to create temp dir: %v", err)
+ }
+ defer os.RemoveAll(tmpDir)
+
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Workspace: tmpDir,
+ Model: "test-model",
+ MaxTokens: 4096,
+ MaxToolIterations: 3,
+ },
+ },
+ }
+
+ msgBus := bus.NewMessageBus()
+ provider := &simpleMockProvider{response: ""}
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ response, err := al.ProcessDirectWithChannel(context.Background(), "hello", "empty-response", "test", "chat1")
+ if err != nil {
+ t.Fatalf("ProcessDirectWithChannel failed: %v", err)
+ }
+ if response != defaultResponse {
+ t.Fatalf("response = %q, want %q", response, defaultResponse)
+ }
+}
+
+func TestAgentLoop_ToolLimitUsesDedicatedFallback(t *testing.T) {
+ tmpDir, err := os.MkdirTemp("", "agent-test-*")
+ if err != nil {
+ t.Fatalf("Failed to create temp dir: %v", err)
+ }
+ defer os.RemoveAll(tmpDir)
+
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Workspace: tmpDir,
+ Model: "test-model",
+ MaxTokens: 4096,
+ MaxToolIterations: 1,
+ },
+ },
+ }
+
+ msgBus := bus.NewMessageBus()
+ provider := &toolLimitOnlyProvider{}
+ al := NewAgentLoop(cfg, msgBus, provider)
+ al.RegisterTool(&toolLimitTestTool{})
+
+ response, err := al.ProcessDirectWithChannel(context.Background(), "hello", "tool-limit", "test", "chat1")
+ if err != nil {
+ t.Fatalf("ProcessDirectWithChannel failed: %v", err)
+ }
+ if response != toolLimitResponse {
+ t.Fatalf("response = %q, want %q", response, toolLimitResponse)
+ }
+
+ defaultAgent := al.registry.GetDefaultAgent()
+ if defaultAgent == nil {
+ t.Fatal("No default agent found")
+ }
+ route := al.registry.ResolveRoute(routing.RouteInput{
+ Channel: "test",
+ Peer: &routing.RoutePeer{
+ Kind: "direct",
+ ID: "cron",
+ },
+ })
+ history := defaultAgent.Sessions.GetHistory(route.SessionKey)
+ if len(history) != 4 {
+ t.Fatalf("history len = %d, want 4", len(history))
+ }
+ assertRoles(t, history, "user", "assistant", "tool", "assistant")
+ if history[3].Content != toolLimitResponse {
+ t.Fatalf("final assistant content = %q, want %q", history[3].Content, toolLimitResponse)
+ }
+}
+
// TestProcessDirectWithChannel_TriggersMCPInitialization verifies that
// ProcessDirectWithChannel triggers MCP initialization when MCP is enabled.
// Note: Manager is only initialized when at least one MCP server is configured
@@ -921,10 +1363,25 @@ func TestHandleReasoning(t *testing.T) {
al, msgBus := newLoop(t)
al.handleReasoning(context.Background(), "reasoning", "telegram", "")
- ctx, cancel := context.WithTimeout(context.Background(), 20*time.Millisecond)
+ ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
defer cancel()
- if msg, ok := msgBus.SubscribeOutbound(ctx); ok {
- t.Fatalf("expected no outbound message, got %+v", msg)
+ for {
+ select {
+ case msg, ok := <-msgBus.OutboundChan():
+ if !ok {
+ t.Fatalf("expected no outbound message, got %+v", msg)
+ }
+ if msg.Content == "reasoning" {
+ t.Fatalf("expected no message for empty chatID, got %+v", msg)
+ }
+ return
+ case <-ctx.Done():
+ t.Log("expected an outbound message, got none within timeout")
+ return
+ default:
+ // Continue to check for message
+ time.Sleep(5 * time.Millisecond) // Avoid busy loop
+ }
}
})
@@ -932,9 +1389,7 @@ func TestHandleReasoning(t *testing.T) {
al, msgBus := newLoop(t)
al.handleReasoning(context.Background(), "hello reasoning", "slack", "channel-1")
- ctx, cancel := context.WithTimeout(context.Background(), 200*time.Millisecond)
- defer cancel()
- msg, ok := msgBus.SubscribeOutbound(ctx)
+ msg, ok := <-msgBus.OutboundChan()
if !ok {
t.Fatal("expected an outbound message")
}
@@ -948,35 +1403,52 @@ func TestHandleReasoning(t *testing.T) {
reasoning := "hello telegram reasoning"
al.handleReasoning(context.Background(), reasoning, "telegram", "tg-chat")
- ctx, cancel := context.WithTimeout(context.Background(), 200*time.Millisecond)
+ ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
defer cancel()
- msg, ok := msgBus.SubscribeOutbound(ctx)
- if !ok {
- t.Fatal("expected outbound message")
- }
+ for {
+ select {
+ case <-ctx.Done():
+ t.Fatal("expected an outbound message, got none within timeout")
+ return
+ case msg, ok := <-msgBus.OutboundChan():
+ if !ok {
+ t.Fatal("expected outbound message")
+ }
- if msg.Channel != "telegram" {
- t.Fatalf("expected telegram channel message, got %+v", msg)
- }
- if msg.ChatID != "tg-chat" {
- t.Fatalf("expected chatID tg-chat, got %+v", msg)
- }
- if msg.Content != reasoning {
- t.Fatalf("content mismatch: got %q want %q", msg.Content, reasoning)
+ if msg.Channel != "telegram" {
+ t.Fatalf("expected telegram channel message, got %+v", msg)
+ }
+ if msg.ChatID != "tg-chat" {
+ t.Fatalf("expected chatID tg-chat, got %+v", msg)
+ }
+ if msg.Content != reasoning {
+ t.Fatalf("content mismatch: got %q want %q", msg.Content, reasoning)
+ }
+ return
+ }
}
})
t.Run("expired ctx", func(t *testing.T) {
al, msgBus := newLoop(t)
reasoning := "hello telegram reasoning"
- ctx, cancel := context.WithCancel(context.Background())
- cancel()
- al.handleReasoning(ctx, reasoning, "telegram", "tg-chat")
- ctx, cancel = context.WithTimeout(context.Background(), 200*time.Millisecond)
- defer cancel()
- msg, ok := msgBus.SubscribeOutbound(ctx)
- if ok {
- t.Fatalf("expected no outbound message, got %+v", msg)
+ al.handleReasoning(context.Background(), reasoning, "telegram", "tg-chat")
+
+ consumeCtx, consumeCancel := context.WithTimeout(context.Background(), 2*time.Second)
+ defer consumeCancel()
+
+ for {
+ select {
+ case msg, ok := <-msgBus.OutboundChan():
+ if !ok {
+ t.Fatalf("expected no outbound message, but received: %+v", msg)
+ }
+ t.Logf("Received unexpected outbound message: %+v", msg)
+ return
+ case <-consumeCtx.Done():
+ t.Fatalf("failed: no message received within timeout")
+ return
+ }
}
})
@@ -1016,20 +1488,23 @@ func TestHandleReasoning(t *testing.T) {
// Drain the bus and verify the reasoning message was NOT published
// (it should have been dropped due to timeout).
- drainCtx, drainCancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
- defer drainCancel()
- foundReasoning := false
+ timeer := time.After(1 * time.Second)
for {
- msg, ok := msgBus.SubscribeOutbound(drainCtx)
- if !ok {
- break
+ select {
+ case <-timeer:
+ t.Logf(
+ "no reasoning message received after draining bus for 1s, as expected,length=%d",
+ len(msgBus.OutboundChan()),
+ )
+ return
+ case msg, ok := <-msgBus.OutboundChan():
+ if !ok {
+ break
+ }
+ if msg.Content == "should timeout" {
+ t.Fatal("expected reasoning message to be dropped when bus is full, but it was published")
+ }
}
- if msg.Content == "should timeout" {
- foundReasoning = true
- }
- }
- if foundReasoning {
- t.Fatal("expected reasoning message to be dropped when bus is full, but it was published")
}
})
}
@@ -1317,3 +1792,84 @@ func TestResolveMediaRefs_MixedImageAndFile(t *testing.T) {
t.Fatalf("expected content %q, got %q", expectedContent, result[0].Content)
}
}
+
+// --- Native search helper tests ---
+
+type nativeSearchProvider struct {
+ supported bool
+}
+
+func (p *nativeSearchProvider) Chat(
+ ctx context.Context, msgs []providers.Message, tools []providers.ToolDefinition,
+ model string, opts map[string]any,
+) (*providers.LLMResponse, error) {
+ return &providers.LLMResponse{Content: "ok"}, nil
+}
+
+func (p *nativeSearchProvider) GetDefaultModel() string { return "test-model" }
+
+func (p *nativeSearchProvider) SupportsNativeSearch() bool { return p.supported }
+
+type plainProvider struct{}
+
+func (p *plainProvider) Chat(
+ ctx context.Context, msgs []providers.Message, tools []providers.ToolDefinition,
+ model string, opts map[string]any,
+) (*providers.LLMResponse, error) {
+ return &providers.LLMResponse{Content: "ok"}, nil
+}
+
+func (p *plainProvider) GetDefaultModel() string { return "test-model" }
+
+func TestIsNativeSearchProvider_Supported(t *testing.T) {
+ if !isNativeSearchProvider(&nativeSearchProvider{supported: true}) {
+ t.Fatal("expected true for provider that supports native search")
+ }
+}
+
+func TestIsNativeSearchProvider_NotSupported(t *testing.T) {
+ if isNativeSearchProvider(&nativeSearchProvider{supported: false}) {
+ t.Fatal("expected false for provider that does not support native search")
+ }
+}
+
+func TestIsNativeSearchProvider_NoInterface(t *testing.T) {
+ if isNativeSearchProvider(&plainProvider{}) {
+ t.Fatal("expected false for provider that does not implement NativeSearchCapable")
+ }
+}
+
+func TestFilterClientWebSearch_RemovesWebSearch(t *testing.T) {
+ defs := []providers.ToolDefinition{
+ {Type: "function", Function: providers.ToolFunctionDefinition{Name: "web_search"}},
+ {Type: "function", Function: providers.ToolFunctionDefinition{Name: "read_file"}},
+ {Type: "function", Function: providers.ToolFunctionDefinition{Name: "exec"}},
+ }
+ result := filterClientWebSearch(defs)
+ if len(result) != 2 {
+ t.Fatalf("len(result) = %d, want 2", len(result))
+ }
+ for _, td := range result {
+ if td.Function.Name == "web_search" {
+ t.Fatal("web_search should be filtered out")
+ }
+ }
+}
+
+func TestFilterClientWebSearch_NoWebSearch(t *testing.T) {
+ defs := []providers.ToolDefinition{
+ {Type: "function", Function: providers.ToolFunctionDefinition{Name: "read_file"}},
+ {Type: "function", Function: providers.ToolFunctionDefinition{Name: "exec"}},
+ }
+ result := filterClientWebSearch(defs)
+ if len(result) != 2 {
+ t.Fatalf("len(result) = %d, want 2", len(result))
+ }
+}
+
+func TestFilterClientWebSearch_EmptyInput(t *testing.T) {
+ result := filterClientWebSearch(nil)
+ if len(result) != 0 {
+ t.Fatalf("len(result) = %d, want 0", len(result))
+ }
+}
diff --git a/pkg/agent/model_resolution.go b/pkg/agent/model_resolution.go
new file mode 100644
index 000000000..140cff718
--- /dev/null
+++ b/pkg/agent/model_resolution.go
@@ -0,0 +1,97 @@
+package agent
+
+import (
+ "fmt"
+ "strings"
+
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/providers"
+)
+
+func buildModelListResolver(cfg *config.Config) func(raw string) (string, bool) {
+ ensureProtocol := func(model string) string {
+ model = strings.TrimSpace(model)
+ if model == "" {
+ return ""
+ }
+ if strings.Contains(model, "/") {
+ return model
+ }
+ return "openai/" + model
+ }
+
+ return func(raw string) (string, bool) {
+ raw = strings.TrimSpace(raw)
+ if raw == "" || cfg == nil {
+ return "", false
+ }
+
+ if mc, err := cfg.GetModelConfig(raw); err == nil && mc != nil && strings.TrimSpace(mc.Model) != "" {
+ return ensureProtocol(mc.Model), true
+ }
+
+ for i := range cfg.ModelList {
+ fullModel := strings.TrimSpace(cfg.ModelList[i].Model)
+ if fullModel == "" {
+ continue
+ }
+ if fullModel == raw {
+ return ensureProtocol(fullModel), true
+ }
+ _, modelID := providers.ExtractProtocol(fullModel)
+ if modelID == raw {
+ return ensureProtocol(fullModel), true
+ }
+ }
+
+ return "", false
+ }
+}
+
+func resolveModelCandidates(
+ cfg *config.Config,
+ defaultProvider string,
+ primary string,
+ fallbacks []string,
+) []providers.FallbackCandidate {
+ return providers.ResolveCandidatesWithLookup(
+ providers.ModelConfig{
+ Primary: primary,
+ Fallbacks: fallbacks,
+ },
+ defaultProvider,
+ buildModelListResolver(cfg),
+ )
+}
+
+func resolvedCandidateModel(candidates []providers.FallbackCandidate, fallback string) string {
+ if len(candidates) > 0 && strings.TrimSpace(candidates[0].Model) != "" {
+ return candidates[0].Model
+ }
+ return fallback
+}
+
+func resolvedCandidateProvider(candidates []providers.FallbackCandidate, fallback string) string {
+ if len(candidates) > 0 && strings.TrimSpace(candidates[0].Provider) != "" {
+ return candidates[0].Provider
+ }
+ return fallback
+}
+
+func resolvedModelConfig(cfg *config.Config, modelName, workspace string) (*config.ModelConfig, error) {
+ if cfg == nil {
+ return nil, fmt.Errorf("config is nil")
+ }
+
+ modelCfg, err := cfg.GetModelConfig(strings.TrimSpace(modelName))
+ if err != nil {
+ return nil, err
+ }
+
+ clone := *modelCfg
+ if clone.Workspace == "" {
+ clone.Workspace = workspace
+ }
+
+ return &clone, nil
+}
diff --git a/pkg/agent/steering.go b/pkg/agent/steering.go
index 70173f42b..ad6613e8c 100644
--- a/pkg/agent/steering.go
+++ b/pkg/agent/steering.go
@@ -9,6 +9,7 @@ import (
"github.com/sipeed/picoclaw/pkg/logger"
"github.com/sipeed/picoclaw/pkg/providers"
"github.com/sipeed/picoclaw/pkg/routing"
+ "github.com/sipeed/picoclaw/pkg/tools"
)
// SteeringMode controls how queued steering messages are dequeued.
@@ -172,7 +173,7 @@ func (sq *steeringQueue) getMode() SteeringMode {
func (al *AgentLoop) Steer(msg providers.Message) error {
scope := ""
agentID := ""
- if ts := al.getActiveTurnState(); ts != nil {
+ if ts := al.getAnyActiveTurnState(); ts != nil {
scope = ts.sessionKey
agentID = ts.agentID
}
@@ -206,7 +207,7 @@ func (al *AgentLoop) enqueueSteeringMessage(scope, agentID string, msg providers
Source: "Steer",
TracePath: "turn.interrupt.received",
}
- if ts := al.getActiveTurnState(); ts != nil {
+ if ts := al.getAnyActiveTurnState(); ts != nil {
meta = ts.eventMeta("Steer", "turn.interrupt.received")
} else {
if strings.TrimSpace(agentID) != "" {
@@ -355,7 +356,7 @@ func (al *AgentLoop) Continue(ctx context.Context, sessionKey, channel, chatID s
}
func (al *AgentLoop) InterruptGraceful(hint string) error {
- ts := al.getActiveTurnState()
+ ts := al.getAnyActiveTurnState()
if ts == nil {
return fmt.Errorf("no active turn")
}
@@ -376,7 +377,7 @@ func (al *AgentLoop) InterruptGraceful(hint string) error {
}
func (al *AgentLoop) InterruptHard() error {
- ts := al.getActiveTurnState()
+ ts := al.getAnyActiveTurnState()
if ts == nil {
return fmt.Errorf("no active turn")
}
@@ -394,3 +395,109 @@ func (al *AgentLoop) InterruptHard() error {
return nil
}
+
+// ====================== SubTurn Result Polling ======================
+
+// dequeuePendingSubTurnResults polls the SubTurn result channel for the given
+// session and returns all available results without blocking.
+// Returns nil if no active turn state exists for this session.
+func (al *AgentLoop) dequeuePendingSubTurnResults(sessionKey string) []*tools.ToolResult {
+ tsInterface, ok := al.activeTurnStates.Load(sessionKey)
+ if !ok {
+ return nil
+ }
+ ts, ok := tsInterface.(*turnState)
+ if !ok {
+ return nil
+ }
+
+ var results []*tools.ToolResult
+ for {
+ select {
+ case result, ok := <-ts.pendingResults:
+ if !ok {
+ return results
+ }
+ if result != nil {
+ results = append(results, result)
+ }
+ default:
+ return results
+ }
+ }
+}
+
+// ====================== Hard Abort ======================
+
+// HardAbort immediately cancels the running agent loop for the given session,
+// cascading the cancellation to all child SubTurns. This is a destructive operation
+// that terminates execution without waiting for graceful cleanup.
+//
+// Use this when the user explicitly requests immediate termination (e.g., "stop now", "abort").
+// For graceful interruption that allows the agent to finish the current tool and summarize,
+// use Steer() instead.
+func (al *AgentLoop) HardAbort(sessionKey string) error {
+ tsInterface, ok := al.activeTurnStates.Load(sessionKey)
+ if !ok {
+ return fmt.Errorf("no active turn state found for session %s", sessionKey)
+ }
+
+ ts, ok := tsInterface.(*turnState)
+ if !ok {
+ return fmt.Errorf("invalid turn state type for session %s", sessionKey)
+ }
+
+ logger.InfoCF("agent", "Hard abort triggered", map[string]any{
+ "session_key": sessionKey,
+ "turn_id": ts.turnID,
+ "depth": ts.depth,
+ "initial_history_length": ts.initialHistoryLength,
+ })
+
+ // IMPORTANT: Trigger cascading cancellation FIRST to stop all child SubTurns
+ // from adding more messages to the session. This prevents race conditions
+ // where rollback happens while children are still writing.
+ // Use isHardAbort=true for hard abort to immediately cancel all children.
+ ts.Finish(true)
+
+ // Roll back session history to the state before the turn started.
+ if ts.session != nil {
+ history := ts.session.GetHistory(sessionKey)
+ if ts.initialHistoryLength < len(history) {
+ ts.session.SetHistory(sessionKey, history[:ts.initialHistoryLength])
+ }
+ }
+
+ return nil
+}
+
+// ====================== Follow-Up Injection ======================
+
+// InjectFollowUp enqueues a message to be automatically processed after the current
+// turn completes. Unlike Steer(), which interrupts the current execution, InjectFollowUp
+// waits for the current turn to finish naturally before processing the message.
+//
+// This is useful for:
+// - Automated workflows that need to chain multiple turns
+// - Background tasks that should run after the main task completes
+// - Scheduled follow-up actions
+//
+// The message will be processed via Continue() when the agent becomes idle.
+func (al *AgentLoop) InjectFollowUp(msg providers.Message) error {
+ // InjectFollowUp uses the same steering queue mechanism as Steer(),
+ // but the semantic difference is in when it's called:
+ // - Steer() is called during active execution to interrupt
+ // - InjectFollowUp() is called when planning future work
+ //
+ // Both end up in the same queue and are processed by Continue()
+ // when the agent is idle.
+ return al.Steer(msg)
+}
+
+// ====================== API Aliases for Design Document Compatibility ======================
+
+// InjectSteering is an alias for Steer() to match the design document naming.
+// It injects a steering message into the currently running agent loop.
+func (al *AgentLoop) InjectSteering(msg providers.Message) error {
+ return al.Steer(msg)
+}
diff --git a/pkg/agent/steering_test.go b/pkg/agent/steering_test.go
index cf2e86904..fe4863f05 100644
--- a/pkg/agent/steering_test.go
+++ b/pkg/agent/steering_test.go
@@ -420,13 +420,14 @@ func TestDrainBusToSteering_RequeuesDifferentScopeMessage(t *testing.T) {
t.Fatalf("expected no steering messages for active scope, got %v", msgs)
}
- requeued, ok := msgBus.ConsumeInbound(context.Background())
- if !ok {
- t.Fatal("expected message to be requeued on the inbound bus")
- }
- if requeued.Channel != otherMsg.Channel || requeued.ChatID != otherMsg.ChatID ||
- requeued.SenderID != otherMsg.SenderID || requeued.Content != otherMsg.Content {
- t.Fatalf("requeued message mismatch: got %+v want %+v", requeued, otherMsg)
+ select {
+ case <-ctx.Done():
+ t.Fatalf("timeout waiting for requeued message on outbound bus")
+ case requeued := <-msgBus.OutboundChan():
+ if requeued.Channel != otherMsg.Channel || requeued.ChatID != otherMsg.ChatID ||
+ requeued.Content != otherMsg.Content {
+ t.Fatalf("requeued message mismatch: got %+v want %+v", requeued, otherMsg)
+ }
}
}
@@ -881,8 +882,10 @@ func TestAgentLoop_Run_AutoContinuesLateSteeringMessage(t *testing.T) {
subCtx, subCancel := context.WithTimeout(context.Background(), 5*time.Second)
defer subCancel()
- out1, ok := msgBus.SubscribeOutbound(subCtx)
- if !ok {
+ var out1 bus.OutboundMessage
+ select {
+ case out1 = <-msgBus.OutboundChan():
+ case <-subCtx.Done():
t.Fatal("expected outbound response")
}
if out1.Content != "continued response" {
@@ -891,8 +894,10 @@ func TestAgentLoop_Run_AutoContinuesLateSteeringMessage(t *testing.T) {
noExtraCtx, cancelNoExtra := context.WithTimeout(context.Background(), 200*time.Millisecond)
defer cancelNoExtra()
- if out2, ok := msgBus.SubscribeOutbound(noExtraCtx); ok {
+ select {
+ case out2 := <-msgBus.OutboundChan():
t.Fatalf("expected stale direct response to be suppressed, got extra outbound %q", out2.Content)
+ case <-noExtraCtx.Done():
}
cancelRun()
diff --git a/pkg/agent/subturn.go b/pkg/agent/subturn.go
new file mode 100644
index 000000000..f5ba412ab
--- /dev/null
+++ b/pkg/agent/subturn.go
@@ -0,0 +1,671 @@
+package agent
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "sync"
+ "sync/atomic"
+ "time"
+
+ "github.com/sipeed/picoclaw/pkg/logger"
+ "github.com/sipeed/picoclaw/pkg/providers"
+ "github.com/sipeed/picoclaw/pkg/tools"
+)
+
+// ====================== Config & Constants ======================
+const (
+ // Default values for SubTurn configuration (used when config is not set or is zero)
+ defaultMaxSubTurnDepth = 3
+ defaultMaxConcurrentSubTurns = 5
+ defaultConcurrencyTimeout = 30 * time.Second
+ defaultSubTurnTimeout = 5 * time.Minute
+ // maxEphemeralHistorySize limits the number of messages stored in ephemeral sessions.
+ // This prevents memory accumulation in long-running sub-turns.
+ maxEphemeralHistorySize = 50
+)
+
+var (
+ ErrDepthLimitExceeded = errors.New("sub-turn depth limit exceeded")
+ ErrInvalidSubTurnConfig = errors.New("invalid sub-turn config")
+ ErrConcurrencyTimeout = errors.New("timeout waiting for concurrency slot")
+)
+
+// getSubTurnConfig returns the effective SubTurn configuration with defaults applied.
+func (al *AgentLoop) getSubTurnConfig() subTurnRuntimeConfig {
+ cfg := al.cfg.Agents.Defaults.SubTurn
+
+ maxDepth := cfg.MaxDepth
+ if maxDepth <= 0 {
+ maxDepth = defaultMaxSubTurnDepth
+ }
+
+ maxConcurrent := cfg.MaxConcurrent
+ if maxConcurrent <= 0 {
+ maxConcurrent = defaultMaxConcurrentSubTurns
+ }
+
+ concurrencyTimeout := time.Duration(cfg.ConcurrencyTimeoutSec) * time.Second
+ if concurrencyTimeout <= 0 {
+ concurrencyTimeout = defaultConcurrencyTimeout
+ }
+
+ defaultTimeout := time.Duration(cfg.DefaultTimeoutMinutes) * time.Minute
+ if defaultTimeout <= 0 {
+ defaultTimeout = defaultSubTurnTimeout
+ }
+
+ return subTurnRuntimeConfig{
+ maxDepth: maxDepth,
+ maxConcurrent: maxConcurrent,
+ concurrencyTimeout: concurrencyTimeout,
+ defaultTimeout: defaultTimeout,
+ defaultTokenBudget: cfg.DefaultTokenBudget,
+ }
+}
+
+// subTurnRuntimeConfig holds the effective runtime configuration for SubTurn execution.
+type subTurnRuntimeConfig struct {
+ maxDepth int
+ maxConcurrent int
+ concurrencyTimeout time.Duration
+ defaultTimeout time.Duration
+ defaultTokenBudget int
+}
+
+// ====================== SubTurn Config ======================
+
+// SubTurnConfig configures the execution of a child sub-turn.
+//
+// Usage Examples:
+//
+// Synchronous sub-turn (Async=false):
+//
+// cfg := SubTurnConfig{
+// Model: "gpt-4o-mini",
+// SystemPrompt: "Analyze this code",
+// Async: false, // Result returned immediately
+// }
+// result, err := SpawnSubTurn(ctx, cfg)
+// // Use result directly here
+// processResult(result)
+//
+// Asynchronous sub-turn (Async=true):
+//
+// cfg := SubTurnConfig{
+// Model: "gpt-4o-mini",
+// SystemPrompt: "Background analysis",
+// Async: true, // Result delivered to channel
+// }
+// result, err := SpawnSubTurn(ctx, cfg)
+// // Result also available in parent's pendingResults channel
+// // Parent turn will poll and process it in a later iteration
+type SubTurnConfig struct {
+ Model string
+ Tools []tools.Tool
+ SystemPrompt string
+ MaxTokens int
+
+ // Async controls the result delivery mechanism:
+ //
+ // When Async = false (synchronous sub-turn):
+ // - The caller blocks until the sub-turn completes
+ // - The result is ONLY returned via the function return value
+ // - The result is NOT delivered to the parent's pendingResults channel
+ // - This prevents double delivery: caller gets result immediately, no need for channel
+ // - Use case: When the caller needs the result immediately to continue execution
+ // - Example: A tool that needs to process the sub-turn result before returning
+ //
+ // When Async = true (asynchronous sub-turn):
+ // - The sub-turn runs in the background (still blocks the caller, but semantically async)
+ // - The result is delivered to the parent's pendingResults channel
+ // - The result is ALSO returned via the function return value (for consistency)
+ // - The parent turn can poll pendingResults in later iterations to process results
+ // - Use case: Fire-and-forget operations, or when results are processed in batches
+ // - Example: Spawning multiple sub-turns in parallel and collecting results later
+ //
+ // IMPORTANT: The Async flag does NOT make the call non-blocking. It only controls
+ // whether the result is delivered via the channel. For true non-blocking execution,
+ // the caller must spawn the sub-turn in a separate goroutine.
+ Async bool
+
+ // Critical indicates this SubTurn's result is important and should continue
+ // running even after the parent turn finishes gracefully.
+ //
+ // When parent finishes gracefully (Finish(false)):
+ // - Critical=true: SubTurn continues running, delivers result as orphan
+ // - Critical=false: SubTurn exits gracefully without error
+ //
+ // When parent finishes with hard abort (Finish(true)):
+ // - All SubTurns are canceled regardless of Critical flag
+ Critical bool
+
+ // Timeout is the maximum duration for this SubTurn.
+ // If the SubTurn runs longer than this, it will be canceled.
+ // Default is 5 minutes (defaultSubTurnTimeout) if not specified.
+ Timeout time.Duration
+
+ // MaxContextRunes limits the context size (in runes) passed to the SubTurn.
+ // This prevents context window overflow by truncating message history before LLM calls.
+ //
+ // Values:
+ // 0 = Auto-calculate based on model's ContextWindow * 0.75 (default, recommended)
+ // -1 = No limit (disable soft truncation, rely only on hard context errors)
+ // >0 = Use specified rune limit
+ //
+ // The soft limit acts as a first line of defense before hitting the provider's
+ // hard context window limit. When exceeded, older messages are intelligently
+ // truncated while preserving system messages and recent context.
+ MaxContextRunes int
+
+ // ActualSystemPrompt is injected as the true 'system' role message for the childAgent.
+ // The legacy SystemPrompt field is actually used as the first 'user' message (task description).
+ ActualSystemPrompt string
+
+ // InitialMessages preloads the ephemeral session history before the agent loop starts.
+ // Used by evaluator-optimizer patterns to pass the full worker context across multiple iterations.
+ InitialMessages []providers.Message
+
+ // InitialTokenBudget is a shared atomic counter for tracking remaining tokens.
+ // If set, the SubTurn will inherit this budget and deduct tokens after each LLM call.
+ // If nil, the SubTurn will inherit the parent's tokenBudget (if any).
+ // Used by team tool to enforce token limits across all team members.
+ InitialTokenBudget *atomic.Int64
+
+ // Can be extended with temperature, topP, etc.
+}
+
+// ====================== Context Keys ======================
+type agentLoopKeyType struct{}
+
+var agentLoopKey = agentLoopKeyType{}
+
+// WithAgentLoop injects AgentLoop into context for tool access
+func WithAgentLoop(ctx context.Context, al *AgentLoop) context.Context {
+ return context.WithValue(ctx, agentLoopKey, al)
+}
+
+// AgentLoopFromContext retrieves AgentLoop from context
+func AgentLoopFromContext(ctx context.Context) *AgentLoop {
+ al, _ := ctx.Value(agentLoopKey).(*AgentLoop)
+ return al
+}
+
+// ====================== Helper Functions ======================
+
+func (al *AgentLoop) generateSubTurnID() string {
+ return fmt.Sprintf("subturn-%d", al.subTurnCounter.Add(1))
+}
+
+// ====================== Core Function: spawnSubTurn ======================
+
+// AgentLoopSpawner implements tools.SubTurnSpawner interface.
+// This allows tools to spawn sub-turns without circular dependency.
+type AgentLoopSpawner struct {
+ al *AgentLoop
+}
+
+// SpawnSubTurn implements tools.SubTurnSpawner interface.
+func (s *AgentLoopSpawner) SpawnSubTurn(
+ ctx context.Context,
+ cfg tools.SubTurnConfig,
+) (*tools.ToolResult, error) {
+ parentTS := turnStateFromContext(ctx)
+ if parentTS == nil {
+ return nil, errors.New(
+ "parent turnState not found in context - cannot spawn sub-turn outside of a turn",
+ )
+ }
+
+ // Convert tools.SubTurnConfig to agent.SubTurnConfig
+ agentCfg := SubTurnConfig{
+ Model: cfg.Model,
+ Tools: cfg.Tools,
+ SystemPrompt: cfg.SystemPrompt,
+ ActualSystemPrompt: cfg.ActualSystemPrompt,
+ InitialMessages: cfg.InitialMessages,
+ InitialTokenBudget: cfg.InitialTokenBudget,
+ MaxTokens: cfg.MaxTokens,
+ Async: cfg.Async,
+ Critical: cfg.Critical,
+ Timeout: cfg.Timeout,
+ MaxContextRunes: cfg.MaxContextRunes,
+ }
+
+ return spawnSubTurn(ctx, s.al, parentTS, agentCfg)
+}
+
+// NewSubTurnSpawner creates a SubTurnSpawner for the given AgentLoop.
+func NewSubTurnSpawner(al *AgentLoop) *AgentLoopSpawner {
+ return &AgentLoopSpawner{al: al}
+}
+
+// SpawnSubTurn is the exported entry point for tools to spawn sub-turns.
+// It retrieves AgentLoop and parent turnState from context and delegates to spawnSubTurn.
+func SpawnSubTurn(ctx context.Context, cfg SubTurnConfig) (*tools.ToolResult, error) {
+ al := AgentLoopFromContext(ctx)
+ if al == nil {
+ return nil, errors.New(
+ "AgentLoop not found in context - ensure context is properly initialized",
+ )
+ }
+
+ parentTS := turnStateFromContext(ctx)
+ if parentTS == nil {
+ return nil, errors.New(
+ "parent turnState not found in context - cannot spawn sub-turn outside of a turn",
+ )
+ }
+
+ return spawnSubTurn(ctx, al, parentTS, cfg)
+}
+
+func spawnSubTurn(
+ ctx context.Context,
+ al *AgentLoop,
+ parentTS *turnState,
+ cfg SubTurnConfig,
+) (result *tools.ToolResult, err error) {
+ // Get effective SubTurn configuration
+ rtCfg := al.getSubTurnConfig()
+
+ // 0. Acquire concurrency semaphore FIRST to ensure it's released even if early validation fails.
+ // Blocks if parent already has maxConcurrentSubTurns running, with a timeout to prevent indefinite blocking.
+ // Also respects context cancellation so we don't block forever if parent is aborted.
+ // NOTE: The semaphore is released immediately after runTurn completes (not in a defer) to
+ // ensure it is freed before the cleanup phase (async result delivery), which may block on
+ // a full pendingResults channel. Holding the semaphore through cleanup would allow the
+ // parent's goroutine to be blocked waiting for a semaphore slot while child turns are
+ // blocked delivering results — a deadlock.
+ var semAcquired bool
+ if parentTS.concurrencySem != nil {
+ // Create a timeout context for semaphore acquisition
+ timeoutCtx, cancel := context.WithTimeout(ctx, rtCfg.concurrencyTimeout)
+ defer cancel()
+
+ select {
+ case parentTS.concurrencySem <- struct{}{}:
+ semAcquired = true
+ defer func() {
+ if semAcquired {
+ <-parentTS.concurrencySem
+ }
+ }()
+ case <-timeoutCtx.Done():
+ // Check parent context first - if it was canceled, propagate that error
+ if ctx.Err() != nil {
+ return nil, ctx.Err()
+ }
+ // Otherwise it's our timeout
+ return nil, fmt.Errorf("%w: all %d slots occupied for %v",
+ ErrConcurrencyTimeout, rtCfg.maxConcurrent, rtCfg.concurrencyTimeout)
+ }
+ }
+
+ // 1. Depth limit check
+ if parentTS.depth >= rtCfg.maxDepth {
+ logger.WarnCF("subturn", "Depth limit exceeded", map[string]any{
+ "parent_id": parentTS.turnID,
+ "depth": parentTS.depth,
+ "max_depth": rtCfg.maxDepth,
+ })
+ return nil, ErrDepthLimitExceeded
+ }
+
+ // 2. Config validation
+ if cfg.Model == "" {
+ return nil, ErrInvalidSubTurnConfig
+ }
+
+ // 3. Determine timeout for child SubTurn
+ timeout := cfg.Timeout
+ if timeout <= 0 {
+ timeout = rtCfg.defaultTimeout
+ }
+
+ // 4. Create INDEPENDENT child context (not derived from parent ctx).
+ // This allows the child to continue running after parent finishes gracefully.
+ // The child has its own timeout for self-protection.
+ childCtx, cancel := context.WithTimeout(context.Background(), timeout)
+ defer cancel()
+
+ childID := al.generateSubTurnID()
+
+ // Get the agent instance from parent, falling back to the default agent.
+ // Wrap it in a shallow copy that uses an ephemeral (in-memory only) session store
+ // so that child turns never pollute or persist to the parent's session history.
+ baseAgent := parentTS.agent
+ if baseAgent == nil {
+ baseAgent = al.registry.GetDefaultAgent()
+ }
+ if baseAgent == nil {
+ return nil, errors.New("parent turnState has no agent instance")
+ }
+ ephemeralStore := newEphemeralSession(nil)
+ agent := *baseAgent // shallow copy
+ agent.Sessions = ephemeralStore
+ // Clone the tool registry so child turn's tool registrations
+ // don't pollute the parent's registry.
+ if baseAgent.Tools != nil {
+ agent.Tools = baseAgent.Tools.Clone()
+ }
+
+ // Create processOptions for the child turn
+ opts := processOptions{
+ SessionKey: childID,
+ Channel: parentTS.channel,
+ ChatID: parentTS.chatID,
+ SenderID: parentTS.opts.SenderID,
+ SenderDisplayName: parentTS.opts.SenderDisplayName,
+ UserMessage: cfg.SystemPrompt, // Task description becomes the first user message
+ SystemPromptOverride: cfg.ActualSystemPrompt,
+ Media: nil,
+ InitialSteeringMessages: cfg.InitialMessages,
+ DefaultResponse: "",
+ EnableSummary: false,
+ SendResponse: false,
+ NoHistory: true, // SubTurns don't use session history
+ SkipInitialSteeringPoll: true,
+ }
+
+ // Create event scope for the child turn
+ scope := al.newTurnEventScope(agent.ID, childID)
+
+ // Create child turnState using the new API
+ childTS := newTurnState(&agent, opts, scope)
+
+ // Set SubTurn-specific fields
+ childTS.cancelFunc = cancel
+ childTS.critical = cfg.Critical
+ childTS.depth = parentTS.depth + 1
+ childTS.parentTurnID = parentTS.turnID
+ childTS.parentTurnState = parentTS
+ childTS.pendingResults = make(chan *tools.ToolResult, 16)
+ childTS.concurrencySem = make(chan struct{}, rtCfg.maxConcurrent)
+ childTS.al = al // back-ref for hard abort cascade
+ childTS.session = ephemeralStore // same store as agent.Sessions
+
+ // Token budget initialization/inheritance
+ // If InitialTokenBudget is explicitly provided (e.g., by team tool), use it.
+ // Otherwise, inherit from parent's tokenBudget (for nested SubTurns).
+ if cfg.InitialTokenBudget != nil {
+ childTS.tokenBudget = cfg.InitialTokenBudget
+ } else if parentTS.tokenBudget != nil {
+ childTS.tokenBudget = parentTS.tokenBudget
+ } else if rtCfg.defaultTokenBudget > 0 {
+ // Apply default token budget from config if no budget is set
+ budget := &atomic.Int64{}
+ budget.Store(int64(rtCfg.defaultTokenBudget))
+ childTS.tokenBudget = budget
+ }
+
+ // IMPORTANT: Put childTS into childCtx so that code inside runTurn can retrieve it
+ childCtx = withTurnState(childCtx, childTS)
+ childCtx = WithAgentLoop(childCtx, al) // Propagate AgentLoop to child turn
+
+ childTS.ctx = childCtx
+
+ // Register child turn state so GetAllActiveTurns/Subagents can find it
+ al.activeTurnStates.Store(childID, childTS)
+ defer al.activeTurnStates.Delete(childID)
+
+ // 5. Establish parent-child relationship (thread-safe)
+ parentTS.mu.Lock()
+ parentTS.childTurnIDs = append(parentTS.childTurnIDs, childID)
+ parentTS.mu.Unlock()
+
+ // 6. Emit Spawn event
+ al.emitEvent(EventKindSubTurnSpawn,
+ childTS.eventMeta("spawnSubTurn", "subturn.spawn"),
+ SubTurnSpawnPayload{
+ AgentID: childTS.agentID,
+ Label: childID,
+ ParentTurnID: parentTS.turnID,
+ },
+ )
+
+ // 7. Defer cleanup: deliver result (for async), emit End event, and recover from panics
+ defer func() {
+ if r := recover(); r != nil {
+ err = fmt.Errorf("subturn panicked: %v", r)
+ result = nil
+ logger.ErrorCF("subturn", "SubTurn panicked", map[string]any{
+ "child_id": childID,
+ "parent_id": parentTS.turnID,
+ "panic": r,
+ })
+ }
+
+ // Result Delivery Strategy (Async vs Sync)
+ if cfg.Async {
+ deliverSubTurnResult(al, parentTS, childID, result)
+ }
+
+ status := "completed"
+ if err != nil {
+ status = "error"
+ }
+ al.emitEvent(EventKindSubTurnEnd,
+ childTS.eventMeta("spawnSubTurn", "subturn.end"),
+ SubTurnEndPayload{
+ AgentID: childTS.agentID,
+ Status: status,
+ },
+ )
+ }()
+
+ // 8. Execute sub-turn via the real agent loop.
+ turnRes, turnErr := al.runTurn(childCtx, childTS)
+
+ // Release the concurrency semaphore immediately after runTurn completes,
+ // before the cleanup defer runs. This prevents a deadlock where:
+ // - All semaphore slots are held by sub-turns in their cleanup phase
+ // - Cleanup blocks on a full pendingResults channel
+ // - The parent goroutine is blocked waiting for a semaphore slot
+ // - The parent cannot consume pendingResults because it is blocked on the semaphore
+ if semAcquired {
+ <-parentTS.concurrencySem
+ semAcquired = false // prevent the defer from double-releasing
+ }
+
+ // Convert turnResult to tools.ToolResult
+ if turnErr != nil {
+ err = turnErr
+ result = &tools.ToolResult{
+ Err: turnErr,
+ ForLLM: fmt.Sprintf("SubTurn failed: %v", turnErr),
+ }
+ } else {
+ result = &tools.ToolResult{
+ ForLLM: turnRes.finalContent,
+ ForUser: turnRes.finalContent,
+ }
+ }
+
+ return result, err
+}
+
+// ====================== Result Delivery ======================
+
+// deliverSubTurnResult delivers a sub-turn result to the parent turn's pendingResults channel.
+//
+// IMPORTANT: This function is ONLY called for asynchronous sub-turns (Async=true).
+// For synchronous sub-turns (Async=false), results are returned directly via the function
+// return value to avoid double delivery.
+//
+// Delivery behavior:
+// - If parent turn is still running: attempts to deliver to pendingResults channel
+// - If channel is full: emits SubTurnOrphanResultEvent (result is lost from channel but tracked)
+// - If parent turn has finished: emits SubTurnOrphanResultEvent (late arrival)
+//
+// Thread safety:
+// - Reads parent state under lock, then releases lock before channel send
+// - Small race window exists but is acceptable (worst case: result becomes orphan)
+//
+// Event emissions:
+// - SubTurnResultDeliveredEvent: successful delivery to channel
+// - SubTurnOrphanResultEvent: delivery failed (parent finished or channel full)
+func deliverSubTurnResult(al *AgentLoop, parentTS *turnState, childID string, result *tools.ToolResult) {
+ // Let GC clean up the pendingResults channel; parent Finish will no longer close it.
+ // We use defer/recover to catch any unlikely channel panics if it were ever closed.
+ defer func() {
+ if r := recover(); r != nil {
+ logger.WarnCF("subturn", "recovered panic sending to pendingResults", map[string]any{
+ "parent_id": parentTS.turnID,
+ "child_id": childID,
+ "recover": r,
+ })
+ if result != nil && al != nil {
+ al.emitEvent(EventKindSubTurnOrphan,
+ parentTS.eventMeta("deliverSubTurnResult", "subturn.orphan"),
+ SubTurnOrphanPayload{ParentTurnID: parentTS.turnID, ChildTurnID: childID, Reason: "panic"},
+ )
+ }
+ }
+ }()
+ parentTS.mu.Lock()
+ isFinished := parentTS.isFinished.Load()
+ resultChan := parentTS.pendingResults
+ parentTS.mu.Unlock()
+
+ // If parent turn has already finished, treat this as an orphan result
+ if isFinished || resultChan == nil {
+ if result != nil && al != nil {
+ al.emitEvent(EventKindSubTurnOrphan,
+ parentTS.eventMeta("deliverSubTurnResult", "subturn.orphan"),
+ SubTurnOrphanPayload{ParentTurnID: parentTS.turnID, ChildTurnID: childID, Reason: "parent_finished"},
+ )
+ }
+ return
+ }
+
+ // Parent Turn is still running → attempt to deliver result
+ // We use a select statement with parentTS.Finished() to ensure that if the
+ // parent turn finishes while we are waiting to send the result (e.g. channel
+ // is full), we don't leak this goroutine by blocking forever.
+ select {
+ case resultChan <- result:
+ // Successfully delivered
+ if al != nil {
+ al.emitEvent(EventKindSubTurnResultDelivered,
+ parentTS.eventMeta("deliverSubTurnResult", "subturn.result_delivered"),
+ SubTurnResultDeliveredPayload{ContentLen: len(result.ForLLM)},
+ )
+ }
+ case <-parentTS.Finished():
+ // Parent finished while we were waiting to deliver.
+ // The result cannot be delivered to the LLM, so it becomes an orphan.
+ logger.WarnCF("subturn", "parent finished before result could be delivered", map[string]any{
+ "parent_id": parentTS.turnID,
+ "child_id": childID,
+ })
+ if result != nil && al != nil {
+ al.emitEvent(
+ EventKindSubTurnOrphan,
+ parentTS.eventMeta("deliverSubTurnResult", "subturn.orphan"),
+ SubTurnOrphanPayload{
+ ParentTurnID: parentTS.turnID,
+ ChildTurnID: childID,
+ Reason: "parent_finished_waiting",
+ },
+ )
+ }
+ }
+}
+
+// ====================== Other Types ======================
+
+// ephemeralSessionStore is an in-memory session.SessionStore used by SubTurns.
+// It does not persist to disk and auto-truncates history to maxEphemeralHistorySize.
+type ephemeralSessionStore struct {
+ mu sync.Mutex
+ history []providers.Message
+ summary string
+}
+
+func newEphemeralSession(initial []providers.Message) ephemeralSessionStoreIface {
+ s := &ephemeralSessionStore{}
+ if len(initial) > 0 {
+ s.history = append(s.history, initial...)
+ }
+ return s
+}
+
+// ephemeralSessionStoreIface is satisfied by *ephemeralSessionStore.
+// Declared so newEphemeralSession can return a typed interface.
+type ephemeralSessionStoreIface interface {
+ AddMessage(sessionKey, role, content string)
+ AddFullMessage(sessionKey string, msg providers.Message)
+ GetHistory(key string) []providers.Message
+ GetSummary(key string) string
+ SetSummary(key, summary string)
+ SetHistory(key string, history []providers.Message)
+ TruncateHistory(key string, keepLast int)
+ Save(key string) error
+ Close() error
+}
+
+func (e *ephemeralSessionStore) AddMessage(_, role, content string) {
+ e.mu.Lock()
+ defer e.mu.Unlock()
+ e.history = append(e.history, providers.Message{Role: role, Content: content})
+ e.truncateLocked()
+}
+
+func (e *ephemeralSessionStore) AddFullMessage(_ string, msg providers.Message) {
+ e.mu.Lock()
+ defer e.mu.Unlock()
+ e.history = append(e.history, msg)
+ e.truncateLocked()
+}
+
+func (e *ephemeralSessionStore) GetHistory(_ string) []providers.Message {
+ e.mu.Lock()
+ defer e.mu.Unlock()
+ out := make([]providers.Message, len(e.history))
+ copy(out, e.history)
+ return out
+}
+
+func (e *ephemeralSessionStore) GetSummary(_ string) string {
+ e.mu.Lock()
+ defer e.mu.Unlock()
+ return e.summary
+}
+
+func (e *ephemeralSessionStore) SetSummary(_, summary string) {
+ e.mu.Lock()
+ defer e.mu.Unlock()
+ e.summary = summary
+}
+
+func (e *ephemeralSessionStore) SetHistory(_ string, history []providers.Message) {
+ e.mu.Lock()
+ defer e.mu.Unlock()
+ e.history = make([]providers.Message, len(history))
+ copy(e.history, history)
+ e.truncateLocked()
+}
+
+func (e *ephemeralSessionStore) TruncateHistory(_ string, keepLast int) {
+ e.mu.Lock()
+ defer e.mu.Unlock()
+ if keepLast <= 0 {
+ e.history = nil
+ return
+ }
+
+ if keepLast >= len(e.history) {
+ return
+ }
+ e.history = e.history[len(e.history)-keepLast:]
+}
+
+func (e *ephemeralSessionStore) Save(_ string) error { return nil }
+func (e *ephemeralSessionStore) Close() error { return nil }
+
+func (e *ephemeralSessionStore) truncateLocked() {
+ if len(e.history) > maxEphemeralHistorySize {
+ e.history = e.history[len(e.history)-maxEphemeralHistorySize:]
+ }
+}
diff --git a/pkg/agent/subturn_test.go b/pkg/agent/subturn_test.go
new file mode 100644
index 000000000..bac786eb3
--- /dev/null
+++ b/pkg/agent/subturn_test.go
@@ -0,0 +1,2067 @@
+package agent
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "sync"
+ "testing"
+ "time"
+
+ "github.com/sipeed/picoclaw/pkg/bus"
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/providers"
+ "github.com/sipeed/picoclaw/pkg/tools"
+)
+
+// Test constants (use defaults from subturn.go)
+const (
+ testMaxConcurrentSubTurns = defaultMaxConcurrentSubTurns
+)
+
+// ====================== Test Helper: Event Collector ======================
+type eventCollector struct {
+ mu sync.Mutex
+ events []Event
+}
+
+func newEventCollector(t *testing.T, al *AgentLoop) (*eventCollector, func()) {
+ t.Helper()
+ c := &eventCollector{}
+ sub := al.SubscribeEvents(16)
+ done := make(chan struct{})
+ go func() {
+ defer close(done)
+ for evt := range sub.C {
+ c.mu.Lock()
+ c.events = append(c.events, evt)
+ c.mu.Unlock()
+ }
+ }()
+ cleanup := func() {
+ al.UnsubscribeEvents(sub.ID)
+ <-done
+ }
+ return c, cleanup
+}
+
+func (c *eventCollector) hasEventOfKind(kind EventKind) bool {
+ c.mu.Lock()
+ defer c.mu.Unlock()
+ for _, e := range c.events {
+ if e.Kind == kind {
+ return true
+ }
+ }
+ return false
+}
+
+// ====================== Main Test Function ======================
+func TestSpawnSubTurn(t *testing.T) {
+ tests := []struct {
+ name string
+ parentDepth int
+ config SubTurnConfig
+ wantErr error
+ wantSpawn bool
+ wantEnd bool
+ wantDepthFail bool
+ }{
+ {
+ name: "Basic success path - Single layer sub-turn",
+ parentDepth: 0,
+ config: SubTurnConfig{
+ Model: "gpt-4o-mini",
+ Tools: []tools.Tool{}, // At least one tool
+ },
+ wantErr: nil,
+ wantSpawn: true,
+ wantEnd: true,
+ },
+ {
+ name: "Nested 2 layers - Normal",
+ parentDepth: 1,
+ config: SubTurnConfig{
+ Model: "gpt-4o-mini",
+ Tools: []tools.Tool{},
+ },
+ wantErr: nil,
+ wantSpawn: true,
+ wantEnd: true,
+ },
+ {
+ name: "Depth limit triggered - 4th layer fails",
+ parentDepth: 3,
+ config: SubTurnConfig{
+ Model: "gpt-4o-mini",
+ Tools: []tools.Tool{},
+ },
+ wantErr: ErrDepthLimitExceeded,
+ wantSpawn: false,
+ wantEnd: false,
+ wantDepthFail: true,
+ },
+ {
+ name: "Invalid config - Empty Model",
+ parentDepth: 0,
+ config: SubTurnConfig{
+ Model: "",
+ Tools: []tools.Tool{},
+ },
+ wantErr: ErrInvalidSubTurnConfig,
+ wantSpawn: false,
+ wantEnd: false,
+ },
+ }
+
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ // Prepare parent Turn
+ parent := &turnState{
+ ctx: context.Background(),
+ turnID: "parent-1",
+ depth: tt.parentDepth,
+ childTurnIDs: []string{},
+ pendingResults: make(chan *tools.ToolResult, 10),
+ session: &ephemeralSessionStore{},
+ agent: al.registry.GetDefaultAgent(),
+ }
+
+ // Subscribe to real EventBus to capture events
+ collector, collectCleanup := newEventCollector(t, al)
+ defer collectCleanup()
+
+ // Execute spawnSubTurn
+ result, err := spawnSubTurn(context.Background(), al, parent, tt.config)
+
+ // Assert errors
+ if tt.wantErr != nil {
+ if err == nil || err != tt.wantErr {
+ t.Errorf("expected error %v, got %v", tt.wantErr, err)
+ }
+ return
+ }
+ if err != nil {
+ t.Errorf("unexpected error: %v", err)
+ return
+ }
+
+ // Verify result
+ if result == nil {
+ t.Error("expected non-nil result")
+ }
+
+ // Verify event emission
+ time.Sleep(10 * time.Millisecond) // let event goroutine flush
+ if tt.wantSpawn {
+ if !collector.hasEventOfKind(EventKindSubTurnSpawn) {
+ t.Error("SubTurnSpawnEvent not emitted")
+ }
+ }
+ if tt.wantEnd {
+ if !collector.hasEventOfKind(EventKindSubTurnEnd) {
+ t.Error("SubTurnEndEvent not emitted")
+ }
+ }
+
+ // Verify turn tree
+ if len(parent.childTurnIDs) == 0 && !tt.wantDepthFail {
+ t.Error("child Turn not added to parent.childTurnIDs")
+ }
+
+ // For synchronous calls (Async=false, the default), result is returned directly
+ // and should NOT be in pendingResults. The result was already verified above.
+ // Only async calls (Async=true) would place results in pendingResults.
+ })
+ }
+}
+
+// ====================== Extra Independent Test: Ephemeral Session Isolation ======================
+func TestSpawnSubTurn_EphemeralSessionIsolation(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ // Parent uses its own ephemeral store pre-seeded with one message
+ parentSession := &ephemeralSessionStore{}
+ parentSession.AddMessage("", "user", "parent msg")
+ parent := &turnState{
+ ctx: context.Background(),
+ turnID: "parent-1",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 4),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ session: parentSession,
+ }
+
+ cfg := SubTurnConfig{Model: "gpt-4o-mini", Tools: []tools.Tool{}}
+
+ originalParentLen := len(parentSession.GetHistory(""))
+
+ _, _ = spawnSubTurn(context.Background(), al, parent, cfg)
+
+ // Parent session must be untouched — child used its own store
+ if got := len(parentSession.GetHistory("")); got != originalParentLen {
+ t.Errorf("parent session polluted: expected %d messages, got %d", originalParentLen, got)
+ }
+
+ // The child's agent.Sessions must NOT be the same pointer as the parent's session.
+ // We verify this indirectly: spawnSubTurn stores childTS in activeTurnStates during
+ // execution (deleted on return), so we can't easily grab childTS after the call.
+ // Instead, confirm that the child session is a distinct ephemeralSessionStore by
+ // checking the parent session key is only used by the parent store.
+ // If isolation is correct, parent.session.GetHistory(childID) is always empty
+ // (the child never wrote to the parent store).
+ al.activeTurnStates.Range(func(k, v any) bool {
+ // No active turns should remain after spawnSubTurn returns
+ t.Errorf("unexpected active turn state left after spawnSubTurn: key=%v", k)
+ return true
+ })
+}
+
+// ====================== Extra Independent Test: Result Delivery Path (Async) ======================
+func TestSpawnSubTurn_ResultDelivery(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ parent := &turnState{
+ ctx: context.Background(),
+ turnID: "parent-1",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 1),
+ session: &ephemeralSessionStore{},
+ }
+
+ // Set Async=true to test async result delivery via pendingResults channel
+ cfg := SubTurnConfig{Model: "gpt-4o-mini", Tools: []tools.Tool{}, Async: true}
+
+ _, _ = spawnSubTurn(context.Background(), al, parent, cfg)
+
+ // Check if pendingResults received the result (only for async calls)
+ select {
+ case res := <-parent.pendingResults:
+ if res == nil {
+ t.Error("received nil result in pendingResults")
+ }
+ default:
+ t.Error("result did not enter pendingResults for async call")
+ }
+}
+
+// ====================== Extra Independent Test: Result Delivery Path (Sync) ======================
+func TestSpawnSubTurn_ResultDeliverySync(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ parent := &turnState{
+ ctx: context.Background(),
+ turnID: "parent-sync-1",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 1),
+ session: &ephemeralSessionStore{},
+ }
+
+ // Sync call (Async=false, the default) - result should be returned directly
+ cfg := SubTurnConfig{Model: "gpt-4o-mini", Tools: []tools.Tool{}, Async: false}
+
+ result, err := spawnSubTurn(context.Background(), al, parent, cfg)
+ if err != nil {
+ t.Fatalf("unexpected error: %v", err)
+ }
+
+ // Result should be returned directly
+ if result == nil {
+ t.Error("expected non-nil result from sync call")
+ }
+
+ // pendingResults should NOT contain the result (no double delivery)
+ select {
+ case <-parent.pendingResults:
+ t.Error("sync call should not place result in pendingResults (double delivery)")
+ default:
+ // Expected - channel should be empty
+ }
+}
+
+// ====================== Extra Independent Test: Orphan Result Routing ======================
+func TestSpawnSubTurn_OrphanResultRouting(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ collector, collectCleanup := newEventCollector(t, al)
+ defer collectCleanup()
+
+ parentCtx, cancelParent := context.WithCancel(context.Background())
+ parent := &turnState{
+ ctx: parentCtx,
+ cancelFunc: cancelParent,
+ turnID: "parent-1",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 1),
+ session: &ephemeralSessionStore{},
+ }
+
+ // Simulate parent finishing before child delivers result
+ parent.Finish(false)
+
+ // Call deliverSubTurnResult directly to simulate a delayed child
+ deliverSubTurnResult(al, parent, "delayed-child", &tools.ToolResult{ForLLM: "late result"})
+
+ time.Sleep(10 * time.Millisecond) // let event goroutine flush
+ // Verify Orphan event is emitted
+ if !collector.hasEventOfKind(EventKindSubTurnOrphan) {
+ t.Error("SubTurnOrphanResultEvent not emitted for finished parent")
+ }
+
+ // Verify history is NOT polluted
+ if len(parent.session.GetHistory("")) != 0 {
+ t.Error("Parent history was polluted by orphan result")
+ }
+}
+
+// ====================== Extra Independent Test: Result Channel Registration ======================
+func TestSubTurnResultChannelRegistration(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ parent := &turnState{
+ ctx: context.Background(),
+ turnID: "parent-reg-1",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 4),
+ session: &ephemeralSessionStore{},
+ }
+
+ cfg := SubTurnConfig{Model: "gpt-4o-mini", Tools: []tools.Tool{}}
+
+ // Before spawn: channel should not be registered
+ if results := al.dequeuePendingSubTurnResults(parent.turnID); results != nil {
+ t.Error("expected no channel before spawnSubTurn")
+ }
+
+ _, _ = spawnSubTurn(context.Background(), al, parent, cfg)
+}
+
+// ====================== Extra Independent Test: Dequeue Pending SubTurn Results ======================
+func TestDequeuePendingSubTurnResults(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ sessionKey := "test-session-dequeue"
+
+ // Empty (no turnState registered) returns nil
+ if results := al.dequeuePendingSubTurnResults(sessionKey); len(results) != 0 {
+ t.Errorf("expected empty results, got %d", len(results))
+ }
+
+ // Register a turnState so dequeuePendingSubTurnResults can find it
+ ts := &turnState{
+ ctx: context.Background(),
+ turnID: sessionKey,
+ depth: 0,
+ session: &ephemeralSessionStore{},
+ pendingResults: make(chan *tools.ToolResult, 4),
+ }
+ al.activeTurnStates.Store(sessionKey, ts)
+ defer al.activeTurnStates.Delete(sessionKey)
+
+ // Put 3 results in
+ ts.pendingResults <- &tools.ToolResult{ForLLM: "result-1"}
+ ts.pendingResults <- &tools.ToolResult{ForLLM: "result-2"}
+ ts.pendingResults <- &tools.ToolResult{ForLLM: "result-3"}
+
+ results := al.dequeuePendingSubTurnResults(sessionKey)
+ if len(results) != 3 {
+ t.Errorf("expected 3 results, got %d", len(results))
+ }
+ if results[0].ForLLM != "result-1" || results[2].ForLLM != "result-3" {
+ t.Error("results order or content mismatch")
+ }
+
+ // Channel should be drained now
+ if results := al.dequeuePendingSubTurnResults(sessionKey); len(results) != 0 {
+ t.Errorf("expected empty after drain, got %d", len(results))
+ }
+
+ // After removing from activeTurnStates, returns nil
+ al.activeTurnStates.Delete(sessionKey)
+ if results := al.dequeuePendingSubTurnResults(sessionKey); results != nil {
+ t.Error("expected nil for unregistered session")
+ }
+}
+
+// ====================== Extra Independent Test: Concurrency Semaphore ======================
+func TestSubTurnConcurrencySemaphore(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ parent := &turnState{
+ ctx: context.Background(),
+ turnID: "parent-concurrency",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 10),
+ session: &ephemeralSessionStore{},
+ concurrencySem: make(chan struct{}, 2), // Only allow 2 concurrent children
+ }
+
+ cfg := SubTurnConfig{Model: "gpt-4o-mini", Tools: []tools.Tool{}}
+
+ // Spawn 2 children — should succeed immediately
+ done := make(chan bool, 3)
+ for i := 0; i < 2; i++ {
+ go func() {
+ _, _ = spawnSubTurn(context.Background(), al, parent, cfg)
+ done <- true
+ }()
+ }
+
+ // Wait a bit to ensure the first 2 are running
+ // (In real scenario they'd be blocked in runTurn, but mockProvider returns immediately)
+ // So we just verify the semaphore doesn't block when under limit
+ <-done
+ <-done
+
+ // Verify semaphore is now full (2/2 slots used, but they already released)
+ // Since mockProvider returns immediately, semaphore is already released
+ // So we can't easily test blocking without a real long-running operation
+
+ // Instead, verify that semaphore exists and has correct capacity
+ if cap(parent.concurrencySem) != 2 {
+ t.Errorf("expected semaphore capacity 2, got %d", cap(parent.concurrencySem))
+ }
+}
+
+// ====================== Extra Independent Test: Hard Abort Cascading ======================
+func TestHardAbortCascading(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ sessionKey := "test-session-abort"
+
+ // Root turn with its own independent context (not derived from child)
+ rootCtx, rootCancel := context.WithCancel(context.Background())
+ rootTS := &turnState{
+ ctx: rootCtx,
+ cancelFunc: rootCancel,
+ turnID: sessionKey,
+ depth: 0,
+ session: &ephemeralSessionStore{},
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, 5),
+ al: al,
+ }
+ al.activeTurnStates.Store(sessionKey, rootTS)
+ defer al.activeTurnStates.Delete(sessionKey)
+
+ // Child turn with an INDEPENDENT context (simulates spawnSubTurn behavior:
+ // context.WithTimeout(context.Background(), ...) — NOT derived from parent).
+ // Cascade must therefore happen via childTurnIDs traversal, not Go context tree.
+ childCtx, childCancel := context.WithCancel(context.Background())
+ childID := "child-independent"
+ childTS := &turnState{
+ ctx: childCtx,
+ cancelFunc: childCancel,
+ turnID: childID,
+ pendingResults: make(chan *tools.ToolResult, 4),
+ al: al,
+ }
+ al.activeTurnStates.Store(childID, childTS)
+ defer al.activeTurnStates.Delete(childID)
+
+ // Wire child into root's childTurnIDs (as spawnSubTurn would do)
+ rootTS.childTurnIDs = append(rootTS.childTurnIDs, childID)
+
+ // Verify neither context is canceled yet
+ select {
+ case <-rootTS.ctx.Done():
+ t.Fatal("root context should not be canceled yet")
+ default:
+ }
+ select {
+ case <-childTS.ctx.Done():
+ t.Fatal("child context should not be canceled yet (independent context)")
+ default:
+ }
+
+ // Trigger Hard Abort via al.HardAbort (goes through steering.go → Finish(true))
+ err := al.HardAbort(sessionKey)
+ if err != nil {
+ t.Fatalf("HardAbort failed: %v", err)
+ }
+
+ // Root context must be canceled
+ select {
+ case <-rootTS.ctx.Done():
+ default:
+ t.Error("root context should be canceled after HardAbort")
+ }
+
+ // Child context must be canceled via childTurnIDs cascade, NOT via Go context tree
+ select {
+ case <-childTS.ctx.Done():
+ default:
+ t.Error("child context should be canceled via childTurnIDs cascade")
+ }
+
+ // HardAbort on non-existent session should return an error
+ if err := al.HardAbort("non-existent-session"); err == nil {
+ t.Error("expected error for non-existent session")
+ }
+}
+
+// TestHardAbortSessionRollback verifies that HardAbort rolls back session history
+// to the state before the turn started, discarding all messages added during the turn.
+func TestHardAbortSessionRollback(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ // Create a session with initial history
+ sess := &ephemeralSessionStore{
+ history: []providers.Message{
+ {Role: "user", Content: "initial message 1"},
+ {Role: "assistant", Content: "initial response 1"},
+ },
+ }
+
+ // Create a root turnState with initialHistoryLength = 2
+ rootTS := &turnState{
+ ctx: context.Background(),
+ turnID: "test-session",
+ depth: 0,
+ session: sess,
+ initialHistoryLength: 2, // Snapshot: 2 messages
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, 5),
+ }
+
+ // Register the turn state
+ al.activeTurnStates.Store("test-session", rootTS)
+
+ // Simulate adding messages during the turn (e.g., user input + assistant response)
+ sess.AddMessage("", "user", "new user message")
+ sess.AddMessage("", "assistant", "new assistant response")
+
+ // Verify history grew to 4 messages
+ if len(sess.GetHistory("")) != 4 {
+ t.Fatalf("expected 4 messages before abort, got %d", len(sess.GetHistory("")))
+ }
+
+ // Trigger HardAbort
+ err := al.HardAbort("test-session")
+ if err != nil {
+ t.Fatalf("HardAbort failed: %v", err)
+ }
+
+ // Verify history rolled back to initial 2 messages
+ finalHistory := sess.GetHistory("")
+ if len(finalHistory) != 2 {
+ t.Errorf("expected history to rollback to 2 messages, got %d", len(finalHistory))
+ }
+
+ // Verify the content matches the initial state
+ if finalHistory[0].Content != "initial message 1" || finalHistory[1].Content != "initial response 1" {
+ t.Error("history content does not match initial state after rollback")
+ }
+}
+
+// TestNestedSubTurnHierarchy verifies that nested SubTurns maintain correct
+// parent-child relationships and depth tracking when recursively calling runAgentLoop.
+func TestNestedSubTurnHierarchy(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ // Track spawned turns and their depths
+ type turnInfo struct {
+ parentID string
+ childID string
+ }
+ var spawnedTurns []turnInfo
+ var mu sync.Mutex
+
+ // Subscribe to real EventBus to capture spawn events
+ sub := al.SubscribeEvents(16)
+ defer al.UnsubscribeEvents(sub.ID)
+ go func() {
+ for evt := range sub.C {
+ if evt.Kind == EventKindSubTurnSpawn {
+ p, _ := evt.Payload.(SubTurnSpawnPayload)
+ mu.Lock()
+ spawnedTurns = append(spawnedTurns, turnInfo{
+ parentID: p.ParentTurnID,
+ childID: p.Label,
+ })
+ mu.Unlock()
+ }
+ }
+ }()
+
+ // Create a root turn
+ rootSession := &ephemeralSessionStore{}
+ rootTS := &turnState{
+ ctx: context.Background(),
+ turnID: "root-turn",
+ depth: 0,
+ session: rootSession,
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, 5),
+ }
+
+ // Spawn a child (depth 1)
+ childCfg := SubTurnConfig{Model: "gpt-4o-mini"}
+ _, err := spawnSubTurn(context.Background(), al, rootTS, childCfg)
+ if err != nil {
+ t.Fatalf("failed to spawn child: %v", err)
+ }
+
+ time.Sleep(10 * time.Millisecond) // let event goroutine flush
+
+ // Verify we captured the spawn event
+ mu.Lock()
+ if len(spawnedTurns) != 1 {
+ t.Fatalf("expected 1 spawn event, got %d", len(spawnedTurns))
+ }
+ if spawnedTurns[0].parentID != "root-turn" {
+ t.Errorf("expected parent ID 'root-turn', got %s", spawnedTurns[0].parentID)
+ }
+ mu.Unlock()
+
+ // Verify root turn has the child in its childTurnIDs
+ rootTS.mu.Lock()
+ if len(rootTS.childTurnIDs) != 1 {
+ t.Errorf("expected root to have 1 child, got %d", len(rootTS.childTurnIDs))
+ }
+ rootTS.mu.Unlock()
+}
+
+// TestDeliverSubTurnResultNoDeadlock verifies that deliverSubTurnResult doesn't
+// deadlock when multiple goroutines are accessing the parent turnState concurrently.
+func TestDeliverSubTurnResultNoDeadlock(t *testing.T) {
+ parent := &turnState{
+ ctx: context.Background(),
+ turnID: "parent-deadlock-test",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 2), // Small buffer to test blocking
+ }
+
+ // Simulate multiple child turns delivering results concurrently
+ var wg sync.WaitGroup
+ numChildren := 10
+
+ for i := 0; i < numChildren; i++ {
+ wg.Add(1)
+ go func(id int) {
+ defer wg.Done()
+ result := &tools.ToolResult{ForLLM: fmt.Sprintf("result-%d", id)}
+ deliverSubTurnResult(nil, parent, fmt.Sprintf("child-%d", id), result)
+ }(i)
+ }
+
+ // Concurrently read from the channel to prevent blocking
+ // and to actually retrieve the matched number of results
+ go func() {
+ for i := 0; i < numChildren; i++ {
+ select {
+ case <-parent.pendingResults:
+ case <-time.After(5 * time.Second):
+ t.Error("timeout waiting for result")
+ return
+ }
+ }
+ }()
+
+ // Wait for all deliveries to complete (with timeout)
+ done := make(chan struct{})
+ go func() {
+ wg.Wait()
+ close(done)
+ }()
+
+ select {
+ case <-done:
+ // Success - no deadlock
+ case <-time.After(3 * time.Second):
+ t.Fatal("deadlock detected: deliverSubTurnResult blocked")
+ }
+}
+
+// TestHardAbortOrderOfOperations verifies that HardAbort calls Finish() before
+// rolling back session history, minimizing the race window where new messages
+// could be added after rollback.
+func TestHardAbortOrderOfOperations(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ sess := &ephemeralSessionStore{
+ history: []providers.Message{
+ {Role: "user", Content: "initial message"},
+ {Role: "assistant", Content: "response 1"},
+ {Role: "user", Content: "follow-up"},
+ },
+ }
+
+ ctx, cancel := context.WithCancel(context.Background())
+ defer cancel()
+
+ rootTS := &turnState{
+ ctx: ctx,
+ cancelFunc: cancel,
+ turnID: "test-session-order",
+ depth: 0,
+ session: sess,
+ initialHistoryLength: 1, // Snapshot: 1 message
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, 5),
+ }
+
+ al.activeTurnStates.Store("test-session-order", rootTS)
+
+ // Trigger HardAbort
+ err := al.HardAbort("test-session-order")
+ if err != nil {
+ t.Fatalf("HardAbort failed: %v", err)
+ }
+
+ // Verify context was canceled (Finish() was called)
+ select {
+ case <-rootTS.ctx.Done():
+ // Good - context was canceled
+ default:
+ t.Error("expected context to be canceled after HardAbort")
+ }
+
+ // Verify history was rolled back
+ finalHistory := sess.GetHistory("")
+ if len(finalHistory) != 1 {
+ t.Errorf("expected history to rollback to 1 message, got %d", len(finalHistory))
+ }
+
+ if finalHistory[0].Content != "initial message" {
+ t.Error("history content does not match initial state after rollback")
+ }
+}
+
+// TestFinishedChannelClosedState verifies that Finish() closes the Finished() channel
+// so that child turns can safely abort waiting.
+func TestFinishedChannelClosedState(t *testing.T) {
+ ctx, cancel := context.WithCancel(context.Background())
+ defer cancel()
+
+ ts := &turnState{
+ ctx: ctx,
+ cancelFunc: cancel,
+ turnID: "test-finished-channel",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 2),
+ }
+
+ // Verify Finished channel is blocking initially
+ select {
+ case <-ts.Finished():
+ t.Fatal("finished channel should block initially")
+ default:
+ // Good
+ }
+
+ // Call Finish() with graceful finish
+ ts.Finish(false)
+
+ // Verify Finished channel is closed
+ select {
+ case _, ok := <-ts.Finished():
+ if ok {
+ t.Error("expected Finished() channel to be closed after Finish()")
+ }
+ default:
+ t.Fatal("expected <-ts.Finished() to not block")
+ }
+
+ // Verify Finish() is idempotent
+ ts.Finish(false) // Should not panic
+
+ // Verify deliverSubTurnResult correctly uses Finished() channel and treats as orphan
+ result := &tools.ToolResult{ForLLM: "late result"}
+ deliverSubTurnResult(nil, ts, "child-1", result) // Will emit orphan due to <-ts.Finished() case
+}
+
+// TestFinalPollCapturesLateResults verifies that the final poll before Finish()
+// captures results that arrive after the last iteration poll.
+func TestFinalPollCapturesLateResults(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ sessionKey := "test-session-final-poll"
+
+ // Register a turnState
+ ts := &turnState{
+ ctx: context.Background(),
+ turnID: sessionKey,
+ depth: 0,
+ session: &ephemeralSessionStore{},
+ pendingResults: make(chan *tools.ToolResult, 4),
+ }
+ al.activeTurnStates.Store(sessionKey, ts)
+ defer al.activeTurnStates.Delete(sessionKey)
+
+ // Simulate results arriving after last iteration poll
+ ts.pendingResults <- &tools.ToolResult{ForLLM: "result 1"}
+ ts.pendingResults <- &tools.ToolResult{ForLLM: "result 2"}
+
+ // Dequeue should capture both results
+ results := al.dequeuePendingSubTurnResults(sessionKey)
+
+ if len(results) != 2 {
+ t.Errorf("expected 2 results, got %d", len(results))
+ }
+
+ // Verify channel is now empty
+ results = al.dequeuePendingSubTurnResults(sessionKey)
+ if len(results) != 0 {
+ t.Errorf("expected 0 results on second poll, got %d", len(results))
+ }
+}
+
+// TestSpawnSubTurn_PanicRecovery verifies that even if runTurn panics,
+// the result is still delivered for async calls and SubTurnEndEvent is emitted.
+func TestSpawnSubTurn_PanicRecovery(t *testing.T) {
+ // Create a panic provider
+ panicProvider := &panicMockProvider{}
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Workspace: t.TempDir(),
+ Model: "test-model",
+ MaxTokens: 4096,
+ MaxToolIterations: 10,
+ },
+ },
+ }
+ al := NewAgentLoop(cfg, bus.NewMessageBus(), panicProvider)
+
+ parent := &turnState{
+ ctx: context.Background(),
+ turnID: "parent-panic",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 1),
+ session: &ephemeralSessionStore{},
+ }
+
+ collector, collectCleanup := newEventCollector(t, al)
+ defer collectCleanup()
+
+ // Test async call - result should still be delivered via channel
+ asyncCfg := SubTurnConfig{Model: "gpt-4o-mini", Tools: []tools.Tool{}, Async: true}
+ result, err := spawnSubTurn(context.Background(), al, parent, asyncCfg)
+
+ // Should return error from panic recovery
+ if err == nil {
+ t.Error("expected error from panic recovery")
+ }
+
+ // Result should be nil because panic occurred before runTurn could return
+ if result != nil {
+ t.Error("expected nil result after panic")
+ }
+
+ time.Sleep(10 * time.Millisecond) // let event goroutine flush
+ // SubTurnEndEvent should still be emitted
+ if !collector.hasEventOfKind(EventKindSubTurnEnd) {
+ t.Error("SubTurnEndEvent not emitted after panic")
+ }
+
+ // For async call, result should still be delivered to channel (even if nil)
+ select {
+ case res := <-parent.pendingResults:
+ // Result was delivered (nil due to panic)
+ _ = res
+ default:
+ t.Error("async result should be delivered to channel even after panic")
+ }
+}
+
+// panicMockProvider is a mock provider that always panics
+type panicMockProvider struct{}
+
+func (m *panicMockProvider) Chat(
+ ctx context.Context,
+ messages []providers.Message,
+ tools []providers.ToolDefinition,
+ model string,
+ opts map[string]any,
+) (*providers.LLMResponse, error) {
+ panic("intentional panic for testing")
+}
+
+func (m *panicMockProvider) GetDefaultModel() string {
+ return "panic-model"
+}
+
+// ====================== Public API Tests ======================
+
+// simpleMockProviderAPI for testing public APIs
+type simpleMockProviderAPI struct {
+ response string
+}
+
+func (m *simpleMockProviderAPI) Chat(
+ ctx context.Context,
+ messages []providers.Message,
+ toolDefs []providers.ToolDefinition,
+ model string,
+ options map[string]any,
+) (*providers.LLMResponse, error) {
+ return &providers.LLMResponse{
+ Content: m.response,
+ }, nil
+}
+
+func (m *simpleMockProviderAPI) GetDefaultModel() string {
+ return "gpt-4o-mini"
+}
+
+// TestGetActiveTurn verifies that GetActiveTurn returns correct turn information
+func TestGetActiveTurn(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Model: "gpt-4o-mini",
+ Provider: "mock",
+ },
+ },
+ }
+ al := NewAgentLoop(cfg, nil, &simpleMockProviderAPI{response: "ok"})
+
+ // Create a root turn state
+ rootCtx := context.Background()
+ rootTS := &turnState{
+ ctx: rootCtx,
+ turnID: "root-turn",
+ parentTurnID: "",
+ depth: 0,
+ childTurnIDs: []string{},
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+
+ sessionKey := "test-session"
+ al.activeTurnStates.Store(sessionKey, rootTS)
+ defer al.activeTurnStates.Delete(sessionKey)
+
+ // Test: GetActiveTurn should return turn info
+ info := al.GetActiveTurnBySession(sessionKey)
+ if info == nil {
+ t.Fatal("GetActiveTurn returned nil for active session")
+ }
+
+ if info.TurnID != "root-turn" {
+ t.Errorf("Expected TurnID 'root-turn', got %q", info.TurnID)
+ }
+
+ if info.Depth != 0 {
+ t.Errorf("Expected Depth 0, got %d", info.Depth)
+ }
+
+ if info.ParentTurnID != "" {
+ t.Errorf("Expected empty ParentTurnID, got %q", info.ParentTurnID)
+ }
+
+ if len(info.ChildTurnIDs) != 0 {
+ t.Errorf("Expected 0 child turns, got %d", len(info.ChildTurnIDs))
+ }
+
+ // Test: GetActiveTurn should return nil for non-existent session
+ nonExistentInfo := al.GetActiveTurnBySession("non-existent-session")
+ if nonExistentInfo != nil {
+ t.Error("GetActiveTurn should return nil for non-existent session")
+ }
+}
+
+// TestGetActiveTurn_WithChildren verifies that child turn IDs are correctly reported
+func TestGetActiveTurn_WithChildren(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Model: "gpt-4o-mini",
+ Provider: "mock",
+ },
+ },
+ }
+ al := NewAgentLoop(cfg, nil, &simpleMockProviderAPI{response: "ok"})
+
+ rootCtx := context.Background()
+ rootTS := &turnState{
+ ctx: rootCtx,
+ turnID: "root-turn",
+ parentTurnID: "",
+ depth: 0,
+ childTurnIDs: []string{"child-1", "child-2"},
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+
+ sessionKey := "test-session-with-children"
+ al.activeTurnStates.Store(sessionKey, rootTS)
+ defer al.activeTurnStates.Delete(sessionKey)
+
+ info := al.GetActiveTurnBySession(sessionKey)
+ if info == nil {
+ t.Fatal("GetActiveTurn returned nil")
+ }
+
+ if len(info.ChildTurnIDs) != 2 {
+ t.Fatalf("Expected 2 child turns, got %d", len(info.ChildTurnIDs))
+ }
+
+ if info.ChildTurnIDs[0] != "child-1" || info.ChildTurnIDs[1] != "child-2" {
+ t.Errorf("Child turn IDs mismatch: got %v", info.ChildTurnIDs)
+ }
+}
+
+// TestTurnStateInfo_ThreadSafety verifies that Info() is thread-safe
+func TestTurnStateInfo_ThreadSafety(t *testing.T) {
+ rootCtx := context.Background()
+ ts := &turnState{
+ ctx: rootCtx,
+ turnID: "test-turn",
+ parentTurnID: "parent",
+ depth: 1,
+ childTurnIDs: []string{},
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+
+ // Concurrently read Info() and modify childTurnIDs
+ done := make(chan bool)
+ go func() {
+ for i := 0; i < 100; i++ {
+ ts.mu.Lock()
+ ts.childTurnIDs = append(ts.childTurnIDs, "child")
+ ts.mu.Unlock()
+ }
+ done <- true
+ }()
+
+ go func() {
+ for i := 0; i < 100; i++ {
+ info := ts.snapshot()
+ if info.TurnID == "" {
+ t.Error("snapshot() returned empty TurnID")
+ }
+ }
+ done <- true
+ }()
+
+ <-done
+ <-done
+}
+
+// TestInjectFollowUp verifies that InjectFollowUp enqueues messages
+func TestInjectFollowUp(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Model: "gpt-4o-mini",
+ Provider: "mock",
+ },
+ },
+ }
+
+ al := NewAgentLoop(cfg, nil, &simpleMockProviderAPI{response: "ok"})
+
+ msg := providers.Message{
+ Role: "user",
+ Content: "Follow-up task",
+ }
+
+ err := al.InjectFollowUp(msg)
+ if err != nil {
+ t.Fatalf("InjectFollowUp failed: %v", err)
+ }
+
+ // Verify message was enqueued
+ if al.steering.len() != 1 {
+ t.Errorf("Expected 1 message in queue, got %d", al.steering.len())
+ }
+}
+
+// TestAPIAliases verifies that API aliases work correctly
+func TestAPIAliases(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Model: "gpt-4o-mini",
+ Provider: "mock",
+ },
+ },
+ }
+
+ al := NewAgentLoop(cfg, nil, &simpleMockProviderAPI{response: "ok"})
+
+ msg := providers.Message{
+ Role: "user",
+ Content: "Test message",
+ }
+
+ // Test InterruptGraceful: requires active turn, so error is expected here
+ _ = al.InterruptGraceful(msg.Content)
+
+ // Test InjectSteering (enqueues a steering message)
+ err := al.InjectSteering(msg)
+ if err != nil {
+ t.Errorf("InjectSteering failed: %v", err)
+ }
+
+ // Also enqueue via Steer to verify second message
+ err = al.Steer(msg)
+ if err != nil {
+ t.Errorf("Steer failed: %v", err)
+ }
+
+ // Verify both messages were enqueued
+ if al.steering.len() != 2 {
+ t.Errorf("Expected 2 messages in queue, got %d", al.steering.len())
+ }
+}
+
+// TestInterruptHard_Alias verifies that InterruptHard is an alias for HardAbort
+func TestInterruptHard_Alias(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Model: "gpt-4o-mini",
+ Provider: "mock",
+ },
+ },
+ }
+ al := NewAgentLoop(cfg, nil, &simpleMockProviderAPI{response: "ok"})
+
+ rootCtx := context.Background()
+ rootTS := &turnState{
+ ctx: rootCtx,
+ turnID: "test-turn",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ initialHistoryLength: 0,
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+
+ sessionKey := "test-session-interrupt"
+ al.activeTurnStates.Store(sessionKey, rootTS)
+
+ // Test InterruptHard (alias for HardAbort)
+ err := al.InterruptHard()
+ if err != nil {
+ t.Errorf("InterruptHard failed: %v", err)
+ }
+
+ // Verify turn was finished (removed from activeTurnStates)
+ info := al.GetActiveTurnBySession(sessionKey)
+ _ = info // turn may still be in map briefly; hard abort sets isFinished on the state
+}
+
+// TestFinish_ConcurrentCalls verifies that calling Finish() concurrently from multiple
+// goroutines is safe and doesn't cause panics or double-close errors.
+func TestFinish_ConcurrentCalls(t *testing.T) {
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-concurrent-finish",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+
+ // Launch multiple goroutines that all call Finish() concurrently
+ const numGoroutines = 10
+ var wg sync.WaitGroup
+ wg.Add(numGoroutines)
+
+ for i := 0; i < numGoroutines; i++ {
+ go func() {
+ defer wg.Done()
+ // This should not panic, even when called concurrently
+ parentTS.Finish(false)
+ }()
+ }
+
+ wg.Wait()
+
+ // Verify the Finished() channel is closed
+ select {
+ case _, ok := <-parentTS.Finished():
+ if ok {
+ t.Error("Expected Finished() channel to be closed")
+ }
+ default:
+ t.Error("Expected Finished() channel to be closed and readable without blocking")
+ }
+
+ // Verify isFinished is set
+ parentTS.mu.Lock()
+ if !parentTS.isFinished.Load() {
+ t.Error("Expected isFinished to be true")
+ }
+ parentTS.mu.Unlock()
+}
+
+// TestDeliverSubTurnResult_RaceWithFinish verifies that deliverSubTurnResult handles
+// the race condition where Finish() is called while results are being delivered.
+func TestDeliverSubTurnResult_RaceWithFinish(t *testing.T) {
+ al, _, _, _, cleanup := newTestAgentLoop(t) //nolint:dogsled
+ defer cleanup()
+
+ // Collect events via real EventBus
+ var mu sync.Mutex
+ var deliveredCount, orphanCount int
+ sub := al.SubscribeEvents(64)
+ defer al.UnsubscribeEvents(sub.ID)
+ go func() {
+ for evt := range sub.C {
+ mu.Lock()
+ switch evt.Kind {
+ case EventKindSubTurnResultDelivered:
+ deliveredCount++
+ case EventKindSubTurnOrphan:
+ orphanCount++
+ }
+ mu.Unlock()
+ }
+ }()
+
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-race-test",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+
+ // Launch goroutines that deliver results while another goroutine calls Finish()
+ const numResults = 20
+ var wg sync.WaitGroup
+ wg.Add(numResults + 1)
+
+ // Goroutine that calls Finish() after a short delay
+ go func() {
+ defer wg.Done()
+ time.Sleep(5 * time.Millisecond)
+ parentTS.Finish(false)
+ }()
+
+ // Goroutines that deliver results
+ for i := 0; i < numResults; i++ {
+ go func(id int) {
+ defer wg.Done()
+ result := &tools.ToolResult{
+ ForLLM: fmt.Sprintf("result-%d", id),
+ }
+ // This should not panic, even if Finish() is called concurrently
+ deliverSubTurnResult(al, parentTS, fmt.Sprintf("child-%d", id), result)
+ }(i)
+ }
+
+ wg.Wait()
+ time.Sleep(20 * time.Millisecond) // let event goroutine flush
+
+ // Get final counts
+ mu.Lock()
+ finalDelivered := deliveredCount
+ finalOrphan := orphanCount
+ mu.Unlock()
+
+ t.Logf("Delivered: %d, Orphan: %d, Total: %d", finalDelivered, finalOrphan, finalDelivered+finalOrphan)
+
+ // With the new drainPendingResults behavior, the total events may be >= numResults
+ // because Finish() drains remaining results from the channel and emits them as orphans.
+ // So we expect:
+ // - Some results were delivered successfully (before Finish())
+ // - Some results became orphans (after Finish() or channel full)
+ // - Some results were in the channel when Finish() was called and got drained as orphans
+ // The total should be at least numResults (could be more due to drain)
+ if finalDelivered+finalOrphan < numResults {
+ t.Errorf("Expected at least %d total events, got %d delivered + %d orphan = %d",
+ numResults, finalDelivered, finalOrphan, finalDelivered+finalOrphan)
+ }
+
+ // Should have at least some orphan results (those that arrived after Finish() or were drained)
+ if finalOrphan == 0 {
+ t.Error("Expected at least some orphan results after Finish()")
+ }
+}
+
+// TestConcurrencySemaphore_Timeout verifies that spawning sub-turns times out
+// when all concurrency slots are occupied for too long.
+// Note: This test uses a shorter timeout by temporarily modifying the constant.
+func TestConcurrencySemaphore_Timeout(t *testing.T) {
+ // This test would take 30 seconds with the default timeout.
+ // Instead, we'll test the mechanism by verifying the timeout context is created correctly.
+ // A full integration test with actual timeout would be too slow for unit tests.
+
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Provider: "mock",
+ },
+ },
+ }
+ msgBus := bus.NewMessageBus()
+ provider := &simpleMockProviderAPI{}
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-timeout-test",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+ defer parentTS.Finish(false)
+
+ // Fill all concurrency slots
+ for i := 0; i < testMaxConcurrentSubTurns; i++ {
+ parentTS.concurrencySem <- struct{}{}
+ }
+
+ // Create a context with a very short timeout for testing
+ testCtx, cancel := context.WithTimeout(ctx, 100*time.Millisecond)
+ defer cancel()
+
+ // Now try to spawn a sub-turn with the short timeout context
+ subTurnCfg := SubTurnConfig{
+ Model: "gpt-4o-mini",
+ Async: false,
+ }
+
+ start := time.Now()
+ _, err := spawnSubTurn(testCtx, al, parentTS, subTurnCfg)
+ elapsed := time.Since(start)
+
+ // Should get a timeout error (either from our timeout context or the internal one)
+ if err == nil {
+ t.Error("Expected timeout error, got nil")
+ }
+
+ // The error should be related to context cancellation or timeout
+ if !errors.Is(err, context.DeadlineExceeded) && !errors.Is(err, ErrConcurrencyTimeout) {
+ t.Logf("Got error: %v (type: %T)", err, err)
+ // This is acceptable - the error might be wrapped
+ }
+
+ // Should timeout quickly (within a reasonable margin)
+ if elapsed > 2*time.Second {
+ t.Errorf("Timeout took too long: %v", elapsed)
+ }
+
+ t.Logf("Timeout occurred after %v with error: %v", elapsed, err)
+
+ // Clean up - drain the semaphore
+ for i := 0; i < testMaxConcurrentSubTurns; i++ {
+ <-parentTS.concurrencySem
+ }
+}
+
+// TestEphemeralSession_AutoTruncate verifies that ephemeral sessions automatically
+// truncate their history to prevent memory accumulation.
+func TestEphemeralSession_AutoTruncate(t *testing.T) {
+ store := newEphemeralSession(nil).(*ephemeralSessionStore)
+
+ // Add more messages than the limit
+ for i := 0; i < maxEphemeralHistorySize+20; i++ {
+ store.AddMessage("test", "user", fmt.Sprintf("message-%d", i))
+ }
+
+ // Verify history is truncated to the limit
+ history := store.GetHistory("test")
+ if len(history) != maxEphemeralHistorySize {
+ t.Errorf("Expected history length %d, got %d", maxEphemeralHistorySize, len(history))
+ }
+
+ // Verify we kept the most recent messages
+ lastMsg := history[len(history)-1]
+ expectedContent := fmt.Sprintf("message-%d", maxEphemeralHistorySize+20-1)
+ if lastMsg.Content != expectedContent {
+ t.Errorf("Expected last message to be %q, got %q", expectedContent, lastMsg.Content)
+ }
+
+ // Verify the oldest messages were discarded
+ firstMsg := history[0]
+ expectedFirstContent := fmt.Sprintf("message-%d", 20) // First 20 were discarded
+ if firstMsg.Content != expectedFirstContent {
+ t.Errorf("Expected first message to be %q, got %q", expectedFirstContent, firstMsg.Content)
+ }
+}
+
+// TestContextWrapping_SingleLayer verifies that we only create one context layer
+// in spawnSubTurn, not multiple redundant layers.
+func TestContextWrapping_SingleLayer(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Provider: "mock",
+ },
+ },
+ }
+ msgBus := bus.NewMessageBus()
+ provider := &simpleMockProviderAPI{}
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-context-test",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+ defer parentTS.Finish(false)
+
+ // Spawn a sub-turn
+ subTurnCfg := SubTurnConfig{
+ Model: "gpt-4o-mini",
+ Async: false,
+ }
+
+ result, err := spawnSubTurn(ctx, al, parentTS, subTurnCfg)
+ if err != nil {
+ t.Fatalf("spawnSubTurn failed: %v", err)
+ }
+
+ if result == nil {
+ t.Error("Expected non-nil result")
+ }
+
+ // Verify the child turn was created with a cancel function
+ // (This is implicit - if the test passes without hanging, the context management is correct)
+ t.Log("Context wrapping test passed - no redundant layers detected")
+}
+
+// TestSyncSubTurn_NoChannelDelivery verifies that synchronous sub-turns
+// do NOT deliver results to the pendingResults channel (only return directly).
+func TestSyncSubTurn_NoChannelDelivery(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Provider: "mock",
+ },
+ },
+ }
+ msgBus := bus.NewMessageBus()
+ provider := &simpleMockProviderAPI{}
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-sync-test",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+ defer parentTS.Finish(false)
+
+ // Spawn a SYNCHRONOUS sub-turn (Async=false)
+ subTurnCfg := SubTurnConfig{
+ Model: "gpt-4o-mini",
+ Async: false, // Synchronous - should NOT deliver to channel
+ }
+
+ result, err := spawnSubTurn(ctx, al, parentTS, subTurnCfg)
+ if err != nil {
+ t.Fatalf("spawnSubTurn failed: %v", err)
+ }
+
+ if result == nil {
+ t.Error("Expected non-nil result from synchronous sub-turn")
+ }
+
+ // Verify the pendingResults channel is EMPTY
+ // (synchronous sub-turns should not deliver to channel)
+ select {
+ case r := <-parentTS.pendingResults:
+ t.Errorf("Expected empty channel for sync sub-turn, but got result: %v", r)
+ default:
+ // Expected: channel is empty
+ t.Log("Verified: synchronous sub-turn did not deliver to channel")
+ }
+
+ // Verify channel length is 0
+ if len(parentTS.pendingResults) != 0 {
+ t.Errorf("Expected channel length 0, got %d", len(parentTS.pendingResults))
+ }
+}
+
+// TestAsyncSubTurn_ChannelDelivery verifies that asynchronous sub-turns
+// DO deliver results to the pendingResults channel.
+func TestAsyncSubTurn_ChannelDelivery(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Provider: "mock",
+ },
+ },
+ }
+ msgBus := bus.NewMessageBus()
+ provider := &simpleMockProviderAPI{}
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-async-test",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+ defer parentTS.Finish(false)
+
+ // Spawn an ASYNCHRONOUS sub-turn (Async=true)
+ subTurnCfg := SubTurnConfig{
+ Model: "gpt-4o-mini",
+ Async: true, // Asynchronous - SHOULD deliver to channel
+ }
+
+ result, err := spawnSubTurn(ctx, al, parentTS, subTurnCfg)
+ if err != nil {
+ t.Fatalf("spawnSubTurn failed: %v", err)
+ }
+
+ if result == nil {
+ t.Error("Expected non-nil result from asynchronous sub-turn")
+ }
+
+ // Verify the pendingResults channel has the result
+ select {
+ case r := <-parentTS.pendingResults:
+ if r == nil {
+ t.Error("Expected non-nil result from channel")
+ }
+ t.Log("Verified: asynchronous sub-turn delivered to channel")
+ case <-time.After(100 * time.Millisecond):
+ t.Error("Expected result in channel for async sub-turn, but channel was empty")
+ }
+}
+
+// TestGrandchildAbort_CascadingCancellation verifies that when a grandparent turn
+// is hard aborted, the cancellation cascades down to grandchild turns.
+func TestGrandchildAbort_CascadingCancellation(t *testing.T) {
+ al, _, _, provider, cleanup := newTestAgentLoop(t)
+ _ = provider
+ defer cleanup()
+
+ // Three independent contexts — none derived from another.
+ // Cascade must happen exclusively through childTurnIDs traversal in Finish(true).
+ gpCtx, gpCancel := context.WithCancel(context.Background())
+ parentCtx, parentCancel := context.WithCancel(context.Background())
+ childCtx, childCancel := context.WithCancel(context.Background())
+
+ childTS := &turnState{
+ ctx: childCtx,
+ cancelFunc: childCancel,
+ turnID: "grandchild",
+ al: al,
+ }
+ parentTS := &turnState{
+ ctx: parentCtx,
+ cancelFunc: parentCancel,
+ turnID: "parent",
+ childTurnIDs: []string{"grandchild"},
+ al: al,
+ }
+ grandparentTS := &turnState{
+ ctx: gpCtx,
+ cancelFunc: gpCancel,
+ turnID: "grandparent",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ childTurnIDs: []string{"parent"},
+ al: al,
+ }
+
+ al.activeTurnStates.Store("grandparent", grandparentTS)
+ al.activeTurnStates.Store("parent", parentTS)
+ al.activeTurnStates.Store("grandchild", childTS)
+ defer al.activeTurnStates.Delete("grandparent")
+ defer al.activeTurnStates.Delete("parent")
+ defer al.activeTurnStates.Delete("grandchild")
+
+ // All contexts must be active before the abort
+ for _, ctx := range []context.Context{gpCtx, parentCtx, childCtx} {
+ select {
+ case <-ctx.Done():
+ t.Fatal("context should not be canceled yet")
+ default:
+ }
+ }
+
+ // Hard abort the grandparent — should cascade to parent and grandchild
+ grandparentTS.Finish(true)
+
+ time.Sleep(10 * time.Millisecond)
+
+ select {
+ case <-gpCtx.Done():
+ t.Log("Grandparent context canceled (expected)")
+ default:
+ t.Error("Grandparent context should be canceled")
+ }
+ select {
+ case <-parentCtx.Done():
+ t.Log("Parent context canceled via cascade (expected)")
+ default:
+ t.Error("Parent context should be canceled via childTurnIDs cascade")
+ }
+ select {
+ case <-childCtx.Done():
+ t.Log("Grandchild context canceled via cascade (expected)")
+ default:
+ t.Error("Grandchild context should be canceled via childTurnIDs cascade")
+ }
+}
+
+// TestSpawnDuringAbort_RaceCondition verifies behavior when trying to spawn
+// a sub-turn while the parent is being aborted.
+func TestSpawnDuringAbort_RaceCondition(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Provider: "mock",
+ },
+ },
+ }
+ msgBus := bus.NewMessageBus()
+ provider := &simpleMockProviderAPI{}
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-abort-race",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+
+ var wg sync.WaitGroup
+ wg.Add(2)
+
+ var spawnErr error
+
+ // Goroutine 1: Try to spawn a sub-turn
+ go func() {
+ defer wg.Done()
+ subTurnCfg := SubTurnConfig{
+ Model: "gpt-4o-mini",
+ Async: false,
+ }
+ _, err := spawnSubTurn(parentTS.ctx, al, parentTS, subTurnCfg)
+ spawnErr = err
+ }()
+
+ // Goroutine 2: Abort the parent almost immediately
+ go func() {
+ defer wg.Done()
+ time.Sleep(1 * time.Millisecond)
+ parentTS.Finish(false)
+ }()
+
+ wg.Wait()
+
+ // The spawn should either succeed (if it started before abort)
+ // or fail with context canceled error (if abort happened first)
+ if spawnErr != nil {
+ if errors.Is(spawnErr, context.Canceled) {
+ t.Logf("Spawn failed with expected context cancellation: %v", spawnErr)
+ } else {
+ t.Logf("Spawn failed with error: %v", spawnErr)
+ }
+ } else {
+ t.Log("Spawn succeeded before abort")
+ }
+
+ // The important thing is that it doesn't panic or deadlock
+ t.Log("Race condition handled gracefully - no panic or deadlock")
+}
+
+// ====================== Slow SubTurn Cancellation Test ======================
+
+// slowMockProvider simulates a slow LLM call that takes a long time to complete.
+// This is used to test the scenario where a parent turn finishes before the child SubTurn.
+type slowMockProvider struct {
+ delay time.Duration
+}
+
+func (m *slowMockProvider) Chat(
+ ctx context.Context,
+ messages []providers.Message,
+ toolDefs []providers.ToolDefinition,
+ model string,
+ options map[string]any,
+) (*providers.LLMResponse, error) {
+ select {
+ case <-time.After(m.delay):
+ // Completed normally after delay
+ return &providers.LLMResponse{
+ Content: "slow response completed",
+ }, nil
+ case <-ctx.Done():
+ // Context was canceled while waiting
+ return nil, ctx.Err()
+ }
+}
+
+func (m *slowMockProvider) GetDefaultModel() string {
+ return "slow-model"
+}
+
+// TestAsyncSubTurn_ParentFinishesEarly simulates the scenario where:
+// 1. Parent spawns an async SubTurn that takes a long time
+// 2. Parent finishes quickly
+// 3. SubTurn should be canceled with context canceled error
+func TestAsyncSubTurn_ParentFinishesEarly(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Provider: "mock",
+ },
+ },
+ }
+ msgBus := bus.NewMessageBus()
+ provider := &slowMockProvider{delay: 5 * time.Second} // SubTurn takes 5 seconds
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ // Capture events via real EventBus
+ var mu sync.Mutex
+ var events []Event
+ sub := al.SubscribeEvents(32)
+ defer al.UnsubscribeEvents(sub.ID)
+ go func() {
+ for evt := range sub.C {
+ mu.Lock()
+ events = append(events, evt)
+ mu.Unlock()
+ }
+ }()
+
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-fast",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+
+ var subTurnErr error
+ var subTurnResult *tools.ToolResult
+ var wg sync.WaitGroup
+
+ // Spawn async SubTurn in a goroutine (it will be slow)
+ wg.Add(1)
+ go func() {
+ defer wg.Done()
+ subTurnCfg := SubTurnConfig{
+ Model: "slow-model",
+ Async: true, // Asynchronous SubTurn
+ }
+ subTurnResult, subTurnErr = spawnSubTurn(parentTS.ctx, al, parentTS, subTurnCfg)
+ }()
+
+ // Parent finishes quickly (after 100ms), while SubTurn is still running
+ time.Sleep(100 * time.Millisecond)
+ t.Log("Parent finishing early...")
+ parentTS.Finish(false)
+
+ // Wait for SubTurn to complete (or be canceled)
+ wg.Wait()
+
+ // Check the result
+ t.Logf("SubTurn error: %v", subTurnErr)
+ t.Logf("SubTurn result: %v", subTurnResult)
+
+ if subTurnErr != nil {
+ if errors.Is(subTurnErr, context.Canceled) {
+ t.Log("✓ SubTurn was canceled as expected (context canceled)")
+ } else {
+ t.Logf("SubTurn failed with other error: %v", subTurnErr)
+ }
+ } else {
+ t.Log("SubTurn completed before parent finished (unlikely but possible)")
+ }
+
+ // Log captured events
+ mu.Lock()
+ t.Logf("Captured %d events:", len(events))
+ for i, e := range events {
+ t.Logf(" Event %d: %s", i+1, e.Kind)
+ }
+ mu.Unlock()
+}
+
+// TestAsyncSubTurn_ParentWaitsForChild simulates the scenario where:
+// 1. Parent spawns an async SubTurn that takes some time
+// 2. Parent WAITS for SubTurn to complete before finishing
+// 3. Both should complete successfully
+func TestAsyncSubTurn_ParentWaitsForChild(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Provider: "mock",
+ },
+ },
+ }
+ msgBus := bus.NewMessageBus()
+ provider := &slowMockProvider{delay: 200 * time.Millisecond} // SubTurn takes 200ms
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-wait",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+
+ var subTurnErr error
+ var subTurnResult *tools.ToolResult
+ var wg sync.WaitGroup
+
+ // Spawn async SubTurn in a goroutine
+ wg.Add(1)
+ go func() {
+ defer wg.Done()
+ subTurnCfg := SubTurnConfig{
+ Model: "slow-model",
+ Async: true,
+ }
+ subTurnResult, subTurnErr = spawnSubTurn(parentTS.ctx, al, parentTS, subTurnCfg)
+ }()
+
+ // Parent WAITS for SubTurn to complete
+ t.Log("Parent waiting for SubTurn...")
+ wg.Wait()
+ t.Log("SubTurn completed, parent now finishing")
+
+ // Now parent can finish safely
+ parentTS.Finish(false)
+
+ // Check the result
+ if subTurnErr != nil {
+ if errors.Is(subTurnErr, context.Canceled) {
+ t.Errorf("SubTurn should NOT have been canceled: %v", subTurnErr)
+ } else {
+ t.Logf("SubTurn failed with error: %v", subTurnErr)
+ }
+ } else {
+ t.Log("✓ SubTurn completed successfully")
+ if subTurnResult != nil {
+ t.Logf("SubTurn result: %s", subTurnResult.ForLLM)
+ }
+ }
+
+ // Check channel delivery
+ select {
+ case r := <-parentTS.pendingResults:
+ if r != nil {
+ t.Logf("✓ Result delivered to channel: %s", r.ForLLM)
+ }
+ case <-time.After(100 * time.Millisecond):
+ t.Log("No result in channel (expected since we waited)")
+ }
+}
+
+// ====================== Graceful vs Hard Finish Tests ======================
+
+// TestFinish_GracefulVsHard verifies the behavior difference between:
+// - Finish(false): graceful finish, signals parentEnded but doesn't cancel children
+// - Finish(true): hard abort, immediately cancels all children
+func TestFinish_GracefulVsHard(t *testing.T) {
+ // Test 1: Graceful finish should set parentEnded but not cancel context
+ t.Run("Graceful_SetsParentEnded", func(t *testing.T) {
+ ctx, cancel := context.WithCancel(context.Background())
+ defer cancel()
+
+ ts := &turnState{
+ ctx: ctx,
+ turnID: "graceful-test",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 16),
+ }
+ ts.ctx, ts.cancelFunc = context.WithCancel(ctx)
+
+ // Finish gracefully
+ ts.Finish(false)
+
+ // Verify parentEnded is set
+ if !ts.parentEnded.Load() {
+ t.Error("parentEnded should be true after graceful finish")
+ }
+
+ // Verify context is NOT canceled (for graceful finish, children continue)
+ // Note: In graceful mode, we don't call cancelFunc()
+ // But since we're using WithCancel on the same ctx, it might be canceled
+ // Let's check that the context is still valid for a moment
+ time.Sleep(10 * time.Millisecond)
+ // Context might be canceled by the deferred cancel() in test, which is fine
+ })
+
+ // Test 2: Hard abort should cancel context immediately
+ t.Run("Hard_CancelsContext", func(t *testing.T) {
+ ctx := context.Background()
+
+ ts := &turnState{
+ ctx: ctx,
+ turnID: "hard-test",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 16),
+ }
+ ts.ctx, ts.cancelFunc = context.WithCancel(ctx)
+
+ // Finish with hard abort
+ ts.Finish(true)
+
+ // Verify context is canceled
+ select {
+ case <-ts.ctx.Done():
+ t.Log("✓ Context canceled after hard abort")
+ default:
+ t.Error("Context should be canceled after hard abort")
+ }
+ })
+
+ // Test 3: IsParentEnded returns correct value
+ t.Run("IsParentEnded", func(t *testing.T) {
+ ctx := context.Background()
+
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-isended-test",
+ depth: 0,
+ pendingResults: make(chan *tools.ToolResult, 16),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+
+ childTS := &turnState{
+ ctx: ctx,
+ turnID: "child-isended-test",
+ depth: 1,
+ parentTurnState: parentTS,
+ pendingResults: make(chan *tools.ToolResult, 16),
+ }
+
+ // Before parent finishes
+ if childTS.IsParentEnded() {
+ t.Error("IsParentEnded should be false before parent finishes")
+ }
+
+ // Finish parent gracefully
+ parentTS.Finish(false)
+
+ // After parent finishes
+ if !childTS.IsParentEnded() {
+ t.Error("IsParentEnded should be true after parent finishes gracefully")
+ }
+ })
+}
+
+// TestSubTurn_IndependentContext verifies that SubTurns use independent contexts
+// that don't get canceled when the parent finishes gracefully.
+func TestSubTurn_IndependentContext(t *testing.T) {
+ cfg := &config.Config{
+ Agents: config.AgentsConfig{
+ Defaults: config.AgentDefaults{
+ Provider: "mock",
+ },
+ },
+ }
+ msgBus := bus.NewMessageBus()
+ provider := &slowMockProvider{delay: 500 * time.Millisecond}
+ al := NewAgentLoop(cfg, msgBus, provider)
+
+ ctx := context.Background()
+ parentTS := &turnState{
+ ctx: ctx,
+ turnID: "parent-independent",
+ depth: 0,
+ session: newEphemeralSession(nil),
+ pendingResults: make(chan *tools.ToolResult, 16),
+ concurrencySem: make(chan struct{}, testMaxConcurrentSubTurns),
+ }
+ parentTS.ctx, parentTS.cancelFunc = context.WithCancel(ctx)
+
+ var subTurnErr error
+ var wg sync.WaitGroup
+
+ // Spawn SubTurn with Critical=true (should continue after parent finishes)
+ wg.Add(1)
+ go func() {
+ defer wg.Done()
+ subTurnCfg := SubTurnConfig{
+ Model: "slow-model",
+ Async: true,
+ Critical: true, // Critical SubTurn should continue
+ }
+ _, subTurnErr = spawnSubTurn(parentTS.ctx, al, parentTS, subTurnCfg)
+ }()
+
+ // Let SubTurn start
+ time.Sleep(50 * time.Millisecond)
+
+ // Parent finishes gracefully (should NOT cancel SubTurn)
+ parentTS.Finish(false)
+ t.Log("Parent finished gracefully, SubTurn should continue")
+
+ // Wait for SubTurn to complete
+ wg.Wait()
+
+ // SubTurn should complete without context canceled error
+ // (because it uses independent context now)
+ if subTurnErr != nil {
+ t.Logf("SubTurn error: %v", subTurnErr)
+ // The error might be context.DeadlineExceeded if timeout is too short
+ // but should NOT be context.Canceled from parent
+ if errors.Is(subTurnErr, context.Canceled) {
+ t.Error("SubTurn should not be canceled by parent's graceful finish")
+ }
+ } else {
+ t.Log("✓ SubTurn completed successfully (independent context)")
+ }
+}
diff --git a/pkg/agent/turn.go b/pkg/agent/turn.go
index 358dae2b4..e4970c519 100644
--- a/pkg/agent/turn.go
+++ b/pkg/agent/turn.go
@@ -4,10 +4,13 @@ import (
"context"
"reflect"
"sync"
+ "sync/atomic"
"time"
"github.com/sipeed/picoclaw/pkg/bus"
"github.com/sipeed/picoclaw/pkg/providers"
+ "github.com/sipeed/picoclaw/pkg/session"
+ "github.com/sipeed/picoclaw/pkg/tools"
)
type TurnPhase string
@@ -22,15 +25,18 @@ const (
)
type ActiveTurnInfo struct {
- TurnID string
- AgentID string
- SessionKey string
- Channel string
- ChatID string
- UserMessage string
- Phase TurnPhase
- Iteration int
- StartedAt time.Time
+ TurnID string
+ AgentID string
+ SessionKey string
+ Channel string
+ ChatID string
+ UserMessage string
+ Phase TurnPhase
+ Iteration int
+ StartedAt time.Time
+ Depth int
+ ParentTurnID string
+ ChildTurnIDs []string
}
type turnResult struct {
@@ -72,10 +78,37 @@ type turnState struct {
restorePointHistory []providers.Message
restorePointSummary string
persistedMessages []providers.Message
+
+ // SubTurn support (from HEAD)
+ depth int // SubTurn depth (0 for root turn)
+ parentTurnID string // Parent turn ID (empty for root turn)
+ childTurnIDs []string // Child turn IDs
+ pendingResults chan *tools.ToolResult // Channel for SubTurn results
+ concurrencySem chan struct{} // Semaphore for limiting concurrent SubTurns
+ isFinished atomic.Bool // Whether this turn has finished
+ session session.SessionStore // Session store reference
+ initialHistoryLength int // Snapshot of history length at turn start
+
+ // Additional SubTurn fields
+ ctx context.Context // Context for this turn
+ cancelFunc context.CancelFunc // Cancel function for this turn's context
+ critical bool // Whether this SubTurn should continue after parent ends
+ parentTurnState *turnState // Reference to parent turnState
+ parentEnded atomic.Bool // Whether parent has ended
+ closeOnce sync.Once // Ensures pendingResults channel is closed once
+ finishedChan chan struct{} // Closed when turn finishes
+
+ // Token budget tracking
+ tokenBudget *atomic.Int64 // Shared token budget counter
+ lastFinishReason string // Last LLM finish_reason
+ lastUsage *providers.UsageInfo // Last LLM usage info
+
+ // Back-reference to the owning AgentLoop (set for SubTurns only, used for hard abort cascade)
+ al *AgentLoop
}
func newTurnState(agent *AgentInstance, opts processOptions, scope turnEventScope) *turnState {
- return &turnState{
+ ts := &turnState{
agent: agent,
opts: opts,
scope: scope,
@@ -89,30 +122,58 @@ func newTurnState(agent *AgentInstance, opts processOptions, scope turnEventScop
phase: TurnPhaseSetup,
startedAt: time.Now(),
}
+
+ // Bind session store and capture initial history length for rollback logic
+ if agent != nil && agent.Sessions != nil {
+ ts.session = agent.Sessions
+ ts.initialHistoryLength = len(agent.Sessions.GetHistory(opts.SessionKey))
+ }
+
+ return ts
}
func (al *AgentLoop) registerActiveTurn(ts *turnState) {
- al.activeTurnMu.Lock()
- defer al.activeTurnMu.Unlock()
- al.activeTurn = ts
+ al.activeTurnStates.Store(ts.sessionKey, ts)
}
func (al *AgentLoop) clearActiveTurn(ts *turnState) {
- al.activeTurnMu.Lock()
- defer al.activeTurnMu.Unlock()
- if al.activeTurn == ts {
- al.activeTurn = nil
- }
+ al.activeTurnStates.Delete(ts.sessionKey)
}
-func (al *AgentLoop) getActiveTurnState() *turnState {
- al.activeTurnMu.RLock()
- defer al.activeTurnMu.RUnlock()
- return al.activeTurn
+func (al *AgentLoop) getActiveTurnState(sessionKey string) *turnState {
+ if val, ok := al.activeTurnStates.Load(sessionKey); ok {
+ return val.(*turnState)
+ }
+ return nil
+}
+
+// getAnyActiveTurnState returns any active turn state (for backward compatibility)
+func (al *AgentLoop) getAnyActiveTurnState() *turnState {
+ var firstTS *turnState
+ al.activeTurnStates.Range(func(key, value any) bool {
+ firstTS = value.(*turnState)
+ return false // stop after first
+ })
+ return firstTS
}
func (al *AgentLoop) GetActiveTurn() *ActiveTurnInfo {
- ts := al.getActiveTurnState()
+ // For backward compatibility, return the first active turn found
+ // In the new architecture, there can be multiple concurrent turns
+ var firstTS *turnState
+ al.activeTurnStates.Range(func(key, value any) bool {
+ firstTS = value.(*turnState)
+ return false // stop after first
+ })
+ if firstTS == nil {
+ return nil
+ }
+ info := firstTS.snapshot()
+ return &info
+}
+
+func (al *AgentLoop) GetActiveTurnBySession(sessionKey string) *ActiveTurnInfo {
+ ts := al.getActiveTurnState(sessionKey)
if ts == nil {
return nil
}
@@ -125,15 +186,18 @@ func (ts *turnState) snapshot() ActiveTurnInfo {
defer ts.mu.RUnlock()
return ActiveTurnInfo{
- TurnID: ts.turnID,
- AgentID: ts.agentID,
- SessionKey: ts.sessionKey,
- Channel: ts.channel,
- ChatID: ts.chatID,
- UserMessage: ts.userMessage,
- Phase: ts.phase,
- Iteration: ts.iteration,
- StartedAt: ts.startedAt,
+ TurnID: ts.turnID,
+ AgentID: ts.agentID,
+ SessionKey: ts.sessionKey,
+ Channel: ts.channel,
+ ChatID: ts.chatID,
+ UserMessage: ts.userMessage,
+ Phase: ts.phase,
+ Iteration: ts.iteration,
+ StartedAt: ts.startedAt,
+ Depth: ts.depth,
+ ParentTurnID: ts.parentTurnID,
+ ChildTurnIDs: append([]string(nil), ts.childTurnIDs...),
}
}
@@ -306,3 +370,112 @@ func (ts *turnState) interruptHintMessage() providers.Message {
Content: content,
}
}
+
+// SubTurn-related methods
+
+// Finish marks the turn as finished and closes the pendingResults channel
+func (ts *turnState) Finish(isHardAbort bool) {
+ ts.isFinished.Store(true)
+
+ // Close pendingResults channel exactly once
+ ts.closeOnce.Do(func() {
+ if ts.pendingResults != nil {
+ close(ts.pendingResults)
+ }
+ ts.mu.Lock()
+ if ts.finishedChan == nil {
+ ts.finishedChan = make(chan struct{})
+ }
+ close(ts.finishedChan)
+ ts.mu.Unlock()
+ })
+
+ // If this is a graceful finish (not hard abort), signal to children
+ if !isHardAbort && ts.parentTurnState == nil {
+ // This is a root turn finishing gracefully
+ ts.parentEnded.Store(true)
+ }
+
+ // Cancel the turn context
+ if ts.cancelFunc != nil {
+ ts.cancelFunc()
+ }
+
+ // Hard abort cascades to all child turns
+ if isHardAbort && ts.al != nil {
+ ts.mu.RLock()
+ children := append([]string(nil), ts.childTurnIDs...)
+ ts.mu.RUnlock()
+ for _, childID := range children {
+ if val, ok := ts.al.activeTurnStates.Load(childID); ok {
+ val.(*turnState).Finish(true)
+ }
+ }
+ }
+}
+
+// Finished returns whether the turn has finished
+func (ts *turnState) Finished() chan struct{} {
+ ts.mu.Lock()
+ defer ts.mu.Unlock()
+ if ts.finishedChan == nil {
+ ts.finishedChan = make(chan struct{})
+ }
+ return ts.finishedChan
+}
+
+// IsParentEnded checks if the parent turn has ended
+func (ts *turnState) IsParentEnded() bool {
+ if ts.parentTurnState == nil {
+ return false
+ }
+ return ts.parentTurnState.parentEnded.Load()
+}
+
+// GetLastFinishReason returns the last LLM finish_reason
+func (ts *turnState) GetLastFinishReason() string {
+ ts.mu.RLock()
+ defer ts.mu.RUnlock()
+ return ts.lastFinishReason
+}
+
+// SetLastFinishReason sets the last LLM finish_reason
+func (ts *turnState) SetLastFinishReason(reason string) {
+ ts.mu.Lock()
+ defer ts.mu.Unlock()
+ ts.lastFinishReason = reason
+}
+
+// GetLastUsage returns the last LLM usage info
+func (ts *turnState) GetLastUsage() *providers.UsageInfo {
+ ts.mu.RLock()
+ defer ts.mu.RUnlock()
+ return ts.lastUsage
+}
+
+// SetLastUsage sets the last LLM usage info
+func (ts *turnState) SetLastUsage(usage *providers.UsageInfo) {
+ ts.mu.Lock()
+ defer ts.mu.Unlock()
+ ts.lastUsage = usage
+}
+
+// Context helper functions for SubTurn
+
+type turnStateKeyType struct{}
+
+var turnStateKey = turnStateKeyType{}
+
+func withTurnState(ctx context.Context, ts *turnState) context.Context {
+ return context.WithValue(ctx, turnStateKey, ts)
+}
+
+func turnStateFromContext(ctx context.Context) *turnState {
+ ts, _ := ctx.Value(turnStateKey).(*turnState)
+ return ts
+}
+
+// TurnStateFromContext retrieves turnState from context (exported for tools)
+func TurnStateFromContext(ctx context.Context) *turnState {
+ return turnStateFromContext(ctx)
+}
diff --git a/pkg/auth/store.go b/pkg/auth/store.go
index 2e55d4877..f7813ca57 100644
--- a/pkg/auth/store.go
+++ b/pkg/auth/store.go
@@ -6,6 +6,7 @@ import (
"path/filepath"
"time"
+ "github.com/sipeed/picoclaw/pkg/config"
"github.com/sipeed/picoclaw/pkg/fileutil"
)
@@ -39,7 +40,7 @@ func (c *AuthCredential) NeedsRefresh() bool {
}
func authFilePath() string {
- if home := os.Getenv("PICOCLAW_HOME"); home != "" {
+ if home := os.Getenv(config.EnvHome); home != "" {
return filepath.Join(home, "auth.json")
}
home, _ := os.UserHomeDir()
diff --git a/pkg/bus/bus.go b/pkg/bus/bus.go
index f5ff9587d..37fcb74c5 100644
--- a/pkg/bus/bus.go
+++ b/pkg/bus/bus.go
@@ -3,6 +3,7 @@ package bus
import (
"context"
"errors"
+ "sync"
"sync/atomic"
"github.com/sipeed/picoclaw/pkg/logger"
@@ -13,12 +14,32 @@ var ErrBusClosed = errors.New("message bus closed")
const defaultBusBufferSize = 64
+// StreamDelegate is implemented by the channel Manager to provide streaming
+// capabilities to the agent loop without tight coupling.
+type StreamDelegate interface {
+ // GetStreamer returns a Streamer for the given channel+chatID if the channel
+ // supports streaming. Returns nil, false if streaming is unavailable.
+ GetStreamer(ctx context.Context, channel, chatID string) (Streamer, bool)
+}
+
+// Streamer pushes incremental content to a streaming-capable channel.
+// Defined here so the agent loop can use it without importing pkg/channels.
+type Streamer interface {
+ Update(ctx context.Context, content string) error
+ Finalize(ctx context.Context, content string) error
+ Cancel(ctx context.Context)
+}
+
type MessageBus struct {
inbound chan InboundMessage
outbound chan OutboundMessage
outboundMedia chan OutboundMediaMessage
- done chan struct{}
- closed atomic.Bool
+
+ closeOnce sync.Once
+ done chan struct{}
+ closed atomic.Bool
+ wg sync.WaitGroup
+ streamDelegate atomic.Value // stores StreamDelegate
}
func NewMessageBus() *MessageBus {
@@ -30,128 +51,104 @@ func NewMessageBus() *MessageBus {
}
}
-func (mb *MessageBus) PublishInbound(ctx context.Context, msg InboundMessage) error {
+func publish[T any](ctx context.Context, mb *MessageBus, ch chan T, msg T) error {
+ // check bus closed before acquiring wg, to avoid unnecessary wg.Add and potential deadlock
if mb.closed.Load() {
return ErrBusClosed
}
- if err := ctx.Err(); err != nil {
- return err
- }
+
+ // check again,before sending message, to avoid sending to closed channel
select {
- case mb.inbound <- msg:
- return nil
- case <-mb.done:
- return ErrBusClosed
case <-ctx.Done():
return ctx.Err()
+ case <-mb.done:
+ return ErrBusClosed
+ default:
+ }
+
+ mb.wg.Add(1)
+ defer mb.wg.Done()
+
+ select {
+ case ch <- msg:
+ return nil
+ case <-ctx.Done():
+ return ctx.Err()
+ case <-mb.done:
+ return ErrBusClosed
}
}
-func (mb *MessageBus) ConsumeInbound(ctx context.Context) (InboundMessage, bool) {
- select {
- case msg, ok := <-mb.inbound:
- return msg, ok
- case <-mb.done:
- return InboundMessage{}, false
- case <-ctx.Done():
- return InboundMessage{}, false
- }
+func (mb *MessageBus) PublishInbound(ctx context.Context, msg InboundMessage) error {
+ return publish(ctx, mb, mb.inbound, msg)
+}
+
+func (mb *MessageBus) InboundChan() <-chan InboundMessage {
+ return mb.inbound
}
func (mb *MessageBus) PublishOutbound(ctx context.Context, msg OutboundMessage) error {
- if mb.closed.Load() {
- return ErrBusClosed
- }
- if err := ctx.Err(); err != nil {
- return err
- }
- select {
- case mb.outbound <- msg:
- return nil
- case <-mb.done:
- return ErrBusClosed
- case <-ctx.Done():
- return ctx.Err()
- }
+ return publish(ctx, mb, mb.outbound, msg)
}
-func (mb *MessageBus) SubscribeOutbound(ctx context.Context) (OutboundMessage, bool) {
- select {
- case msg, ok := <-mb.outbound:
- return msg, ok
- case <-mb.done:
- return OutboundMessage{}, false
- case <-ctx.Done():
- return OutboundMessage{}, false
- }
+func (mb *MessageBus) OutboundChan() <-chan OutboundMessage {
+ return mb.outbound
}
func (mb *MessageBus) PublishOutboundMedia(ctx context.Context, msg OutboundMediaMessage) error {
- if mb.closed.Load() {
- return ErrBusClosed
- }
- if err := ctx.Err(); err != nil {
- return err
- }
- select {
- case mb.outboundMedia <- msg:
- return nil
- case <-mb.done:
- return ErrBusClosed
- case <-ctx.Done():
- return ctx.Err()
- }
+ return publish(ctx, mb, mb.outboundMedia, msg)
}
-func (mb *MessageBus) SubscribeOutboundMedia(ctx context.Context) (OutboundMediaMessage, bool) {
- select {
- case msg, ok := <-mb.outboundMedia:
- return msg, ok
- case <-mb.done:
- return OutboundMediaMessage{}, false
- case <-ctx.Done():
- return OutboundMediaMessage{}, false
+func (mb *MessageBus) OutboundMediaChan() <-chan OutboundMediaMessage {
+ return mb.outboundMedia
+}
+
+// SetStreamDelegate registers a StreamDelegate (typically the channel Manager).
+func (mb *MessageBus) SetStreamDelegate(d StreamDelegate) {
+ mb.streamDelegate.Store(d)
+}
+
+// GetStreamer returns a Streamer for the given channel+chatID via the delegate.
+func (mb *MessageBus) GetStreamer(ctx context.Context, channel, chatID string) (Streamer, bool) {
+ if d, ok := mb.streamDelegate.Load().(StreamDelegate); ok && d != nil {
+ return d.GetStreamer(ctx, channel, chatID)
}
+ return nil, false
}
func (mb *MessageBus) Close() {
- if mb.closed.CompareAndSwap(false, true) {
+ mb.closeOnce.Do(func() {
+ // notify all blocked publishers to exit
close(mb.done)
- // Drain buffered channels so messages aren't silently lost.
- // Channels are NOT closed to avoid send-on-closed panics from concurrent publishers.
+ // because every publisher will check mb.closed before acquiring wg
+ // so we can be sure that new publishers will not be added new messages after this point
+ mb.closed.Store(true)
+
+ // wait for all ongoing Publish calls to finish, ensuring all messages have been sent to channels or exited
+ mb.wg.Wait()
+
+ // close channels safely
+ close(mb.inbound)
+ close(mb.outbound)
+ close(mb.outboundMedia)
+
+ // clean up any remaining messages in channels
drained := 0
- for {
- select {
- case <-mb.inbound:
- drained++
- default:
- goto doneInbound
- }
+ for range mb.inbound {
+ drained++
}
- doneInbound:
- for {
- select {
- case <-mb.outbound:
- drained++
- default:
- goto doneOutbound
- }
+ for range mb.outbound {
+ drained++
}
- doneOutbound:
- for {
- select {
- case <-mb.outboundMedia:
- drained++
- default:
- goto doneMedia
- }
+ for range mb.outboundMedia {
+ drained++
}
- doneMedia:
+
if drained > 0 {
logger.DebugCF("bus", "Drained buffered messages during close", map[string]any{
"count": drained,
})
}
- }
+ })
}
diff --git a/pkg/bus/bus_test.go b/pkg/bus/bus_test.go
index e07b8c7fe..9b6324ca6 100644
--- a/pkg/bus/bus_test.go
+++ b/pkg/bus/bus_test.go
@@ -24,7 +24,7 @@ func TestPublishConsume(t *testing.T) {
t.Fatalf("PublishInbound failed: %v", err)
}
- got, ok := mb.ConsumeInbound(ctx)
+ got, ok := <-mb.InboundChan()
if !ok {
t.Fatal("ConsumeInbound returned ok=false")
}
@@ -52,7 +52,7 @@ func TestPublishOutboundSubscribe(t *testing.T) {
t.Fatalf("PublishOutbound failed: %v", err)
}
- got, ok := mb.SubscribeOutbound(ctx)
+ got, ok := <-mb.OutboundChan()
if !ok {
t.Fatal("SubscribeOutbound returned ok=false")
}
@@ -108,27 +108,48 @@ func TestPublishOutbound_BusClosed(t *testing.T) {
func TestConsumeInbound_ContextCancel(t *testing.T) {
mb := NewMessageBus()
+
defer mb.Close()
- ctx, cancel := context.WithCancel(context.Background())
- cancel()
+ for i := range defaultBusBufferSize {
+ if err := mb.PublishInbound(context.Background(), InboundMessage{Content: "fill"}); err != nil {
+ t.Fatalf("fill failed at %d: %v", i, err)
+ }
+ }
- _, ok := mb.ConsumeInbound(ctx)
- if ok {
- t.Fatal("expected ok=false when context is canceled")
+ ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
+ defer cancel()
+ mb.PublishInbound(ctx, InboundMessage{Content: "ContextCancel"})
+
+ select {
+ case <-ctx.Done():
+ t.Log("context canceled, as expected")
+
+ case msg, ok := <-mb.InboundChan():
+ if !ok {
+ t.Fatal("expected ok=false when context is canceled")
+ }
+ if msg.Content == "ContextCancel" {
+ t.Fatalf("expected content 'ContextCancel', got %q", msg.Content)
+ }
}
}
func TestConsumeInbound_BusClosed(t *testing.T) {
mb := NewMessageBus()
- mb.Close()
- ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
- defer cancel()
+ timer := time.AfterFunc(100*time.Millisecond, func() {
+ mb.Close()
+ })
- _, ok := mb.ConsumeInbound(ctx)
- if ok {
- t.Fatal("expected ok=false when bus is closed")
+ select {
+ case <-timer.C:
+ t.Log("context canceled, as expected")
+
+ case _, ok := <-mb.InboundChan():
+ if ok {
+ t.Fatal("expected ok=false when context is canceled")
+ }
}
}
@@ -136,10 +157,7 @@ func TestSubscribeOutbound_BusClosed(t *testing.T) {
mb := NewMessageBus()
mb.Close()
- ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
- defer cancel()
-
- _, ok := mb.SubscribeOutbound(ctx)
+ _, ok := <-mb.OutboundChan()
if ok {
t.Fatal("expected ok=false when bus is closed")
}
diff --git a/pkg/channels/base.go b/pkg/channels/base.go
index edb5b6f08..882e72d08 100644
--- a/pkg/channels/base.go
+++ b/pkg/channels/base.go
@@ -275,14 +275,18 @@ func (c *BaseChannel) HandleMessage(
// Auto-trigger typing indicator, message reaction, and placeholder before publishing.
// Each capability is independent — all three may fire for the same message.
+ // Note: even when streaming is available, we still show typing + placeholder on inbound.
+ // If streaming actually activates, preSend will skip the placeholder edit (streamActive map)
+ // and the typing stop will still be called. This avoids the problem of compile-time interface
+ // checks incorrectly skipping indicators when streaming may not work at runtime.
if c.owner != nil && c.placeholderRecorder != nil {
- // Typing — independent pipeline
+ // Typing
if tc, ok := c.owner.(TypingCapable); ok {
if stop, err := tc.StartTyping(ctx, chatID); err == nil {
c.placeholderRecorder.RecordTypingStop(c.name, chatID, stop)
}
}
- // Reaction — independent pipeline
+ // Reaction
if rc, ok := c.owner.(ReactionCapable); ok && messageID != "" {
if undo, err := rc.ReactToMessage(ctx, chatID, messageID); err == nil {
c.placeholderRecorder.RecordReactionUndo(c.name, chatID, undo)
diff --git a/pkg/channels/feishu/common.go b/pkg/channels/feishu/common.go
index fbe085b73..4952394b7 100644
--- a/pkg/channels/feishu/common.go
+++ b/pkg/channels/feishu/common.go
@@ -84,3 +84,64 @@ func stripMentionPlaceholders(content string, mentions []*larkim.MentionEvent) s
content = mentionPlaceholderRegex.ReplaceAllString(content, "")
return strings.TrimSpace(content)
}
+
+// extractCardImageKeys recursively extracts all image keys from a Feishu interactive card.
+// Image keys are used to download images from Feishu API.
+// Returns two slices: Feishu-hosted keys and external URLs.
+func extractCardImageKeys(rawContent string) (feishuKeys []string, externalURLs []string) {
+ if rawContent == "" {
+ return nil, nil
+ }
+
+ var card map[string]any
+ if err := json.Unmarshal([]byte(rawContent), &card); err != nil {
+ return nil, nil
+ }
+
+ extractImageKeysRecursive(card, &feishuKeys, &externalURLs)
+ return feishuKeys, externalURLs
+}
+
+// isExternalURL returns true if the string is an external HTTP/HTTPS URL.
+func isExternalURL(s string) bool {
+ return strings.HasPrefix(s, "http://") || strings.HasPrefix(s, "https://")
+}
+
+// extractImageKeysRecursive traverses card structure to find all image keys.
+// Collects both Feishu-hosted keys and external URLs separately.
+func extractImageKeysRecursive(v any, feishuKeys, externalURLs *[]string) {
+ switch val := v.(type) {
+ case map[string]any:
+ // Check if this is an img element
+ if tag, ok := val["tag"].(string); ok {
+ switch tag {
+ case "img":
+ // Try img_key first (always Feishu-hosted)
+ if imgKey, ok := val["img_key"].(string); ok && imgKey != "" {
+ *feishuKeys = append(*feishuKeys, imgKey)
+ }
+ // Check src - could be Feishu key or external URL
+ if src, ok := val["src"].(string); ok && src != "" {
+ if isExternalURL(src) {
+ *externalURLs = append(*externalURLs, src)
+ } else {
+ *feishuKeys = append(*feishuKeys, src)
+ }
+ }
+ case "icon":
+ // Icon elements use icon_key
+ if iconKey, ok := val["icon_key"].(string); ok && iconKey != "" {
+ *feishuKeys = append(*feishuKeys, iconKey)
+ }
+ }
+ }
+ // Recurse into all nested structures
+ for _, child := range val {
+ extractImageKeysRecursive(child, feishuKeys, externalURLs)
+ }
+ case []any:
+ for _, item := range val {
+ extractImageKeysRecursive(item, feishuKeys, externalURLs)
+ }
+ }
+}
diff --git a/pkg/channels/feishu/common_test.go b/pkg/channels/feishu/common_test.go
index fefc9f7c1..ff4af0148 100644
--- a/pkg/channels/feishu/common_test.go
+++ b/pkg/channels/feishu/common_test.go
@@ -290,3 +290,119 @@ func TestStripMentionPlaceholders(t *testing.T) {
})
}
}
+
+func TestExtractCardImageKeys(t *testing.T) {
+ tests := []struct {
+ name string
+ content string
+ wantFeishuKeys []string
+ wantExternalURLs []string
+ }{
+ {
+ name: "empty content",
+ content: "",
+ wantFeishuKeys: nil,
+ wantExternalURLs: nil,
+ },
+ {
+ name: "invalid JSON",
+ content: "not json",
+ wantFeishuKeys: nil,
+ wantExternalURLs: nil,
+ },
+ {
+ name: "card with no images",
+ content: `{"schema":"2.0","body":{"elements":[{"tag":"markdown","content":"text"}]}}`,
+ wantFeishuKeys: nil,
+ wantExternalURLs: nil,
+ },
+ {
+ name: "single image with img_key",
+ content: `{"elements":[{"tag":"img","img_key":"img_abc123"}]}`,
+ wantFeishuKeys: []string{"img_abc123"},
+ wantExternalURLs: nil,
+ },
+ {
+ name: "single image with src as Feishu key",
+ content: `{"elements":[{"tag":"img","src":"img_xyz789"}]}`,
+ wantFeishuKeys: []string{"img_xyz789"},
+ wantExternalURLs: nil,
+ },
+ {
+ name: "multiple images",
+ content: `{"elements":[{"tag":"img","img_key":"img_1"},{"tag":"div","text":{"content":"text"}},{"tag":"img","img_key":"img_2"}]}`,
+ wantFeishuKeys: []string{"img_1", "img_2"},
+ wantExternalURLs: nil,
+ },
+ {
+ name: "nested image in columns",
+ content: `{"elements":[{"tag":"div","columns":[{"tag":"img","img_key":"img_col1"},{"tag":"img","img_key":"img_col2"}]}]}`,
+ wantFeishuKeys: []string{"img_col1", "img_col2"},
+ wantExternalURLs: nil,
+ },
+ {
+ name: "image in action",
+ content: `{"elements":[{"tag":"action","actions":[{"tag":"img","img_key":"img_action"}]}]}`,
+ wantFeishuKeys: []string{"img_action"},
+ wantExternalURLs: nil,
+ },
+ {
+ name: "icon element",
+ content: `{"elements":[{"tag":"icon","icon_key":"icon_123"}]}`,
+ wantFeishuKeys: []string{"icon_123"},
+ wantExternalURLs: nil,
+ },
+ {
+ name: "complex card with text and images",
+ content: `{"header":{"title":{"content":"Title"}},"elements":[{"tag":"div","text":{"content":"Description"}},{"tag":"img","img_key":"img_main"}]}`,
+ wantFeishuKeys: []string{"img_main"},
+ wantExternalURLs: nil,
+ },
+ {
+ name: "external URL in src",
+ content: `{"elements":[{"tag":"img","src":"https://example.com/image.png"}]}`,
+ wantFeishuKeys: nil,
+ wantExternalURLs: []string{"https://example.com/image.png"},
+ },
+ {
+ name: "mixed Feishu keys and external URLs",
+ content: `{"elements":[{"tag":"img","img_key":"img_feishu"},{"tag":"img","src":"https://cdn.example.com/external.jpg"},{"tag":"img","src":"img_another"}]}`,
+ wantFeishuKeys: []string{"img_feishu", "img_another"},
+ wantExternalURLs: []string{"https://cdn.example.com/external.jpg"},
+ },
+ {
+ name: "multiple external URLs",
+ content: `{"elements":[{"tag":"img","src":"https://a.com/1.png"},{"tag":"img","src":"http://b.com/2.jpg"}]}`,
+ wantFeishuKeys: nil,
+ wantExternalURLs: []string{"https://a.com/1.png", "http://b.com/2.jpg"},
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ gotFeishuKeys, gotExternalURLs := extractCardImageKeys(tt.content)
+
+ // Compare Feishu keys
+ if len(gotFeishuKeys) != len(tt.wantFeishuKeys) {
+ t.Errorf("extractCardImageKeys() feishuKeys = %v, want %v", gotFeishuKeys, tt.wantFeishuKeys)
+ return
+ }
+ for i, v := range gotFeishuKeys {
+ if v != tt.wantFeishuKeys[i] {
+ t.Errorf("extractCardImageKeys() feishuKeys[%d] = %q, want %q", i, v, tt.wantFeishuKeys[i])
+ }
+ }
+
+ // Compare external URLs
+ if len(gotExternalURLs) != len(tt.wantExternalURLs) {
+ t.Errorf("extractCardImageKeys() externalURLs = %v, want %v", gotExternalURLs, tt.wantExternalURLs)
+ return
+ }
+ for i, v := range gotExternalURLs {
+ if v != tt.wantExternalURLs[i] {
+ t.Errorf("extractCardImageKeys() externalURLs[%d] = %q, want %q", i, v, tt.wantExternalURLs[i])
+ }
+ }
+ })
+ }
+}
diff --git a/pkg/channels/feishu/feishu_64.go b/pkg/channels/feishu/feishu_64.go
index 5dbbcf0af..37a74718a 100644
--- a/pkg/channels/feishu/feishu_64.go
+++ b/pkg/channels/feishu/feishu_64.go
@@ -11,6 +11,7 @@ import (
"net/http"
"os"
"path/filepath"
+ "strings"
"sync"
"sync/atomic"
@@ -29,11 +30,17 @@ import (
"github.com/sipeed/picoclaw/pkg/utils"
)
+// errCodeTenantTokenInvalid is the Feishu API error code for an expired/revoked
+// tenant_access_token. The Lark SDK's built-in retry does not clear its cache
+// on this error, so we do it ourselves.
+const errCodeTenantTokenInvalid = 99991663
+
type FeishuChannel struct {
*channels.BaseChannel
- config config.FeishuConfig
- client *lark.Client
- wsClient *larkws.Client
+ config config.FeishuConfig
+ client *lark.Client
+ wsClient *larkws.Client
+ tokenCache *tokenCache // custom cache that supports invalidation
botOpenID atomic.Value // stores string; populated lazily for @mention detection
@@ -47,10 +54,16 @@ func NewFeishuChannel(cfg config.FeishuConfig, bus *bus.MessageBus) (*FeishuChan
channels.WithReasoningChannelID(cfg.ReasoningChannelID),
)
+ tc := newTokenCache()
+ opts := []lark.ClientOptionFunc{lark.WithTokenCache(tc)}
+ if cfg.IsLark {
+ opts = append(opts, lark.WithOpenBaseUrl(lark.LarkBaseUrl))
+ }
ch := &FeishuChannel{
BaseChannel: base,
config: cfg,
- client: lark.NewClient(cfg.AppID, cfg.AppSecret),
+ tokenCache: tc,
+ client: lark.NewClient(cfg.AppID, cfg.AppSecret, opts...),
}
ch.SetOwner(ch)
return ch, nil
@@ -75,10 +88,15 @@ func (c *FeishuChannel) Start(ctx context.Context) error {
c.mu.Lock()
c.cancel = cancel
+ domain := lark.FeishuBaseUrl
+ if c.config.IsLark {
+ domain = lark.LarkBaseUrl
+ }
c.wsClient = larkws.NewClient(
c.config.AppID,
c.config.AppSecret,
larkws.WithEventHandler(dispatcher),
+ larkws.WithDomain(domain),
)
wsClient := c.wsClient
c.mu.Unlock()
@@ -112,6 +130,7 @@ func (c *FeishuChannel) Stop(ctx context.Context) error {
}
// Send sends a message using Interactive Card format for markdown rendering.
+// Falls back to plain text message if card sending fails (e.g., table limit exceeded).
func (c *FeishuChannel) Send(ctx context.Context, msg bus.OutboundMessage) error {
if !c.IsRunning() {
return channels.ErrNotRunning
@@ -124,9 +143,38 @@ func (c *FeishuChannel) Send(ctx context.Context, msg bus.OutboundMessage) error
// Build interactive card with markdown content
cardContent, err := buildMarkdownCard(msg.Content)
if err != nil {
- return fmt.Errorf("feishu send: card build failed: %w", err)
+ // If card build fails, fall back to plain text
+ return c.sendText(ctx, msg.ChatID, msg.Content)
}
- return c.sendCard(ctx, msg.ChatID, cardContent)
+
+ // First attempt: try sending as interactive card
+ err = c.sendCard(ctx, msg.ChatID, cardContent)
+ if err == nil {
+ return nil
+ }
+
+ // Check if error is due to card table limit (error code 11310)
+ // See: https://open.feishu.cn/document/server-docs/im-api/message-content-description/create_json
+ errMsg := err.Error()
+ isCardLimitError := strings.Contains(errMsg, "11310")
+
+ if isCardLimitError {
+ logger.WarnCF("feishu", "Card send failed (table limit), falling back to text message", map[string]any{
+ "chat_id": msg.ChatID,
+ "error": errMsg,
+ })
+
+ // Second attempt: fall back to plain text message
+ textErr := c.sendText(ctx, msg.ChatID, msg.Content)
+ if textErr == nil {
+ return nil
+ }
+ // If text also fails, return the text error
+ return textErr
+ }
+
+ // For other errors, return the original card error
+ return err
}
// EditMessage implements channels.MessageEditor.
@@ -147,6 +195,7 @@ func (c *FeishuChannel) EditMessage(ctx context.Context, chatID, messageID, cont
return fmt.Errorf("feishu edit: %w", err)
}
if !resp.Success() {
+ c.invalidateTokenOnAuthError(resp.Code)
return fmt.Errorf("feishu edit api error (code=%d msg=%s)", resp.Code, resp.Msg)
}
return nil
@@ -186,6 +235,7 @@ func (c *FeishuChannel) SendPlaceholder(ctx context.Context, chatID string) (str
return "", fmt.Errorf("feishu placeholder send: %w", err)
}
if !resp.Success() {
+ c.invalidateTokenOnAuthError(resp.Code)
return "", fmt.Errorf("feishu placeholder api error (code=%d msg=%s)", resp.Code, resp.Msg)
}
@@ -226,6 +276,7 @@ func (c *FeishuChannel) ReactToMessage(ctx context.Context, chatID, messageID st
return func() {}, fmt.Errorf("feishu react: %w", err)
}
if !resp.Success() {
+ c.invalidateTokenOnAuthError(resp.Code)
logger.ErrorCF("feishu", "Reaction API error", map[string]any{
"emoji": chosenEmoji,
"message_id": messageID,
@@ -373,6 +424,15 @@ func (c *FeishuChannel) handleMessageReceive(ctx context.Context, event *larkim.
mediaRefs = c.downloadInboundMedia(ctx, chatID, messageID, messageType, rawContent, store)
}
+ // For interactive cards, pass external image URLs via media refs.
+ // Keep content as valid raw JSON for downstream parsing.
+ if messageType == larkim.MsgTypeInteractive {
+ _, externalURLs := extractCardImageKeys(rawContent)
+ if len(externalURLs) > 0 {
+ mediaRefs = append(mediaRefs, externalURLs...)
+ }
+ }
+
// Append media tags to content (like Telegram does)
content = appendMediaTags(content, messageType, mediaRefs)
@@ -451,6 +511,7 @@ func (c *FeishuChannel) fetchBotOpenID(ctx context.Context) error {
return fmt.Errorf("bot info parse: %w", err)
}
if result.Code != 0 {
+ c.invalidateTokenOnAuthError(result.Code)
return fmt.Errorf("bot info api error (code=%d)", result.Code)
}
if result.Bot.OpenID == "" {
@@ -507,6 +568,10 @@ func extractContent(messageType, rawContent string) string {
// Pass raw JSON to LLM — structured rich text is more informative than flattened plain text
return rawContent
+ case larkim.MsgTypeInteractive:
+ // Pass raw JSON to LLM — structured card is more informative than flattened text
+ return rawContent
+
case larkim.MsgTypeImage:
// Image messages don't have text content
return ""
@@ -544,6 +609,18 @@ func (c *FeishuChannel) downloadInboundMedia(
refs = append(refs, ref)
}
+ case larkim.MsgTypeInteractive:
+ // Extract and download images embedded in interactive cards
+ feishuKeys, _ := extractCardImageKeys(rawContent)
+ // Download Feishu-hosted images via API
+ for _, imageKey := range feishuKeys {
+ ref := c.downloadResource(ctx, messageID, imageKey, "image", ".jpg", store, scope)
+ if ref != "" {
+ refs = append(refs, ref)
+ }
+ }
+ // External URLs are passed directly to LLM, not downloaded
+
case larkim.MsgTypeFile, larkim.MsgTypeAudio, larkim.MsgTypeMedia:
fileKey := extractFileKey(rawContent)
if fileKey == "" {
@@ -593,6 +670,7 @@ func (c *FeishuChannel) downloadResource(
return ""
}
if !resp.Success() {
+ c.invalidateTokenOnAuthError(resp.Code)
logger.ErrorCF("feishu", "Resource download api error", map[string]any{
"code": resp.Code,
"msg": resp.Msg,
@@ -618,7 +696,7 @@ func (c *FeishuChannel) downloadResource(
}
// Write to the shared picoclaw_media directory using a unique name to avoid collisions.
- mediaDir := filepath.Join(os.TempDir(), "picoclaw_media")
+ mediaDir := media.TempDir()
if mkdirErr := os.MkdirAll(mediaDir, 0o700); mkdirErr != nil {
logger.ErrorCF("feishu", "Failed to create media directory", map[string]any{
"error": mkdirErr.Error(),
@@ -663,11 +741,18 @@ func (c *FeishuChannel) downloadResource(
}
// appendMediaTags appends media type tags to content (like Telegram's "[image: photo]").
+// For interactive cards, media tags are not appended because content is raw JSON
+// and appending would produce invalid JSON format.
func appendMediaTags(content, messageType string, mediaRefs []string) string {
if len(mediaRefs) == 0 {
return content
}
+ // Don't append tags to JSON content (interactive cards) - would produce invalid JSON
+ if messageType == larkim.MsgTypeInteractive {
+ return content
+ }
+
var tag string
switch messageType {
case larkim.MsgTypeImage:
@@ -705,6 +790,7 @@ func (c *FeishuChannel) sendCard(ctx context.Context, chatID, cardContent string
}
if !resp.Success() {
+ c.invalidateTokenOnAuthError(resp.Code)
return fmt.Errorf("feishu api error (code=%d msg=%s): %w", resp.Code, resp.Msg, channels.ErrTemporary)
}
@@ -715,6 +801,35 @@ func (c *FeishuChannel) sendCard(ctx context.Context, chatID, cardContent string
return nil
}
+// sendText sends a plain text message to a chat (fallback when card fails).
+func (c *FeishuChannel) sendText(ctx context.Context, chatID, text string) error {
+ content, _ := json.Marshal(map[string]string{"text": text})
+
+ req := larkim.NewCreateMessageReqBuilder().
+ ReceiveIdType(larkim.ReceiveIdTypeChatId).
+ Body(larkim.NewCreateMessageReqBodyBuilder().
+ ReceiveId(chatID).
+ MsgType(larkim.MsgTypeText).
+ Content(string(content)).
+ Build()).
+ Build()
+
+ resp, err := c.client.Im.V1.Message.Create(ctx, req)
+ if err != nil {
+ return fmt.Errorf("feishu send text: %w", channels.ErrTemporary)
+ }
+
+ if !resp.Success() {
+ return fmt.Errorf("feishu text api error (code=%d msg=%s): %w", resp.Code, resp.Msg, channels.ErrTemporary)
+ }
+
+ logger.DebugCF("feishu", "Feishu text message sent (fallback)", map[string]any{
+ "chat_id": chatID,
+ })
+
+ return nil
+}
+
// sendImage uploads an image and sends it as a message.
func (c *FeishuChannel) sendImage(ctx context.Context, chatID string, file *os.File) error {
// Upload image to get image_key
@@ -730,6 +845,7 @@ func (c *FeishuChannel) sendImage(ctx context.Context, chatID string, file *os.F
return fmt.Errorf("feishu image upload: %w", err)
}
if !uploadResp.Success() {
+ c.invalidateTokenOnAuthError(uploadResp.Code)
return fmt.Errorf("feishu image upload api error (code=%d msg=%s)", uploadResp.Code, uploadResp.Msg)
}
if uploadResp.Data == nil || uploadResp.Data.ImageKey == nil {
@@ -754,6 +870,7 @@ func (c *FeishuChannel) sendImage(ctx context.Context, chatID string, file *os.F
return fmt.Errorf("feishu image send: %w", err)
}
if !resp.Success() {
+ c.invalidateTokenOnAuthError(resp.Code)
return fmt.Errorf("feishu image send api error (code=%d msg=%s)", resp.Code, resp.Msg)
}
return nil
@@ -784,6 +901,7 @@ func (c *FeishuChannel) sendFile(ctx context.Context, chatID string, file *os.Fi
return fmt.Errorf("feishu file upload: %w", err)
}
if !uploadResp.Success() {
+ c.invalidateTokenOnAuthError(uploadResp.Code)
return fmt.Errorf("feishu file upload api error (code=%d msg=%s)", uploadResp.Code, uploadResp.Msg)
}
if uploadResp.Data == nil || uploadResp.Data.FileKey == nil {
@@ -808,6 +926,7 @@ func (c *FeishuChannel) sendFile(ctx context.Context, chatID string, file *os.Fi
return fmt.Errorf("feishu file send: %w", err)
}
if !resp.Success() {
+ c.invalidateTokenOnAuthError(resp.Code)
return fmt.Errorf("feishu file send api error (code=%d msg=%s)", resp.Code, resp.Msg)
}
return nil
@@ -830,3 +949,14 @@ func extractFeishuSenderID(sender *larkim.EventSender) string {
return ""
}
+
+// invalidateTokenOnAuthError clears the cached tenant_access_token when the
+// Feishu API reports it as invalid (99991663), so the next request fetches a
+// fresh one. The Lark SDK's built-in retry does not clear the cache, causing
+// all API calls to fail until the token naturally expires (~2 hours).
+func (c *FeishuChannel) invalidateTokenOnAuthError(code int) {
+ if code == errCodeTenantTokenInvalid {
+ c.tokenCache.InvalidateAll()
+ logger.WarnCF("feishu", "Invalidated cached token due to auth error", nil)
+ }
+}
diff --git a/pkg/channels/feishu/feishu_64_test.go b/pkg/channels/feishu/feishu_64_test.go
index dc3eab2e7..9010abf69 100644
--- a/pkg/channels/feishu/feishu_64_test.go
+++ b/pkg/channels/feishu/feishu_64_test.go
@@ -75,6 +75,24 @@ func TestExtractContent(t *testing.T) {
rawContent: "",
want: "",
},
+ {
+ name: "interactive card returns raw JSON",
+ messageType: "interactive",
+ rawContent: `{"schema":"2.0","body":{"elements":[{"tag":"markdown","content":"Hello from card"}]}}`,
+ want: `{"schema":"2.0","body":{"elements":[{"tag":"markdown","content":"Hello from card"}]}}`,
+ },
+ {
+ name: "interactive card with complex structure returns raw JSON",
+ messageType: "interactive",
+ rawContent: `{"header":{"title":{"tag":"plain_text","content":"Title"}},"elements":[{"tag":"div","text":{"tag":"lark_md","content":"Card content"}}]}`,
+ want: `{"header":{"title":{"tag":"plain_text","content":"Title"}},"elements":[{"tag":"div","text":{"tag":"lark_md","content":"Card content"}}]}`,
+ },
+ {
+ name: "interactive card invalid JSON returns as-is",
+ messageType: "interactive",
+ rawContent: `not valid json`,
+ want: `not valid json`,
+ },
}
for _, tt := range tests {
@@ -151,6 +169,13 @@ func TestAppendMediaTags(t *testing.T) {
mediaRefs: []string{"ref1"},
want: "something [attachment]",
},
+ {
+ name: "interactive card with images returns content unchanged",
+ content: `{"schema":"2.0","body":{"elements":[{"tag":"img","img_key":"img_123"}]}}`,
+ messageType: "interactive",
+ mediaRefs: []string{"ref1"},
+ want: `{"schema":"2.0","body":{"elements":[{"tag":"img","img_key":"img_123"}]}}`,
+ },
}
for _, tt := range tests {
diff --git a/pkg/channels/feishu/token_cache.go b/pkg/channels/feishu/token_cache.go
new file mode 100644
index 000000000..00acbc084
--- /dev/null
+++ b/pkg/channels/feishu/token_cache.go
@@ -0,0 +1,52 @@
+package feishu
+
+import (
+ "context"
+ "sync"
+ "time"
+)
+
+// tokenCache implements larkcore.Cache with an extra InvalidateAll method.
+// This works around a bug in the Lark SDK v3 where the built-in token retry
+// loop does not clear stale tokens from cache on auth errors.
+type tokenCache struct {
+ mu sync.RWMutex
+ store map[string]*tokenEntry
+}
+
+type tokenEntry struct {
+ value string
+ expireAt time.Time
+}
+
+func newTokenCache() *tokenCache {
+ return &tokenCache{store: make(map[string]*tokenEntry)}
+}
+
+func (c *tokenCache) Set(_ context.Context, key, value string, ttl time.Duration) error {
+ c.mu.Lock()
+ defer c.mu.Unlock()
+ c.store[key] = &tokenEntry{value: value, expireAt: time.Now().Add(ttl)}
+ return nil
+}
+
+func (c *tokenCache) Get(_ context.Context, key string) (string, error) {
+ c.mu.Lock()
+ defer c.mu.Unlock()
+ e, ok := c.store[key]
+ if !ok {
+ return "", nil
+ }
+ if e.expireAt.Before(time.Now()) {
+ delete(c.store, key)
+ return "", nil
+ }
+ return e.value, nil
+}
+
+// InvalidateAll removes all cached tokens, forcing fresh acquisition.
+func (c *tokenCache) InvalidateAll() {
+ c.mu.Lock()
+ defer c.mu.Unlock()
+ clear(c.store)
+}
diff --git a/pkg/channels/interfaces.go b/pkg/channels/interfaces.go
index b3a493761..0cfd435b0 100644
--- a/pkg/channels/interfaces.go
+++ b/pkg/channels/interfaces.go
@@ -3,6 +3,7 @@ package channels
import (
"context"
+ "github.com/sipeed/picoclaw/pkg/bus"
"github.com/sipeed/picoclaw/pkg/commands"
)
@@ -19,6 +20,11 @@ type MessageEditor interface {
EditMessage(ctx context.Context, chatID string, messageID string, content string) error
}
+// MessageDeleter — channels that can delete a message by ID.
+type MessageDeleter interface {
+ DeleteMessage(ctx context.Context, chatID string, messageID string) error
+}
+
// ReactionCapable — channels that can add a reaction (e.g. 👀) to an inbound message.
// ReactToMessage adds a reaction and returns an undo function to remove it.
// The undo function MUST be idempotent and safe to call multiple times.
@@ -35,6 +41,18 @@ type PlaceholderCapable interface {
SendPlaceholder(ctx context.Context, chatID string) (messageID string, err error)
}
+// StreamingCapable — channels that can show partial LLM output in real-time.
+// The channel SHOULD gracefully degrade if the platform rejects streaming
+// (e.g. Telegram bot without forum mode). In that case, Update becomes a no-op
+// and Finalize still delivers the final message.
+type StreamingCapable interface {
+ BeginStream(ctx context.Context, chatID string) (Streamer, error)
+}
+
+// Streamer is defined in pkg/bus to avoid circular imports.
+// This alias keeps channel implementations using channels.Streamer unchanged.
+type Streamer = bus.Streamer
+
// PlaceholderRecorder is injected into channels by Manager.
// Channels call these methods on inbound to register typing/placeholder state.
// Manager uses the registered state on outbound to stop typing and edit placeholders.
diff --git a/pkg/channels/manager.go b/pkg/channels/manager.go
index df430e4d3..ff3fa399c 100644
--- a/pkg/channels/manager.go
+++ b/pkg/channels/manager.go
@@ -86,9 +86,11 @@ type Manager struct {
mux *http.ServeMux
httpServer *http.Server
mu sync.RWMutex
- placeholders sync.Map // "channel:chatID" → placeholderID (string)
- typingStops sync.Map // "channel:chatID" → func()
- reactionUndos sync.Map // "channel:chatID" → reactionEntry
+ placeholders sync.Map // "channel:chatID" → placeholderID (string)
+ typingStops sync.Map // "channel:chatID" → func()
+ reactionUndos sync.Map // "channel:chatID" → reactionEntry
+ streamActive sync.Map // "channel:chatID" → true (set when streamer.Finalize sent the message)
+ channelHashes map[string]string // channel name → config hash
}
type asyncTask struct {
@@ -135,6 +137,19 @@ func (m *Manager) RecordTypingStop(channel, chatID string, stop func()) {
}
}
+// InvokeTypingStop invokes the registered typing stop function for the given channel and chatID.
+// It is safe to call even when no typing indicator is active (no-op).
+// Used by the agent loop to stop typing when processing completes (success, error, or panic),
+// regardless of whether an outbound message is published.
+func (m *Manager) InvokeTypingStop(channel, chatID string) {
+ key := channel + ":" + chatID
+ if v, loaded := m.typingStops.LoadAndDelete(key); loaded {
+ if entry, ok := v.(typingEntry); ok {
+ entry.stop()
+ }
+ }
+}
+
// RecordReactionUndo registers a reaction undo function for later invocation.
// Implements PlaceholderRecorder.
func (m *Manager) RecordReactionUndo(channel, chatID string, undo func()) {
@@ -143,7 +158,7 @@ func (m *Manager) RecordReactionUndo(channel, chatID string, undo func()) {
}
// preSend handles typing stop, reaction undo, and placeholder editing before sending a message.
-// Returns true if the message was edited into a placeholder (skip Send).
+// Returns true if the message was already delivered (skip Send).
func (m *Manager) preSend(ctx context.Context, name string, msg bus.OutboundMessage, ch Channel) bool {
key := name + ":" + msg.ChatID
@@ -161,7 +176,22 @@ func (m *Manager) preSend(ctx context.Context, name string, msg bus.OutboundMess
}
}
- // 3. Try editing placeholder
+ // 3. If a stream already finalized this message, delete the placeholder and skip send
+ if _, loaded := m.streamActive.LoadAndDelete(key); loaded {
+ if v, loaded := m.placeholders.LoadAndDelete(key); loaded {
+ if entry, ok := v.(placeholderEntry); ok && entry.id != "" {
+ // Prefer deleting the placeholder (cleaner UX than editing to same content)
+ if deleter, ok := ch.(MessageDeleter); ok {
+ deleter.DeleteMessage(ctx, msg.ChatID, entry.id) // best effort
+ } else if editor, ok := ch.(MessageEditor); ok {
+ editor.EditMessage(ctx, msg.ChatID, entry.id, msg.Content) // fallback
+ }
+ }
+ }
+ return true
+ }
+
+ // 4. Try editing placeholder
if v, loaded := m.placeholders.LoadAndDelete(key); loaded {
if entry, ok := v.(placeholderEntry); ok && entry.id != "" {
if editor, ok := ch.(MessageEditor); ok {
@@ -178,20 +208,74 @@ func (m *Manager) preSend(ctx context.Context, name string, msg bus.OutboundMess
func NewManager(cfg *config.Config, messageBus *bus.MessageBus, store media.MediaStore) (*Manager, error) {
m := &Manager{
- channels: make(map[string]Channel),
- workers: make(map[string]*channelWorker),
- bus: messageBus,
- config: cfg,
- mediaStore: store,
+ channels: make(map[string]Channel),
+ workers: make(map[string]*channelWorker),
+ bus: messageBus,
+ config: cfg,
+ mediaStore: store,
+ channelHashes: make(map[string]string),
}
- if err := m.initChannels(); err != nil {
+ // Register as streaming delegate so the agent loop can obtain streamers
+ messageBus.SetStreamDelegate(m)
+
+ if err := m.initChannels(&cfg.Channels); err != nil {
return nil, err
}
+ // Store initial config hashes for all channels
+ m.channelHashes = toChannelHashes(cfg)
+
return m, nil
}
+// GetStreamer implements bus.StreamDelegate.
+// It checks if the named channel supports streaming and returns a Streamer.
+func (m *Manager) GetStreamer(ctx context.Context, channelName, chatID string) (bus.Streamer, bool) {
+ m.mu.RLock()
+ ch, exists := m.channels[channelName]
+ m.mu.RUnlock()
+
+ if !exists {
+ return nil, false
+ }
+
+ sc, ok := ch.(StreamingCapable)
+ if !ok {
+ return nil, false
+ }
+
+ streamer, err := sc.BeginStream(ctx, chatID)
+ if err != nil {
+ logger.DebugCF("channels", "Streaming unavailable, falling back to placeholder", map[string]any{
+ "channel": channelName,
+ "error": err.Error(),
+ })
+ return nil, false
+ }
+
+ // Mark streamActive on Finalize so preSend knows to clean up the placeholder
+ key := channelName + ":" + chatID
+ return &finalizeHookStreamer{
+ Streamer: streamer,
+ onFinalize: func() { m.streamActive.Store(key, true) },
+ }, true
+}
+
+// finalizeHookStreamer wraps a Streamer to run a hook on Finalize.
+type finalizeHookStreamer struct {
+ Streamer
+ onFinalize func()
+}
+
+func (s *finalizeHookStreamer) Finalize(ctx context.Context, content string) error {
+ if err := s.Streamer.Finalize(ctx, content); err != nil {
+ return err
+ }
+ s.onFinalize()
+ return nil
+}
+
// initChannel is a helper that looks up a factory by name and creates the channel.
func (m *Manager) initChannel(name, displayName string) {
f, ok := getFactory(name)
@@ -232,15 +316,15 @@ func (m *Manager) initChannel(name, displayName string) {
}
}
-func (m *Manager) initChannels() error {
+func (m *Manager) initChannels(channels *config.ChannelsConfig) error {
logger.InfoC("channels", "Initializing channel manager")
- if m.config.Channels.Telegram.Enabled && m.config.Channels.Telegram.Token != "" {
+ if channels.Telegram.Enabled && channels.Telegram.Token != "" {
m.initChannel("telegram", "Telegram")
}
- if m.config.Channels.WhatsApp.Enabled {
- waCfg := m.config.Channels.WhatsApp
+ if channels.WhatsApp.Enabled {
+ waCfg := channels.WhatsApp
if waCfg.UseNative {
m.initChannel("whatsapp_native", "WhatsApp Native")
} else if waCfg.BridgeURL != "" {
@@ -248,62 +332,68 @@ func (m *Manager) initChannels() error {
}
}
- if m.config.Channels.Feishu.Enabled {
+ if channels.Feishu.Enabled {
m.initChannel("feishu", "Feishu")
}
- if m.config.Channels.Discord.Enabled && m.config.Channels.Discord.Token != "" {
+ if channels.Discord.Enabled && channels.Discord.Token != "" {
m.initChannel("discord", "Discord")
}
- if m.config.Channels.MaixCam.Enabled {
+ if channels.MaixCam.Enabled {
m.initChannel("maixcam", "MaixCam")
}
- if m.config.Channels.QQ.Enabled {
+ if channels.QQ.Enabled {
m.initChannel("qq", "QQ")
}
- if m.config.Channels.DingTalk.Enabled && m.config.Channels.DingTalk.ClientID != "" {
+ if channels.DingTalk.Enabled && channels.DingTalk.ClientID != "" {
m.initChannel("dingtalk", "DingTalk")
}
- if m.config.Channels.Slack.Enabled && m.config.Channels.Slack.BotToken != "" {
+ if channels.Slack.Enabled && channels.Slack.BotToken != "" {
m.initChannel("slack", "Slack")
}
- if m.config.Channels.Matrix.Enabled &&
+ if channels.Matrix.Enabled &&
m.config.Channels.Matrix.Homeserver != "" &&
m.config.Channels.Matrix.UserID != "" &&
m.config.Channels.Matrix.AccessToken != "" {
m.initChannel("matrix", "Matrix")
}
- if m.config.Channels.LINE.Enabled && m.config.Channels.LINE.ChannelAccessToken != "" {
+ if channels.LINE.Enabled && channels.LINE.ChannelAccessToken != "" {
m.initChannel("line", "LINE")
}
- if m.config.Channels.OneBot.Enabled && m.config.Channels.OneBot.WSUrl != "" {
+ if channels.OneBot.Enabled && channels.OneBot.WSUrl != "" {
m.initChannel("onebot", "OneBot")
}
- if m.config.Channels.WeCom.Enabled && m.config.Channels.WeCom.Token != "" {
+ if channels.WeCom.Enabled && channels.WeCom.Token != "" {
m.initChannel("wecom", "WeCom")
}
- if m.config.Channels.WeComAIBot.Enabled && m.config.Channels.WeComAIBot.Token != "" {
+ if m.config.Channels.WeComAIBot.Enabled &&
+ ((m.config.Channels.WeComAIBot.BotID != "" && m.config.Channels.WeComAIBot.Secret != "") ||
+ m.config.Channels.WeComAIBot.Token != "") {
m.initChannel("wecom_aibot", "WeCom AI Bot")
}
- if m.config.Channels.WeComApp.Enabled && m.config.Channels.WeComApp.CorpID != "" {
+ if channels.WeComApp.Enabled && channels.WeComApp.CorpID != "" {
m.initChannel("wecom_app", "WeCom App")
}
- if m.config.Channels.Pico.Enabled && m.config.Channels.Pico.Token != "" {
+ if channels.Pico.Enabled && channels.Pico.Token != "" {
m.initChannel("pico", "Pico")
}
- if m.config.Channels.IRC.Enabled && m.config.Channels.IRC.Server != "" {
+ if channels.PicoClient.Enabled && channels.PicoClient.URL != "" {
+ m.initChannel("pico_client", "Pico Client")
+ }
+
+ if channels.IRC.Enabled && channels.IRC.Server != "" {
m.initChannel("irc", "IRC")
}
@@ -357,7 +447,6 @@ func (m *Manager) StartAll(ctx context.Context) error {
if len(m.channels) == 0 {
logger.WarnC("channels", "No channels enabled")
- return errors.New("no channels enabled")
}
logger.InfoC("channels", "Starting all channels")
@@ -397,7 +486,7 @@ func (m *Manager) StartAll(ctx context.Context) error {
"addr": m.httpServer.Addr,
})
if err := m.httpServer.ListenAndServe(); err != nil && err != http.ErrServerClosed {
- logger.ErrorCF("channels", "Shared HTTP server error", map[string]any{
+ logger.FatalCF("channels", "Shared HTTP server error", map[string]any{
"error": err.Error(),
})
}
@@ -586,7 +675,7 @@ func (m *Manager) sendWithRetry(ctx context.Context, name string, w *channelWork
func dispatchLoop[M any](
ctx context.Context,
m *Manager,
- subscribe func(context.Context) (M, bool),
+ ch <-chan M,
getChannel func(M) string,
enqueue func(context.Context, *channelWorker, M) bool,
startMsg, stopMsg, unknownMsg, noWorkerMsg string,
@@ -594,35 +683,41 @@ func dispatchLoop[M any](
logger.InfoC("channels", startMsg)
for {
- msg, ok := subscribe(ctx)
- if !ok {
+ select {
+ case <-ctx.Done():
logger.InfoC("channels", stopMsg)
return
- }
- channel := getChannel(msg)
-
- // Silently skip internal channels
- if constants.IsInternalChannel(channel) {
- continue
- }
-
- m.mu.RLock()
- _, exists := m.channels[channel]
- w, wExists := m.workers[channel]
- m.mu.RUnlock()
-
- if !exists {
- logger.WarnCF("channels", unknownMsg, map[string]any{"channel": channel})
- continue
- }
-
- if wExists && w != nil {
- if !enqueue(ctx, w, msg) {
+ case msg, ok := <-ch:
+ if !ok {
+ logger.InfoC("channels", stopMsg)
return
}
- } else if exists {
- logger.WarnCF("channels", noWorkerMsg, map[string]any{"channel": channel})
+
+ channel := getChannel(msg)
+
+ // Silently skip internal channels
+ if constants.IsInternalChannel(channel) {
+ continue
+ }
+
+ m.mu.RLock()
+ _, exists := m.channels[channel]
+ w, wExists := m.workers[channel]
+ m.mu.RUnlock()
+
+ if !exists {
+ logger.WarnCF("channels", unknownMsg, map[string]any{"channel": channel})
+ continue
+ }
+
+ if wExists && w != nil {
+ if !enqueue(ctx, w, msg) {
+ return
+ }
+ } else if exists {
+ logger.WarnCF("channels", noWorkerMsg, map[string]any{"channel": channel})
+ }
}
}
}
@@ -630,7 +725,7 @@ func dispatchLoop[M any](
func (m *Manager) dispatchOutbound(ctx context.Context) {
dispatchLoop(
ctx, m,
- m.bus.SubscribeOutbound,
+ m.bus.OutboundChan(),
func(msg bus.OutboundMessage) string { return msg.Channel },
func(ctx context.Context, w *channelWorker, msg bus.OutboundMessage) bool {
select {
@@ -650,7 +745,7 @@ func (m *Manager) dispatchOutbound(ctx context.Context) {
func (m *Manager) dispatchOutboundMedia(ctx context.Context) {
dispatchLoop(
ctx, m,
- m.bus.SubscribeOutboundMedia,
+ m.bus.OutboundMediaChan(),
func(msg bus.OutboundMediaMessage) string { return msg.Channel },
func(ctx context.Context, w *channelWorker, msg bus.OutboundMediaMessage) bool {
select {
@@ -820,6 +915,68 @@ func (m *Manager) GetEnabledChannels() []string {
return names
}
+// Reload updates the config reference without restarting channels.
+// This is used when channel config hasn't changed but other parts of the config have.
+func (m *Manager) Reload(ctx context.Context, cfg *config.Config) error {
+ m.mu.Lock()
+ defer m.mu.Unlock()
+ list := toChannelHashes(cfg)
+ added, removed := compareChannels(m.channelHashes, list)
+ for _, name := range removed {
+ // Stop all channels
+ channel := m.channels[name]
+ logger.InfoCF("channels", "Stopping channel", map[string]any{
+ "channel": name,
+ })
+ if err := channel.Stop(ctx); err != nil {
+ logger.ErrorCF("channels", "Error stopping channel", map[string]any{
+ "channel": name,
+ "error": err.Error(),
+ })
+ }
+ go func() {
+ m.UnregisterChannel(name)
+ }()
+ }
+ dispatchCtx, cancel := context.WithCancel(ctx)
+ m.dispatchTask = &asyncTask{cancel: cancel}
+ cc, err := toChannelConfig(cfg, added)
+ if err != nil {
+ logger.ErrorC("channels", fmt.Sprintf("toChannelConfig error: %v", err))
+ return err
+ }
+ err = m.initChannels(cc)
+ if err != nil {
+ logger.ErrorC("channels", fmt.Sprintf("initChannels error: %v", err))
+ return err
+ }
+ for _, name := range added {
+ channel := m.channels[name]
+ logger.InfoCF("channels", "Starting channel", map[string]any{
+ "channel": name,
+ })
+ if err := channel.Start(ctx); err != nil {
+ logger.ErrorCF("channels", "Failed to start channel", map[string]any{
+ "channel": name,
+ "error": err.Error(),
+ })
+ continue
+ }
+ // Lazily create worker only after channel starts successfully
+ w := newChannelWorker(name, channel)
+ m.workers[name] = w
+ go m.runWorker(dispatchCtx, name, w)
+ go m.runMediaWorker(dispatchCtx, name, w)
+ go func() {
+ m.RegisterChannel(name, channel)
+ }()
+ }
+
+ m.config = cfg
+ m.channelHashes = toChannelHashes(cfg)
+ return nil
+}
+
func (m *Manager) RegisterChannel(name string, channel Channel) {
m.mu.Lock()
defer m.mu.Unlock()
diff --git a/pkg/channels/manager_channel.go b/pkg/channels/manager_channel.go
new file mode 100644
index 000000000..57cb05412
--- /dev/null
+++ b/pkg/channels/manager_channel.go
@@ -0,0 +1,86 @@
+package channels
+
+import (
+ "crypto/md5"
+ "encoding/hex"
+ "encoding/json"
+
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/logger"
+)
+
+func toChannelHashes(cfg *config.Config) map[string]string {
+ result := make(map[string]string)
+ ch := cfg.Channels
+ // should not be error
+ marshal, _ := json.Marshal(ch)
+ var channelConfig map[string]map[string]any
+ _ = json.Unmarshal(marshal, &channelConfig)
+
+ for key, value := range channelConfig {
+ if !value["enabled"].(bool) {
+ continue
+ }
+ valueBytes, _ := json.Marshal(value)
+ hash := md5.Sum(valueBytes)
+ result[key] = hex.EncodeToString(hash[:])
+ }
+
+ return result
+}
+
+func compareChannels(old, news map[string]string) (added, removed []string) {
+ for key, newHash := range news {
+ if oldHash, ok := old[key]; ok {
+ if newHash != oldHash {
+ removed = append(removed, key)
+ added = append(added, key)
+ }
+ } else {
+ added = append(added, key)
+ }
+ }
+ for key := range old {
+ if _, ok := news[key]; !ok {
+ removed = append(removed, key)
+ }
+ }
+ return added, removed
+}
+
+func toChannelConfig(cfg *config.Config, list []string) (*config.ChannelsConfig, error) {
+ result := &config.ChannelsConfig{}
+ ch := cfg.Channels
+ // should not be error
+ marshal, _ := json.Marshal(ch)
+ var channelConfig map[string]map[string]any
+ _ = json.Unmarshal(marshal, &channelConfig)
+ temp := make(map[string]map[string]any, 0)
+
+ for key, value := range channelConfig {
+ found := false
+ for _, s := range list {
+ if key == s {
+ found = true
+ break
+ }
+ }
+ if !found || !value["enabled"].(bool) {
+ continue
+ }
+ temp[key] = value
+ }
+
+ marshal, err := json.Marshal(temp)
+ if err != nil {
+ logger.Errorf("marshal error: %v", err)
+ return nil, err
+ }
+ err = json.Unmarshal(marshal, result)
+ if err != nil {
+ logger.Errorf("unmarshal error: %v", err)
+ return nil, err
+ }
+
+ return result, nil
+}
diff --git a/pkg/channels/manager_channel_test.go b/pkg/channels/manager_channel_test.go
new file mode 100644
index 000000000..651764c4f
--- /dev/null
+++ b/pkg/channels/manager_channel_test.go
@@ -0,0 +1,51 @@
+package channels
+
+import (
+ "testing"
+
+ "github.com/stretchr/testify/assert"
+
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/logger"
+)
+
+func TestToChannelHashes(t *testing.T) {
+ logger.SetLevel(logger.DEBUG)
+ cfg := config.DefaultConfig()
+ results := toChannelHashes(cfg)
+ assert.Equal(t, 0, len(results))
+ logger.Debugf("results: %v", results)
+ cfg2 := config.DefaultConfig()
+ cfg2.Channels.DingTalk.Enabled = true
+ results2 := toChannelHashes(cfg2)
+ assert.Equal(t, 1, len(results2))
+ logger.Debugf("results2: %v", results2)
+ added, removed := compareChannels(results, results2)
+ assert.EqualValues(t, []string{"dingtalk"}, added)
+ assert.EqualValues(t, []string(nil), removed)
+ cfg3 := config.DefaultConfig()
+ cfg3.Channels.Telegram.Enabled = true
+ results3 := toChannelHashes(cfg3)
+ assert.Equal(t, 1, len(results3))
+ logger.Debugf("results3: %v", results3)
+ added, removed = compareChannels(results2, results3)
+ assert.EqualValues(t, []string{"dingtalk"}, removed)
+ assert.EqualValues(t, []string{"telegram"}, added)
+ cfg3.Channels.Telegram.Token = "114314"
+ results4 := toChannelHashes(cfg3)
+ assert.Equal(t, 1, len(results4))
+ logger.Debugf("results4: %v", results4)
+ added, removed = compareChannels(results3, results4)
+ assert.EqualValues(t, []string{"telegram"}, removed)
+ assert.EqualValues(t, []string{"telegram"}, added)
+ cc, err := toChannelConfig(cfg3, added)
+ assert.NoError(t, err)
+ logger.Debugf("cc: %#v", cc.Telegram)
+ assert.Equal(t, "114314", cc.Telegram.Token)
+ assert.Equal(t, true, cc.Telegram.Enabled)
+ cc, err = toChannelConfig(cfg2, added)
+ assert.NoError(t, err)
+ logger.Debugf("cc: %#v", cc.Telegram)
+ assert.Equal(t, "", cc.Telegram.Token)
+ assert.Equal(t, false, cc.Telegram.Enabled)
+}
diff --git a/pkg/channels/manager_test.go b/pkg/channels/manager_test.go
index e0f55288a..7dfec9ebf 100644
--- a/pkg/channels/manager_test.go
+++ b/pkg/channels/manager_test.go
@@ -511,6 +511,43 @@ func TestPreSend_PlaceholderEditFails_FallsThrough(t *testing.T) {
}
}
+func TestInvokeTypingStop_CallsRegisteredStop(t *testing.T) {
+ m := newTestManager()
+ var stopCalled bool
+
+ m.RecordTypingStop("telegram", "chat123", func() {
+ stopCalled = true
+ })
+
+ m.InvokeTypingStop("telegram", "chat123")
+
+ if !stopCalled {
+ t.Fatal("expected typing stop func to be called")
+ }
+}
+
+func TestInvokeTypingStop_NoOpWhenNoEntry(t *testing.T) {
+ m := newTestManager()
+ // Should not panic
+ m.InvokeTypingStop("telegram", "nonexistent")
+}
+
+func TestInvokeTypingStop_Idempotent(t *testing.T) {
+ m := newTestManager()
+ var callCount int
+
+ m.RecordTypingStop("telegram", "chat123", func() {
+ callCount++
+ })
+
+ m.InvokeTypingStop("telegram", "chat123")
+ m.InvokeTypingStop("telegram", "chat123") // Second call: entry already removed, no-op
+
+ if callCount != 1 {
+ t.Fatalf("expected stop to be called once, got %d", callCount)
+ }
+}
+
func TestPreSend_TypingStopCalled(t *testing.T) {
m := newTestManager()
var stopCalled bool
diff --git a/pkg/channels/matrix/matrix.go b/pkg/channels/matrix/matrix.go
index bec5dfdac..4cbe95c5c 100644
--- a/pkg/channels/matrix/matrix.go
+++ b/pkg/channels/matrix/matrix.go
@@ -35,8 +35,6 @@ const (
roomKindCacheTTL = 5 * time.Minute
roomKindCacheCleanupPeriod = 1 * time.Minute
roomKindCacheMaxEntries = 2048
-
- matrixMediaTempDirName = "picoclaw_media"
)
var matrixMentionHrefRegexp = regexp.MustCompile(`(?i)]+href=["']([^"']+)["']`)
@@ -1105,7 +1103,7 @@ func (c *MatrixChannel) stripSelfMention(text string) string {
}
func matrixMediaTempDir() (string, error) {
- mediaDir := filepath.Join(os.TempDir(), matrixMediaTempDirName)
+ mediaDir := media.TempDir()
if err := os.MkdirAll(mediaDir, 0o700); err != nil {
return "", err
}
diff --git a/pkg/channels/matrix/matrix_test.go b/pkg/channels/matrix/matrix_test.go
index 07a35c021..7484c8d87 100644
--- a/pkg/channels/matrix/matrix_test.go
+++ b/pkg/channels/matrix/matrix_test.go
@@ -15,6 +15,7 @@ import (
"maunium.net/go/mautrix/id"
"github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/media"
)
func TestMatrixLocalpartMentionRegexp(t *testing.T) {
@@ -165,7 +166,7 @@ func TestMatrixMediaTempDir(t *testing.T) {
if err != nil {
t.Fatalf("matrixMediaTempDir failed: %v", err)
}
- if filepath.Base(dir) != matrixMediaTempDirName {
+ if filepath.Base(dir) != media.TempDirName {
t.Fatalf("unexpected media dir base: %q", filepath.Base(dir))
}
diff --git a/pkg/channels/pico/client.go b/pkg/channels/pico/client.go
new file mode 100644
index 000000000..2c335050d
--- /dev/null
+++ b/pkg/channels/pico/client.go
@@ -0,0 +1,319 @@
+package pico
+
+import (
+ "context"
+ "encoding/json"
+ "fmt"
+ "net/http"
+ "strings"
+ "sync"
+ "time"
+
+ "github.com/google/uuid"
+ "github.com/gorilla/websocket"
+
+ "github.com/sipeed/picoclaw/pkg/bus"
+ "github.com/sipeed/picoclaw/pkg/channels"
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/identity"
+ "github.com/sipeed/picoclaw/pkg/logger"
+)
+
+// PicoClientChannel connects to a remote Pico Protocol WebSocket server.
+type PicoClientChannel struct {
+ *channels.BaseChannel
+ config config.PicoClientConfig
+ conn *picoConn
+ mu sync.Mutex
+ ctx context.Context
+ cancel context.CancelFunc
+}
+
+// NewPicoClientChannel creates a new Pico Protocol client channel.
+func NewPicoClientChannel(
+ cfg config.PicoClientConfig,
+ messageBus *bus.MessageBus,
+) (*PicoClientChannel, error) {
+ if cfg.URL == "" {
+ return nil, fmt.Errorf("pico_client url is required")
+ }
+
+ base := channels.NewBaseChannel("pico_client", cfg, messageBus, cfg.AllowFrom)
+
+ return &PicoClientChannel{
+ BaseChannel: base,
+ config: cfg,
+ }, nil
+}
+
+// Start dials the remote server and begins reading.
+func (c *PicoClientChannel) Start(ctx context.Context) error {
+ logger.InfoC("pico_client", "Starting Pico Client channel")
+ c.ctx, c.cancel = context.WithCancel(ctx)
+
+ if err := c.dial(); err != nil {
+ c.cancel()
+ return fmt.Errorf("pico_client initial connect: %w", err)
+ }
+
+ c.SetRunning(true)
+ go c.reconnectLoop()
+
+ logger.InfoCF("pico_client", "Connected", map[string]any{"url": c.config.URL})
+ return nil
+}
+
+// Stop closes the connection.
+func (c *PicoClientChannel) Stop(ctx context.Context) error {
+ logger.InfoC("pico_client", "Stopping Pico Client channel")
+ c.SetRunning(false)
+ if c.cancel != nil {
+ c.cancel()
+ }
+ c.mu.Lock()
+ if c.conn != nil {
+ c.conn.close()
+ }
+ c.mu.Unlock()
+ logger.InfoC("pico_client", "Pico Client channel stopped")
+ return nil
+}
+
+func (c *PicoClientChannel) dial() error {
+ header := http.Header{}
+ if c.config.Token != "" {
+ header.Set("Authorization", "Bearer "+c.config.Token)
+ }
+
+ ws, resp, err := websocket.DefaultDialer.DialContext(c.ctx, c.config.URL, header)
+ if resp != nil && resp.Body != nil {
+ resp.Body.Close()
+ }
+ if err != nil {
+ return err
+ }
+
+ connCtx, connCancel := context.WithCancel(c.ctx)
+
+ pc := &picoConn{
+ id: uuid.New().String(),
+ conn: ws,
+ sessionID: c.config.SessionID,
+ cancel: connCancel,
+ }
+ if pc.sessionID == "" {
+ pc.sessionID = uuid.New().String()
+ }
+
+ c.mu.Lock()
+ c.conn = pc
+ c.mu.Unlock()
+
+ go c.readLoop(connCtx, pc)
+ return nil
+}
+
+// reconnectLoop re-dials when the connection drops.
+func (c *PicoClientChannel) reconnectLoop() {
+ for {
+ select {
+ case <-c.ctx.Done():
+ return
+ default:
+ }
+
+ c.mu.Lock()
+ pc := c.conn
+ c.mu.Unlock()
+
+ if pc == nil || pc.closed.Load() {
+ backoff := 5 * time.Second
+ logger.InfoC("pico_client", "Reconnecting...")
+ if err := c.dial(); err != nil {
+ logger.WarnCF("pico_client", "Reconnect failed", map[string]any{
+ "error": err.Error(),
+ })
+ select {
+ case <-c.ctx.Done():
+ return
+ case <-time.After(backoff):
+ }
+ continue
+ }
+ logger.InfoC("pico_client", "Reconnected")
+ }
+
+ select {
+ case <-c.ctx.Done():
+ return
+ case <-time.After(1 * time.Second):
+ }
+ }
+}
+
+func (c *PicoClientChannel) readLoop(connCtx context.Context, pc *picoConn) {
+ defer pc.close()
+
+ readTimeout := time.Duration(c.config.ReadTimeout) * time.Second
+ if readTimeout <= 0 {
+ readTimeout = 60 * time.Second
+ }
+
+ _ = pc.conn.SetReadDeadline(time.Now().Add(readTimeout))
+ pc.conn.SetPongHandler(func(string) error {
+ return pc.conn.SetReadDeadline(time.Now().Add(readTimeout))
+ })
+
+ pingInterval := time.Duration(c.config.PingInterval) * time.Second
+ if pingInterval <= 0 {
+ pingInterval = 30 * time.Second
+ }
+ go c.pingLoop(connCtx, pc, pingInterval)
+
+ for {
+ select {
+ case <-connCtx.Done():
+ return
+ default:
+ }
+
+ _, raw, err := pc.conn.ReadMessage()
+ if err != nil {
+ if websocket.IsUnexpectedCloseError(
+ err,
+ websocket.CloseGoingAway,
+ websocket.CloseNormalClosure,
+ ) {
+ logger.DebugCF("pico_client", "Read error", map[string]any{
+ "error": err.Error(),
+ })
+ }
+ return
+ }
+
+ _ = pc.conn.SetReadDeadline(time.Now().Add(readTimeout))
+
+ var msg PicoMessage
+ if err := json.Unmarshal(raw, &msg); err != nil {
+ continue
+ }
+
+ c.handleInbound(pc, msg)
+ }
+}
+
+func (c *PicoClientChannel) pingLoop(connCtx context.Context, pc *picoConn, interval time.Duration) {
+ ticker := time.NewTicker(interval)
+ defer ticker.Stop()
+ for {
+ select {
+ case <-connCtx.Done():
+ return
+ case <-ticker.C:
+ if pc.closed.Load() {
+ return
+ }
+ pc.writeMu.Lock()
+ err := pc.conn.WriteMessage(websocket.PingMessage, nil)
+ pc.writeMu.Unlock()
+ if err != nil {
+ return
+ }
+ }
+ }
+}
+
+// handleInbound processes messages from the remote server.
+// In client mode the server sends message.create (responses) and the client
+// sends message.send (user input). We treat message.create from the server
+// as inbound user messages to feed into the agent loop.
+func (c *PicoClientChannel) handleInbound(pc *picoConn, msg PicoMessage) {
+ switch msg.Type {
+ case TypePong:
+ // response to our ping, ignore
+ case TypeMessageCreate:
+ // Server sent us a message — treat as inbound
+ c.handleServerMessage(pc, msg)
+ default:
+ logger.DebugCF("pico_client", "Ignoring message type", map[string]any{
+ "type": msg.Type,
+ })
+ }
+}
+
+func (c *PicoClientChannel) handleServerMessage(pc *picoConn, msg PicoMessage) {
+ content, _ := msg.Payload["content"].(string)
+ if strings.TrimSpace(content) == "" {
+ return
+ }
+
+ sessionID := msg.SessionID
+ if sessionID == "" {
+ sessionID = pc.sessionID
+ }
+
+ chatID := "pico_client:" + sessionID
+ senderID := "pico-remote"
+ peer := bus.Peer{Kind: "direct", ID: chatID}
+
+ sender := bus.SenderInfo{
+ Platform: "pico_client",
+ PlatformID: senderID,
+ CanonicalID: identity.BuildCanonicalID("pico_client", senderID),
+ }
+
+ if !c.IsAllowedSender(sender) {
+ return
+ }
+
+ c.HandleMessage(c.ctx, peer, msg.ID, senderID, chatID, content, nil, map[string]string{
+ "platform": "pico_client",
+ "session_id": sessionID,
+ }, sender)
+}
+
+// Send sends a message to the remote server.
+func (c *PicoClientChannel) Send(ctx context.Context, msg bus.OutboundMessage) error {
+ if !c.IsRunning() {
+ return channels.ErrNotRunning
+ }
+ c.mu.Lock()
+ pc := c.conn
+ c.mu.Unlock()
+ if pc == nil || pc.closed.Load() {
+ return channels.ErrSendFailed
+ }
+
+ outMsg := newMessage(TypeMessageSend, map[string]any{
+ "content": msg.Content,
+ })
+ outMsg.SessionID = strings.TrimPrefix(msg.ChatID, "pico_client:")
+ return pc.writeJSON(outMsg)
+}
+
+// StartTyping implements channels.TypingCapable.
+func (c *PicoClientChannel) StartTyping(ctx context.Context, chatID string) (func(), error) {
+ c.mu.Lock()
+ pc := c.conn
+ c.mu.Unlock()
+ if pc == nil || pc.closed.Load() {
+ return func() {}, nil
+ }
+
+ startMsg := newMessage(TypeTypingStart, nil)
+ startMsg.SessionID = strings.TrimPrefix(chatID, "pico_client:")
+ if err := pc.writeJSON(startMsg); err != nil {
+ return func() {}, err
+ }
+ return func() {
+ c.mu.Lock()
+ currentPC := c.conn
+ c.mu.Unlock()
+ if currentPC == nil {
+ return
+ }
+ stopMsg := newMessage(TypeTypingStop, nil)
+ stopMsg.SessionID = strings.TrimPrefix(chatID, "pico_client:")
+ currentPC.writeJSON(stopMsg)
+ }, nil
+}
diff --git a/pkg/channels/pico/client_test.go b/pkg/channels/pico/client_test.go
new file mode 100644
index 000000000..118c9abea
--- /dev/null
+++ b/pkg/channels/pico/client_test.go
@@ -0,0 +1,264 @@
+package pico
+
+import (
+ "context"
+ "encoding/json"
+ "errors"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "testing"
+ "time"
+
+ "github.com/gorilla/websocket"
+
+ "github.com/sipeed/picoclaw/pkg/bus"
+ "github.com/sipeed/picoclaw/pkg/channels"
+ "github.com/sipeed/picoclaw/pkg/config"
+)
+
+func TestNewPicoClientChannel_MissingURL(t *testing.T) {
+ _, err := NewPicoClientChannel(config.PicoClientConfig{}, bus.NewMessageBus())
+ if err == nil {
+ t.Fatal("expected error for missing URL")
+ }
+ if !strings.Contains(err.Error(), "url is required") {
+ t.Fatalf("unexpected error: %v", err)
+ }
+}
+
+func TestNewPicoClientChannel_OK(t *testing.T) {
+ ch, err := NewPicoClientChannel(config.PicoClientConfig{
+ URL: "ws://localhost:9999/ws",
+ }, bus.NewMessageBus())
+ if err != nil {
+ t.Fatalf("unexpected error: %v", err)
+ }
+ if ch.Name() != "pico_client" {
+ t.Fatalf("name = %q, want pico_client", ch.Name())
+ }
+}
+
+func TestSend_NotRunning(t *testing.T) {
+ ch, err := NewPicoClientChannel(config.PicoClientConfig{
+ URL: "ws://localhost:9999/ws",
+ }, bus.NewMessageBus())
+ if err != nil {
+ t.Fatal(err)
+ }
+ err = ch.Send(context.Background(), bus.OutboundMessage{Content: "hi"})
+ if !errors.Is(err, channels.ErrNotRunning) {
+ t.Fatalf("expected ErrNotRunning, got %v", err)
+ }
+}
+
+// testServer starts a WS server that echoes message.send back as message.create.
+func testServer(t *testing.T, token string) *httptest.Server {
+ t.Helper()
+ upgrader := websocket.Upgrader{CheckOrigin: func(*http.Request) bool { return true }}
+
+ return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if token != "" {
+ auth := r.Header.Get("Authorization")
+ if auth != "Bearer "+token {
+ http.Error(w, "unauthorized", http.StatusUnauthorized)
+ return
+ }
+ }
+
+ conn, err := upgrader.Upgrade(w, r, nil)
+ if err != nil {
+ t.Logf("upgrade error: %v", err)
+ return
+ }
+ defer conn.Close()
+
+ for {
+ _, raw, err := conn.ReadMessage()
+ if err != nil {
+ return
+ }
+
+ var msg PicoMessage
+ if err := json.Unmarshal(raw, &msg); err != nil {
+ continue
+ }
+
+ if msg.Type == TypeMessageSend {
+ reply := newMessage(TypeMessageCreate, msg.Payload)
+ reply.SessionID = msg.SessionID
+ if err := conn.WriteJSON(reply); err != nil {
+ return
+ }
+ }
+ }
+ }))
+}
+
+func wsURL(httpURL string) string {
+ return "ws" + strings.TrimPrefix(httpURL, "http")
+}
+
+func TestClientChannel_ConnectAndSend(t *testing.T) {
+ srv := testServer(t, "test-token")
+ defer srv.Close()
+
+ mb := bus.NewMessageBus()
+ ch, err := NewPicoClientChannel(config.PicoClientConfig{
+ URL: wsURL(srv.URL),
+ Token: "test-token",
+ SessionID: "sess-1",
+ PingInterval: 60,
+ ReadTimeout: 10,
+ }, mb)
+ if err != nil {
+ t.Fatal(err)
+ }
+
+ ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
+ defer cancel()
+
+ if err = ch.Start(ctx); err != nil {
+ t.Fatalf("Start: %v", err)
+ }
+ defer ch.Stop(ctx)
+
+ // Send a message
+ err = ch.Send(ctx, bus.OutboundMessage{
+ ChatID: "pico_client:sess-1",
+ Content: "hello",
+ })
+ if err != nil {
+ t.Fatalf("Send: %v", err)
+ }
+}
+
+func TestClientChannel_AuthFailure(t *testing.T) {
+ srv := testServer(t, "correct-token")
+ defer srv.Close()
+
+ ch, err := NewPicoClientChannel(config.PicoClientConfig{
+ URL: wsURL(srv.URL),
+ Token: "wrong-token",
+ }, bus.NewMessageBus())
+ if err != nil {
+ t.Fatal(err)
+ }
+
+ ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
+ defer cancel()
+
+ err = ch.Start(ctx)
+ if err == nil {
+ ch.Stop(ctx)
+ t.Fatal("expected auth failure")
+ }
+}
+
+func TestClientChannel_ReceivesServerMessage(t *testing.T) {
+ srv := testServer(t, "")
+ defer srv.Close()
+
+ mb := bus.NewMessageBus()
+
+ ch, err := NewPicoClientChannel(config.PicoClientConfig{
+ URL: wsURL(srv.URL),
+ SessionID: "sess-echo",
+ ReadTimeout: 10,
+ }, mb)
+ if err != nil {
+ t.Fatal(err)
+ }
+
+ ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
+ defer cancel()
+
+ if err = ch.Start(ctx); err != nil {
+ t.Fatalf("Start: %v", err)
+ }
+ defer ch.Stop(ctx)
+
+ // Send a message; the echo server replies with message.create
+ err = ch.Send(ctx, bus.OutboundMessage{
+ ChatID: "pico_client:sess-echo",
+ Content: "ping",
+ })
+ if err != nil {
+ t.Fatalf("Send: %v", err)
+ }
+
+ // The echoed message.create is processed by handleServerMessage which
+ // calls HandleMessage → PublishInbound. Consume it from the bus.
+ select {
+ case msg := <-mb.InboundChan():
+ if msg.Content != "ping" {
+ t.Fatalf("received = %q, want %q", msg.Content, "ping")
+ }
+ case <-ctx.Done():
+ t.Fatal("timed out waiting for echoed message")
+ }
+}
+
+func TestClientChannel_StartTyping(t *testing.T) {
+ srv := testServer(t, "")
+ defer srv.Close()
+
+ ch, err := NewPicoClientChannel(config.PicoClientConfig{
+ URL: wsURL(srv.URL),
+ SessionID: "sess-type",
+ ReadTimeout: 10,
+ }, bus.NewMessageBus())
+ if err != nil {
+ t.Fatal(err)
+ }
+
+ ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
+ defer cancel()
+
+ if err = ch.Start(ctx); err != nil {
+ t.Fatalf("Start: %v", err)
+ }
+ defer ch.Stop(ctx)
+
+ stop, err := ch.StartTyping(ctx, "pico_client:sess-type")
+ if err != nil {
+ t.Fatalf("StartTyping: %v", err)
+ }
+ stop() // should not panic
+}
+
+func TestSend_ClosedConnection(t *testing.T) {
+ srv := testServer(t, "")
+ defer srv.Close()
+
+ ch, err := NewPicoClientChannel(config.PicoClientConfig{
+ URL: wsURL(srv.URL),
+ SessionID: "sess-close",
+ ReadTimeout: 10,
+ }, bus.NewMessageBus())
+ if err != nil {
+ t.Fatal(err)
+ }
+
+ ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
+ defer cancel()
+
+ if err = ch.Start(ctx); err != nil {
+ t.Fatalf("Start: %v", err)
+ }
+
+ // Force close the underlying connection
+ ch.mu.Lock()
+ ch.conn.close()
+ ch.mu.Unlock()
+
+ err = ch.Send(ctx, bus.OutboundMessage{
+ ChatID: "pico_client:sess-close",
+ Content: "should fail",
+ })
+ if !errors.Is(err, channels.ErrSendFailed) {
+ t.Fatalf("expected ErrSendFailed, got %v", err)
+ }
+
+ ch.Stop(ctx)
+}
diff --git a/pkg/channels/pico/init.go b/pkg/channels/pico/init.go
index 96d764418..0319279d8 100644
--- a/pkg/channels/pico/init.go
+++ b/pkg/channels/pico/init.go
@@ -10,4 +10,7 @@ func init() {
channels.RegisterFactory("pico", func(cfg *config.Config, b *bus.MessageBus) (channels.Channel, error) {
return NewPicoChannel(cfg.Channels.Pico, b)
})
+ channels.RegisterFactory("pico_client", func(cfg *config.Config, b *bus.MessageBus) (channels.Channel, error) {
+ return NewPicoClientChannel(cfg.Channels.PicoClient, b)
+ })
}
diff --git a/pkg/channels/pico/pico.go b/pkg/channels/pico/pico.go
index 8d8b62a67..77e7bbdb6 100644
--- a/pkg/channels/pico/pico.go
+++ b/pkg/channels/pico/pico.go
@@ -27,6 +27,7 @@ type picoConn struct {
sessionID string
writeMu sync.Mutex
closed atomic.Bool
+ cancel context.CancelFunc // cancels per-connection goroutines (e.g. pingLoop)
}
// writeJSON sends a JSON message to the connection with write locking.
@@ -42,6 +43,9 @@ func (pc *picoConn) writeJSON(v any) error {
// close closes the connection.
func (pc *picoConn) close() {
if pc.closed.CompareAndSwap(false, true) {
+ if pc.cancel != nil {
+ pc.cancel()
+ }
pc.conn.Close()
}
}
@@ -251,7 +255,13 @@ func (c *PicoChannel) handleWebSocket(w http.ResponseWriter, r *http.Request) {
return
}
- conn, err := c.upgrader.Upgrade(w, r, nil)
+ // Echo the matched subprotocol back so the browser accepts the upgrade.
+ var responseHeader http.Header
+ if proto := c.matchedSubprotocol(r); proto != "" {
+ responseHeader = http.Header{"Sec-WebSocket-Protocol": {proto}}
+ }
+
+ conn, err := c.upgrader.Upgrade(w, r, responseHeader)
if err != nil {
logger.ErrorCF("pico", "WebSocket upgrade failed", map[string]any{
"error": err.Error(),
@@ -282,8 +292,10 @@ func (c *PicoChannel) handleWebSocket(w http.ResponseWriter, r *http.Request) {
go c.readLoop(pc)
}
-// authenticate checks the Bearer token from the Authorization header.
-// Query parameter authentication is only allowed when AllowTokenQuery is explicitly enabled.
+// authenticate checks the request for a valid token:
+// 1. Authorization: Bearer header
+// 2. Sec-WebSocket-Protocol "token." (for browsers that can't set headers)
+// 3. Query parameter "token" (only when AllowTokenQuery is on)
func (c *PicoChannel) authenticate(r *http.Request) bool {
token := c.config.Token
if token == "" {
@@ -298,6 +310,11 @@ func (c *PicoChannel) authenticate(r *http.Request) bool {
}
}
+ // Check Sec-WebSocket-Protocol subprotocol ("token.")
+ if c.matchedSubprotocol(r) != "" {
+ return true
+ }
+
// Check query parameter only when explicitly allowed
if c.config.AllowTokenQuery {
if r.URL.Query().Get("token") == token {
@@ -308,6 +325,18 @@ func (c *PicoChannel) authenticate(r *http.Request) bool {
return false
}
+// matchedSubprotocol returns the "token." subprotocol that matches
+// the configured token, or "" if none do.
+func (c *PicoChannel) matchedSubprotocol(r *http.Request) string {
+ token := c.config.Token
+ for _, proto := range websocket.Subprotocols(r) {
+ if after, ok := strings.CutPrefix(proto, "token."); ok && after == token {
+ return proto
+ }
+ }
+ return ""
+}
+
// readLoop reads messages from a WebSocket connection.
func (c *PicoChannel) readLoop(pc *picoConn) {
defer func() {
diff --git a/pkg/channels/qq/botgo_logger.go b/pkg/channels/qq/botgo_logger.go
new file mode 100644
index 000000000..e1d2462a3
--- /dev/null
+++ b/pkg/channels/qq/botgo_logger.go
@@ -0,0 +1,41 @@
+package qq
+
+import (
+ "fmt"
+ "strings"
+
+ "github.com/sipeed/picoclaw/pkg/logger"
+)
+
+// botGoLogger preserves useful SDK info logs while demoting noisy heartbeat
+// traffic to DEBUG so long-running QQ sessions do not spam the console.
+type botGoLogger struct {
+ *logger.Logger
+}
+
+func newBotGoLogger(component string) *botGoLogger {
+ return &botGoLogger{Logger: logger.NewLogger(component)}
+}
+
+func (b *botGoLogger) Info(v ...any) {
+ message := fmt.Sprint(v...)
+ if shouldDemoteBotGoInfo(message) {
+ b.Logger.Debug(message)
+ return
+ }
+ b.Logger.Info(message)
+}
+
+func (b *botGoLogger) Infof(format string, v ...any) {
+ message := fmt.Sprintf(format, v...)
+ if shouldDemoteBotGoInfo(message) {
+ b.Logger.Debug(message)
+ return
+ }
+ b.Logger.Info(message)
+}
+
+func shouldDemoteBotGoInfo(message string) bool {
+ return strings.Contains(message, " write Heartbeat message") ||
+ strings.Contains(message, " receive HeartbeatAck message")
+}
diff --git a/pkg/channels/qq/qq.go b/pkg/channels/qq/qq.go
index 4cb4db3c6..1a48369f8 100644
--- a/pkg/channels/qq/qq.go
+++ b/pkg/channels/qq/qq.go
@@ -2,7 +2,15 @@ package qq
import (
"context"
+ "encoding/base64"
+ "encoding/json"
+ "errors"
"fmt"
+ "net/http"
+ "net/url"
+ "os"
+ "path"
+ "path/filepath"
"regexp"
"strings"
"sync"
@@ -10,9 +18,10 @@ import (
"time"
"github.com/tencent-connect/botgo"
+ "github.com/tencent-connect/botgo/constant"
"github.com/tencent-connect/botgo/dto"
"github.com/tencent-connect/botgo/event"
- "github.com/tencent-connect/botgo/openapi"
+ "github.com/tencent-connect/botgo/openapi/options"
"github.com/tencent-connect/botgo/token"
"golang.org/x/oauth2"
@@ -21,6 +30,8 @@ import (
"github.com/sipeed/picoclaw/pkg/config"
"github.com/sipeed/picoclaw/pkg/identity"
"github.com/sipeed/picoclaw/pkg/logger"
+ "github.com/sipeed/picoclaw/pkg/media"
+ "github.com/sipeed/picoclaw/pkg/utils"
)
const (
@@ -29,16 +40,29 @@ const (
dedupMaxSize = 10000 // hard cap on dedup map entries
typingResend = 8 * time.Second
typingSeconds = 10
+ bytesPerMiB = 1024 * 1024
)
+type qqAPI interface {
+ WS(ctx context.Context, params map[string]string, body string) (*dto.WebsocketAP, error)
+ PostGroupMessage(
+ ctx context.Context, groupID string, msg dto.APIMessage, opt ...options.Option,
+ ) (*dto.Message, error)
+ PostC2CMessage(
+ ctx context.Context, userID string, msg dto.APIMessage, opt ...options.Option,
+ ) (*dto.Message, error)
+ Transport(ctx context.Context, method, url string, body any) ([]byte, error)
+}
+
type QQChannel struct {
*channels.BaseChannel
config config.QQConfig
- api openapi.OpenAPI
+ api qqAPI
tokenSource oauth2.TokenSource
ctx context.Context
cancel context.CancelFunc
sessionManager botgo.SessionManager
+ downloadFn func(urlStr, filename string) string
// Chat routing: track whether a chatID is group or direct.
chatType sync.Map // chatID → "group" | "direct"
@@ -78,7 +102,7 @@ func (c *QQChannel) Start(ctx context.Context) error {
return fmt.Errorf("QQ app_id and app_secret not configured")
}
- botgo.SetLogger(logger.NewLogger("botgo"))
+ botgo.SetLogger(newBotGoLogger("botgo"))
logger.InfoC("qq", "Starting QQ bot (WebSocket mode)")
// Reinitialize shutdown signal for clean restart.
@@ -199,20 +223,7 @@ func (c *QQChannel) Send(ctx context.Context, msg bus.OutboundMessage) error {
msgToCreate.Content = ""
}
- // Attach passive reply msg_id and msg_seq if available.
- if v, ok := c.lastMsgID.Load(msg.ChatID); ok {
- if msgID, ok := v.(string); ok && msgID != "" {
- msgToCreate.MsgID = msgID
-
- // Increment msg_seq atomically for multi-part replies.
- if counterVal, ok := c.msgSeqCounters.Load(msg.ChatID); ok {
- if counter, ok := counterVal.(*atomic.Uint64); ok {
- seq := counter.Add(1)
- msgToCreate.MsgSeq = uint32(seq)
- }
- }
- }
- }
+ c.applyPassiveReplyMetadata(msg.ChatID, msgToCreate)
// Sanitize URLs in group messages to avoid QQ's URL blacklist rejection.
if chatKind == "group" {
@@ -305,9 +316,9 @@ func (c *QQChannel) StartTyping(ctx context.Context, chatID string) (func(), err
}
// SendMedia implements the channels.MediaSender interface.
-// QQ RichMediaMessage requires an HTTP/HTTPS URL — local file paths are not supported.
-// If part.Ref is already an http(s) URL it is used directly; otherwise we try
-// the media store, and skip with a warning if the resolved path is not an HTTP URL.
+// QQ group/C2C media sending is a two-step flow:
+// 1. Upload media to /files using a remote URL or base64-encoded local bytes.
+// 2. Send a msg_type=7 message using the returned file_info.
func (c *QQChannel) SendMedia(ctx context.Context, msg bus.OutboundMediaMessage) error {
if !c.IsRunning() {
return channels.ErrNotRunning
@@ -316,69 +327,24 @@ func (c *QQChannel) SendMedia(ctx context.Context, msg bus.OutboundMediaMessage)
chatKind := c.getChatKind(msg.ChatID)
for _, part := range msg.Parts {
- // If the ref is already an HTTP(S) URL, use it directly.
- mediaURL := part.Ref
- if !isHTTPURL(mediaURL) {
- // Try resolving through media store.
- store := c.GetMediaStore()
- if store == nil {
- logger.WarnCF("qq", "QQ media requires HTTP/HTTPS URL, no media store available", map[string]any{
- "ref": part.Ref,
- })
- continue
+ fileInfo, err := c.uploadMedia(ctx, chatKind, msg.ChatID, part)
+ if err != nil {
+ logger.ErrorCF("qq", "Failed to upload media", map[string]any{
+ "type": part.Type,
+ "chat_id": msg.ChatID,
+ "error": err.Error(),
+ })
+ if errors.Is(err, channels.ErrSendFailed) {
+ return err
}
-
- resolved, err := store.Resolve(part.Ref)
- if err != nil {
- logger.ErrorCF("qq", "Failed to resolve media ref", map[string]any{
- "ref": part.Ref,
- "error": err.Error(),
- })
- continue
- }
-
- if !isHTTPURL(resolved) {
- logger.WarnCF("qq", "QQ media requires HTTP/HTTPS URL, local files not supported", map[string]any{
- "ref": part.Ref,
- "resolved": resolved,
- })
- continue
- }
-
- mediaURL = resolved
+ return fmt.Errorf("qq send media: %w", channels.ErrTemporary)
}
- // Map part type to QQ file type: 1=image, 2=video, 3=audio, 4=file.
- var fileType uint64
- switch part.Type {
- case "image":
- fileType = 1
- case "video":
- fileType = 2
- case "audio":
- fileType = 3
- default:
- fileType = 4 // file
- }
-
- richMedia := &dto.RichMediaMessage{
- FileType: fileType,
- URL: mediaURL,
- SrvSendMsg: true,
- }
-
- var sendErr error
- if chatKind == "group" {
- _, sendErr = c.api.PostGroupMessage(ctx, msg.ChatID, richMedia)
- } else {
- _, sendErr = c.api.PostC2CMessage(ctx, msg.ChatID, richMedia)
- }
-
- if sendErr != nil {
+ if err := c.sendUploadedMedia(ctx, chatKind, msg.ChatID, part, fileInfo); err != nil {
logger.ErrorCF("qq", "Failed to send media", map[string]any{
"type": part.Type,
"chat_id": msg.ChatID,
- "error": sendErr.Error(),
+ "error": err.Error(),
})
return fmt.Errorf("qq send media: %w", channels.ErrTemporary)
}
@@ -387,6 +353,161 @@ func (c *QQChannel) SendMedia(ctx context.Context, msg bus.OutboundMediaMessage)
return nil
}
+type qqMediaUpload struct {
+ FileType uint64 `json:"file_type"`
+ URL string `json:"url,omitempty"`
+ FileData string `json:"file_data,omitempty"`
+ SrvSendMsg bool `json:"srv_send_msg,omitempty"`
+}
+
+func (c *QQChannel) uploadMedia(
+ ctx context.Context,
+ chatKind, chatID string,
+ part bus.MediaPart,
+) ([]byte, error) {
+ payload, err := c.buildMediaUpload(part)
+ if err != nil {
+ return nil, err
+ }
+
+ body, err := c.api.Transport(ctx, http.MethodPost, c.mediaUploadURL(chatKind, chatID), payload)
+ if err != nil {
+ return nil, err
+ }
+
+ var uploaded dto.Message
+ if err := json.Unmarshal(body, &uploaded); err != nil {
+ return nil, fmt.Errorf("qq decode media upload response: %w", err)
+ }
+ if len(uploaded.FileInfo) == 0 {
+ return nil, fmt.Errorf("qq upload media: missing file_info")
+ }
+
+ return uploaded.FileInfo, nil
+}
+
+func (c *QQChannel) buildMediaUpload(part bus.MediaPart) (*qqMediaUpload, error) {
+ payload := &qqMediaUpload{
+ FileType: qqFileType(part.Type),
+ }
+
+ mediaRef := part.Ref
+ if isHTTPURL(mediaRef) {
+ payload.URL = mediaRef
+ return payload, nil
+ }
+
+ store := c.GetMediaStore()
+ if store == nil {
+ return nil, fmt.Errorf("no media store available: %w", channels.ErrSendFailed)
+ }
+
+ resolved, err := store.Resolve(part.Ref)
+ if err != nil {
+ return nil, fmt.Errorf("qq resolve media ref %q: %v: %w", part.Ref, err, channels.ErrSendFailed)
+ }
+
+ if isHTTPURL(resolved) {
+ payload.URL = resolved
+ return payload, nil
+ }
+
+ if limitBytes := c.maxBase64FileSizeBytes(); limitBytes > 0 {
+ info, statErr := os.Stat(resolved)
+ if statErr != nil {
+ return nil, fmt.Errorf("qq stat local media %q: %v: %w", resolved, statErr, channels.ErrSendFailed)
+ }
+ if info.Size() > limitBytes {
+ return nil, fmt.Errorf(
+ "qq local media %q exceeds max_base64_file_size_mib (%d > %d bytes): %w",
+ resolved,
+ info.Size(),
+ limitBytes,
+ channels.ErrSendFailed,
+ )
+ }
+ }
+
+ data, err := os.ReadFile(resolved)
+ if err != nil {
+ return nil, fmt.Errorf("qq read local media %q: %v: %w", resolved, err, channels.ErrSendFailed)
+ }
+
+ payload.FileData = base64.StdEncoding.EncodeToString(data)
+ return payload, nil
+}
+
+func (c *QQChannel) sendUploadedMedia(
+ ctx context.Context,
+ chatKind, chatID string,
+ part bus.MediaPart,
+ fileInfo []byte,
+) error {
+ msg := &dto.MessageToCreate{
+ Content: part.Caption,
+ MsgType: dto.RichMediaMsg,
+ Media: &dto.MediaInfo{
+ FileInfo: fileInfo,
+ },
+ }
+ c.applyPassiveReplyMetadata(chatID, msg)
+
+ if chatKind == "group" && msg.Content != "" {
+ msg.Content = sanitizeURLs(msg.Content)
+ }
+
+ if chatKind == "group" {
+ _, err := c.api.PostGroupMessage(ctx, chatID, msg)
+ return err
+ }
+ _, err := c.api.PostC2CMessage(ctx, chatID, msg)
+ return err
+}
+
+func (c *QQChannel) applyPassiveReplyMetadata(chatID string, msg *dto.MessageToCreate) {
+ if v, ok := c.lastMsgID.Load(chatID); ok {
+ if msgID, ok := v.(string); ok && msgID != "" {
+ msg.MsgID = msgID
+
+ // Increment msg_seq atomically for multi-part replies.
+ if counterVal, ok := c.msgSeqCounters.Load(chatID); ok {
+ if counter, ok := counterVal.(*atomic.Uint64); ok {
+ seq := counter.Add(1)
+ msg.MsgSeq = uint32(seq)
+ }
+ }
+ }
+ }
+}
+
+func (c *QQChannel) mediaUploadURL(chatKind, chatID string) string {
+ base := constant.APIDomain
+ if chatKind == "group" {
+ return fmt.Sprintf("%s/v2/groups/%s/files", base, chatID)
+ }
+ return fmt.Sprintf("%s/v2/users/%s/files", base, chatID)
+}
+
+func qqFileType(partType string) uint64 {
+ switch partType {
+ case "image":
+ return 1
+ case "video":
+ return 2
+ case "audio":
+ return 3
+ default:
+ return 4
+ }
+}
+
+func (c *QQChannel) maxBase64FileSizeBytes() int64 {
+ if c.config.MaxBase64FileSizeMiB <= 0 {
+ return 0
+ }
+ return c.config.MaxBase64FileSizeMiB * bytesPerMiB
+}
+
// handleC2CMessage handles QQ private messages.
func (c *QQChannel) handleC2CMessage() event.C2CMessageEventHandler {
return func(event *dto.WSPayload, data *dto.WSC2CMessageData) error {
@@ -404,16 +525,30 @@ func (c *QQChannel) handleC2CMessage() event.C2CMessageEventHandler {
return nil
}
- // extract message content
- content := data.Content
- if content == "" {
- logger.DebugC("qq", "Received empty message, ignoring")
+ sender := bus.SenderInfo{
+ Platform: "qq",
+ PlatformID: data.Author.ID,
+ CanonicalID: identity.BuildCanonicalID("qq", data.Author.ID),
+ }
+
+ if !c.IsAllowedSender(sender) {
+ return nil
+ }
+
+ content := strings.TrimSpace(data.Content)
+ mediaPaths, attachmentNotes := c.extractInboundAttachments(senderID, data.ID, data.Attachments)
+ for _, note := range attachmentNotes {
+ content = appendContent(content, note)
+ }
+ if content == "" && len(mediaPaths) == 0 {
+ logger.DebugC("qq", "Received empty C2C message with no attachments, ignoring")
return nil
}
logger.InfoCF("qq", "Received C2C message", map[string]any{
- "sender": senderID,
- "length": len(content),
+ "sender": senderID,
+ "length": len(content),
+ "media_count": len(mediaPaths),
})
// Store chat routing context.
@@ -427,23 +562,13 @@ func (c *QQChannel) handleC2CMessage() event.C2CMessageEventHandler {
"account_id": senderID,
}
- sender := bus.SenderInfo{
- Platform: "qq",
- PlatformID: data.Author.ID,
- CanonicalID: identity.BuildCanonicalID("qq", data.Author.ID),
- }
-
- if !c.IsAllowedSender(sender) {
- return nil
- }
-
c.HandleMessage(c.ctx,
bus.Peer{Kind: "direct", ID: senderID},
data.ID,
senderID,
senderID,
content,
- []string{},
+ mediaPaths,
metadata,
sender,
)
@@ -469,24 +594,38 @@ func (c *QQChannel) handleGroupATMessage() event.GroupATMessageEventHandler {
return nil
}
- // extract message content (remove @ bot part)
- content := data.Content
- if content == "" {
- logger.DebugC("qq", "Received empty group message, ignoring")
+ sender := bus.SenderInfo{
+ Platform: "qq",
+ PlatformID: data.Author.ID,
+ CanonicalID: identity.BuildCanonicalID("qq", data.Author.ID),
+ }
+
+ if !c.IsAllowedSender(sender) {
return nil
}
- // GroupAT event means bot is always mentioned; apply group trigger filtering
+ content := strings.TrimSpace(data.Content)
+ mediaPaths, attachmentNotes := c.extractInboundAttachments(data.GroupID, data.ID, data.Attachments)
+ for _, note := range attachmentNotes {
+ content = appendContent(content, note)
+ }
+
+ // GroupAT event means bot is always mentioned; apply group trigger filtering.
respond, cleaned := c.ShouldRespondInGroup(true, content)
if !respond {
return nil
}
content = cleaned
+ if content == "" && len(mediaPaths) == 0 {
+ logger.DebugC("qq", "Received empty group message with no attachments, ignoring")
+ return nil
+ }
logger.InfoCF("qq", "Received group AT message", map[string]any{
- "sender": senderID,
- "group": data.GroupID,
- "length": len(content),
+ "sender": senderID,
+ "group": data.GroupID,
+ "length": len(content),
+ "media_count": len(mediaPaths),
})
// Store chat routing context using GroupID as chatID.
@@ -501,23 +640,13 @@ func (c *QQChannel) handleGroupATMessage() event.GroupATMessageEventHandler {
"group_id": data.GroupID,
}
- sender := bus.SenderInfo{
- Platform: "qq",
- PlatformID: data.Author.ID,
- CanonicalID: identity.BuildCanonicalID("qq", data.Author.ID),
- }
-
- if !c.IsAllowedSender(sender) {
- return nil
- }
-
c.HandleMessage(c.ctx,
bus.Peer{Kind: "group", ID: data.GroupID},
data.ID,
senderID,
data.GroupID,
content,
- []string{},
+ mediaPaths,
metadata,
sender,
)
@@ -526,6 +655,157 @@ func (c *QQChannel) handleGroupATMessage() event.GroupATMessageEventHandler {
}
}
+func (c *QQChannel) extractInboundAttachments(
+ chatID, messageID string,
+ attachments []*dto.MessageAttachment,
+) ([]string, []string) {
+ if len(attachments) == 0 {
+ return nil, nil
+ }
+
+ scope := channels.BuildMediaScope("qq", chatID, messageID)
+ mediaPaths := make([]string, 0, len(attachments))
+ notes := make([]string, 0, len(attachments))
+
+ storeMedia := func(localPath string, attachment *dto.MessageAttachment) string {
+ if store := c.GetMediaStore(); store != nil {
+ ref, err := store.Store(localPath, media.MediaMeta{
+ Filename: qqAttachmentFilename(attachment),
+ ContentType: attachment.ContentType,
+ Source: "qq",
+ }, scope)
+ if err == nil {
+ return ref
+ }
+ }
+ return localPath
+ }
+
+ for _, attachment := range attachments {
+ if attachment == nil {
+ continue
+ }
+
+ filename := qqAttachmentFilename(attachment)
+ if localPath := c.downloadAttachment(attachment.URL, filename); localPath != "" {
+ mediaPaths = append(mediaPaths, storeMedia(localPath, attachment))
+ } else if attachment.URL != "" {
+ mediaPaths = append(mediaPaths, attachment.URL)
+ }
+
+ notes = append(notes, qqAttachmentNote(attachment))
+ }
+
+ return mediaPaths, notes
+}
+
+func (c *QQChannel) downloadAttachment(urlStr, filename string) string {
+ if urlStr == "" {
+ return ""
+ }
+ if c.downloadFn != nil {
+ return c.downloadFn(urlStr, filename)
+ }
+
+ return utils.DownloadFile(urlStr, filename, utils.DownloadOptions{
+ LoggerPrefix: "qq",
+ ExtraHeaders: c.downloadHeaders(),
+ })
+}
+
+func (c *QQChannel) downloadHeaders() map[string]string {
+ headers := map[string]string{}
+
+ if c.config.AppID != "" {
+ headers["X-Union-Appid"] = c.config.AppID
+ }
+
+ if c.tokenSource != nil {
+ if tk, err := c.tokenSource.Token(); err == nil && tk.AccessToken != "" {
+ auth := strings.TrimSpace(tk.TokenType + " " + tk.AccessToken)
+ if auth != "" {
+ headers["Authorization"] = auth
+ }
+ }
+ }
+
+ if len(headers) == 0 {
+ return nil
+ }
+ return headers
+}
+
+func qqAttachmentFilename(attachment *dto.MessageAttachment) string {
+ if attachment == nil {
+ return "attachment"
+ }
+ if attachment.FileName != "" {
+ return attachment.FileName
+ }
+ if attachment.URL != "" {
+ if parsed, err := url.Parse(attachment.URL); err == nil {
+ if base := path.Base(parsed.Path); base != "" && base != "." && base != "/" {
+ return base
+ }
+ }
+ }
+
+ switch qqAttachmentKind(attachment) {
+ case "image":
+ return "image"
+ case "audio":
+ return "audio"
+ case "video":
+ return "video"
+ default:
+ return "attachment"
+ }
+}
+
+func qqAttachmentKind(attachment *dto.MessageAttachment) string {
+ if attachment == nil {
+ return "file"
+ }
+
+ contentType := strings.ToLower(attachment.ContentType)
+ filename := strings.ToLower(attachment.FileName)
+
+ switch {
+ case strings.HasPrefix(contentType, "image/"):
+ return "image"
+ case strings.HasPrefix(contentType, "video/"):
+ return "video"
+ case strings.HasPrefix(contentType, "audio/"), contentType == "application/ogg", contentType == "application/x-ogg":
+ return "audio"
+ }
+
+ switch filepath.Ext(filename) {
+ case ".jpg", ".jpeg", ".png", ".gif", ".webp", ".bmp", ".svg":
+ return "image"
+ case ".mp4", ".avi", ".mov", ".webm", ".mkv":
+ return "video"
+ case ".mp3", ".wav", ".ogg", ".m4a", ".flac", ".aac", ".wma", ".opus", ".silk":
+ return "audio"
+ default:
+ return "file"
+ }
+}
+
+func qqAttachmentNote(attachment *dto.MessageAttachment) string {
+ filename := qqAttachmentFilename(attachment)
+
+ switch qqAttachmentKind(attachment) {
+ case "image":
+ return fmt.Sprintf("[image: %s]", filename)
+ case "audio":
+ return fmt.Sprintf("[audio: %s]", filename)
+ case "video":
+ return fmt.Sprintf("[video: %s]", filename)
+ default:
+ return fmt.Sprintf("[file: %s]", filename)
+ }
+}
+
// isDuplicate checks whether a message has been seen within the TTL window.
// It also enforces a hard cap on map size by evicting oldest entries.
func (c *QQChannel) isDuplicate(messageID string) bool {
@@ -587,6 +867,16 @@ func isHTTPURL(s string) bool {
return strings.HasPrefix(s, "http://") || strings.HasPrefix(s, "https://")
}
+func appendContent(content, suffix string) string {
+ if suffix == "" {
+ return content
+ }
+ if content == "" {
+ return suffix
+ }
+ return content + "\n" + suffix
+}
+
// urlPattern matches URLs with explicit http(s):// scheme.
// Only scheme-prefixed URLs are matched to avoid false positives on bare text
// like version numbers (e.g., "1.2.3") or domain-like fragments.
diff --git a/pkg/channels/qq/qq_test.go b/pkg/channels/qq/qq_test.go
index 3ceee0d09..3cb3d39bd 100644
--- a/pkg/channels/qq/qq_test.go
+++ b/pkg/channels/qq/qq_test.go
@@ -2,13 +2,22 @@ package qq
import (
"context"
+ "encoding/base64"
+ "encoding/json"
+ "errors"
+ "os"
+ "strings"
+ "sync/atomic"
"testing"
"time"
"github.com/tencent-connect/botgo/dto"
+ "github.com/tencent-connect/botgo/openapi/options"
"github.com/sipeed/picoclaw/pkg/bus"
"github.com/sipeed/picoclaw/pkg/channels"
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/media"
)
func TestHandleC2CMessage_IncludesAccountIDMetadata(t *testing.T) {
@@ -34,11 +43,454 @@ func TestHandleC2CMessage_IncludesAccountIDMetadata(t *testing.T) {
ctx, cancel := context.WithTimeout(context.Background(), time.Second)
defer cancel()
- inbound, ok := messageBus.ConsumeInbound(ctx)
- if !ok {
- t.Fatal("expected inbound message")
- }
- if inbound.Metadata["account_id"] != "7750283E123456" {
- t.Fatalf("account_id metadata = %q, want %q", inbound.Metadata["account_id"], "7750283E123456")
+ for {
+ select {
+ case <-ctx.Done():
+ t.Fatal("timeout waiting for inbound message")
+ return
+ case inbound, ok := <-messageBus.InboundChan():
+ if !ok {
+ t.Fatal("expected inbound message")
+ }
+ if inbound.Metadata["account_id"] != "7750283E123456" {
+ t.Fatalf("account_id metadata = %q, want %q", inbound.Metadata["account_id"], "7750283E123456")
+ }
+ return
+ }
}
}
+
+func TestHandleC2CMessage_AttachmentOnlyPublishesMedia(t *testing.T) {
+ messageBus := bus.NewMessageBus()
+ store := media.NewFileMediaStore()
+ localPath := writeTempFile(t, t.TempDir(), "image.png", []byte("fake-image"))
+
+ ch := &QQChannel{
+ BaseChannel: channels.NewBaseChannel("qq", nil, messageBus, nil),
+ dedup: make(map[string]time.Time),
+ done: make(chan struct{}),
+ ctx: context.Background(),
+ downloadFn: func(urlStr, filename string) string {
+ if filename != "image.png" {
+ t.Fatalf("download filename = %q, want image.png", filename)
+ }
+ return localPath
+ },
+ }
+ ch.SetMediaStore(store)
+
+ err := ch.handleC2CMessage()(nil, &dto.WSC2CMessageData{
+ ID: "msg-attachment",
+ Content: "",
+ Author: &dto.User{
+ ID: "7750283E123456",
+ },
+ Attachments: []*dto.MessageAttachment{{
+ URL: "https://example.com/image.png",
+ FileName: "image.png",
+ ContentType: "image/png",
+ }},
+ })
+ if err != nil {
+ t.Fatalf("handleC2CMessage() error = %v", err)
+ }
+
+ inbound := waitInboundMessage(t, messageBus)
+ if inbound.Content != "[image: image.png]" {
+ t.Fatalf("inbound.Content = %q", inbound.Content)
+ }
+ if len(inbound.Media) != 1 {
+ t.Fatalf("len(inbound.Media) = %d, want 1", len(inbound.Media))
+ }
+ if !strings.HasPrefix(inbound.Media[0], "media://") {
+ t.Fatalf("inbound.Media[0] = %q, want media:// ref", inbound.Media[0])
+ }
+ _, meta, err := store.ResolveWithMeta(inbound.Media[0])
+ if err != nil {
+ t.Fatalf("ResolveWithMeta() error = %v", err)
+ }
+ if meta.Filename != "image.png" {
+ t.Fatalf("meta.Filename = %q, want image.png", meta.Filename)
+ }
+ if meta.ContentType != "image/png" {
+ t.Fatalf("meta.ContentType = %q, want image/png", meta.ContentType)
+ }
+}
+
+func TestHandleGroupATMessage_AttachmentOnlyPublishesMedia(t *testing.T) {
+ messageBus := bus.NewMessageBus()
+ store := media.NewFileMediaStore()
+ localPath := writeTempFile(t, t.TempDir(), "report.pdf", []byte("fake-pdf"))
+
+ ch := &QQChannel{
+ BaseChannel: channels.NewBaseChannel("qq", nil, messageBus, nil),
+ dedup: make(map[string]time.Time),
+ done: make(chan struct{}),
+ ctx: context.Background(),
+ downloadFn: func(urlStr, filename string) string {
+ if filename != "report.pdf" {
+ t.Fatalf("download filename = %q, want report.pdf", filename)
+ }
+ return localPath
+ },
+ }
+ ch.SetMediaStore(store)
+
+ err := ch.handleGroupATMessage()(nil, &dto.WSGroupATMessageData{
+ ID: "group-attachment",
+ GroupID: "group-1",
+ Content: "",
+ Author: &dto.User{
+ ID: "7750283E123456",
+ },
+ Attachments: []*dto.MessageAttachment{{
+ URL: "https://example.com/report.pdf",
+ FileName: "report.pdf",
+ ContentType: "application/pdf",
+ }},
+ })
+ if err != nil {
+ t.Fatalf("handleGroupATMessage() error = %v", err)
+ }
+
+ inbound := waitInboundMessage(t, messageBus)
+ if inbound.Content != "[file: report.pdf]" {
+ t.Fatalf("inbound.Content = %q", inbound.Content)
+ }
+ if len(inbound.Media) != 1 {
+ t.Fatalf("len(inbound.Media) = %d, want 1", len(inbound.Media))
+ }
+ if !strings.HasPrefix(inbound.Media[0], "media://") {
+ t.Fatalf("inbound.Media[0] = %q, want media:// ref", inbound.Media[0])
+ }
+ if inbound.Peer.Kind != "group" || inbound.Peer.ID != "group-1" {
+ t.Fatalf("inbound.Peer = %+v, want group/group-1", inbound.Peer)
+ }
+}
+
+func TestSendMedia_UploadsLocalFileAsBase64(t *testing.T) {
+ messageBus := bus.NewMessageBus()
+ store := media.NewFileMediaStore()
+
+ tmpFile, err := os.CreateTemp(t.TempDir(), "qq-media-*.png")
+ if err != nil {
+ t.Fatalf("CreateTemp() error = %v", err)
+ }
+ defer tmpFile.Close()
+
+ content := []byte("local-image-data")
+ if _, writeErr := tmpFile.Write(content); writeErr != nil {
+ t.Fatalf("Write() error = %v", writeErr)
+ }
+
+ ref, err := store.Store(tmpFile.Name(), media.MediaMeta{
+ Filename: "reply.png",
+ ContentType: "image/png",
+ }, "qq:test")
+ if err != nil {
+ t.Fatalf("Store() error = %v", err)
+ }
+
+ api := &fakeQQAPI{
+ transportResp: mustJSON(t, dto.Message{FileInfo: []byte("uploaded-file-info")}),
+ }
+ ch := &QQChannel{
+ BaseChannel: channels.NewBaseChannel("qq", nil, messageBus, nil),
+ api: api,
+ dedup: make(map[string]time.Time),
+ done: make(chan struct{}),
+ ctx: context.Background(),
+ }
+ ch.SetRunning(true)
+ ch.SetMediaStore(store)
+ ch.chatType.Store("group-1", "group")
+ ch.lastMsgID.Store("group-1", "msg-1")
+ ch.msgSeqCounters.Store("group-1", new(atomic.Uint64))
+
+ err = ch.SendMedia(context.Background(), bus.OutboundMediaMessage{
+ ChatID: "group-1",
+ Parts: []bus.MediaPart{{
+ Type: "image",
+ Ref: ref,
+ Caption: "see https://example.com/image",
+ }},
+ })
+ if err != nil {
+ t.Fatalf("SendMedia() error = %v", err)
+ }
+
+ if len(api.transportCalls) != 1 {
+ t.Fatalf("transportCalls = %d, want 1", len(api.transportCalls))
+ }
+ upload := api.transportCalls[0]
+ if upload.method != "POST" {
+ t.Fatalf("upload method = %q, want POST", upload.method)
+ }
+ if upload.url != "https://api.sgroup.qq.com/v2/groups/group-1/files" {
+ t.Fatalf("upload url = %q", upload.url)
+ }
+ if upload.body.URL != "" {
+ t.Fatalf("upload URL = %q, want empty", upload.body.URL)
+ }
+ wantBase64 := base64.StdEncoding.EncodeToString(content)
+ if upload.body.FileData != wantBase64 {
+ t.Fatalf("upload file_data = %q, want %q", upload.body.FileData, wantBase64)
+ }
+ if upload.body.FileType != 1 {
+ t.Fatalf("upload file_type = %d, want 1", upload.body.FileType)
+ }
+
+ if len(api.groupMessages) != 1 {
+ t.Fatalf("groupMessages = %d, want 1", len(api.groupMessages))
+ }
+ msg, ok := api.groupMessages[0].(*dto.MessageToCreate)
+ if !ok {
+ t.Fatalf("groupMessages[0] type = %T, want *dto.MessageToCreate", api.groupMessages[0])
+ }
+ if msg.MsgType != dto.RichMediaMsg {
+ t.Fatalf("msg.MsgType = %d, want %d", msg.MsgType, dto.RichMediaMsg)
+ }
+ if msg.MsgID != "msg-1" {
+ t.Fatalf("msg.MsgID = %q, want msg-1", msg.MsgID)
+ }
+ if msg.MsgSeq != 1 {
+ t.Fatalf("msg.MsgSeq = %d, want 1", msg.MsgSeq)
+ }
+ if msg.Content != "see https://example。com/image" {
+ t.Fatalf("msg.Content = %q", msg.Content)
+ }
+ if msg.Media == nil || string(msg.Media.FileInfo) != "uploaded-file-info" {
+ t.Fatalf("msg.Media.FileInfo = %q, want uploaded-file-info", string(msg.Media.FileInfo))
+ }
+}
+
+func TestSendMedia_UsesRemoteURLUploadForC2C(t *testing.T) {
+ messageBus := bus.NewMessageBus()
+ api := &fakeQQAPI{
+ transportResp: mustJSON(t, dto.Message{FileInfo: []byte("remote-file-info")}),
+ }
+ ch := &QQChannel{
+ BaseChannel: channels.NewBaseChannel("qq", nil, messageBus, nil),
+ api: api,
+ dedup: make(map[string]time.Time),
+ done: make(chan struct{}),
+ ctx: context.Background(),
+ }
+ ch.SetRunning(true)
+ ch.chatType.Store("user-1", "direct")
+
+ err := ch.SendMedia(context.Background(), bus.OutboundMediaMessage{
+ ChatID: "user-1",
+ Parts: []bus.MediaPart{{
+ Type: "file",
+ Ref: "https://cdn.example.com/report.pdf",
+ }},
+ })
+ if err != nil {
+ t.Fatalf("SendMedia() error = %v", err)
+ }
+
+ if len(api.transportCalls) != 1 {
+ t.Fatalf("transportCalls = %d, want 1", len(api.transportCalls))
+ }
+ upload := api.transportCalls[0]
+ if upload.url != "https://api.sgroup.qq.com/v2/users/user-1/files" {
+ t.Fatalf("upload url = %q", upload.url)
+ }
+ if upload.body.URL != "https://cdn.example.com/report.pdf" {
+ t.Fatalf("upload URL = %q", upload.body.URL)
+ }
+ if upload.body.FileData != "" {
+ t.Fatalf("upload file_data = %q, want empty", upload.body.FileData)
+ }
+ if upload.body.FileType != 4 {
+ t.Fatalf("upload file_type = %d, want 4", upload.body.FileType)
+ }
+
+ if len(api.c2cMessages) != 1 {
+ t.Fatalf("c2cMessages = %d, want 1", len(api.c2cMessages))
+ }
+ msg, ok := api.c2cMessages[0].(*dto.MessageToCreate)
+ if !ok {
+ t.Fatalf("c2cMessages[0] type = %T, want *dto.MessageToCreate", api.c2cMessages[0])
+ }
+ if msg.MsgType != dto.RichMediaMsg {
+ t.Fatalf("msg.MsgType = %d, want %d", msg.MsgType, dto.RichMediaMsg)
+ }
+ if msg.Media == nil || string(msg.Media.FileInfo) != "remote-file-info" {
+ t.Fatalf("msg.Media.FileInfo = %q, want remote-file-info", string(msg.Media.FileInfo))
+ }
+}
+
+func TestSendMedia_ReturnsSendFailedWithoutMediaStore(t *testing.T) {
+ messageBus := bus.NewMessageBus()
+ ch := &QQChannel{
+ BaseChannel: channels.NewBaseChannel("qq", nil, messageBus, nil),
+ api: &fakeQQAPI{},
+ dedup: make(map[string]time.Time),
+ done: make(chan struct{}),
+ ctx: context.Background(),
+ }
+ ch.SetRunning(true)
+ ch.chatType.Store("group-1", "group")
+
+ err := ch.SendMedia(context.Background(), bus.OutboundMediaMessage{
+ ChatID: "group-1",
+ Parts: []bus.MediaPart{{
+ Type: "image",
+ Ref: "media://missing",
+ }},
+ })
+ if !errors.Is(err, channels.ErrSendFailed) {
+ t.Fatalf("SendMedia() error = %v, want ErrSendFailed", err)
+ }
+}
+
+func TestSendMedia_ReturnsSendFailedWhenLocalFileExceedsBase64MiBLimit(t *testing.T) {
+ messageBus := bus.NewMessageBus()
+ store := media.NewFileMediaStore()
+
+ tmpFile, err := os.CreateTemp(t.TempDir(), "qq-media-too-large-*.bin")
+ if err != nil {
+ t.Fatalf("CreateTemp() error = %v", err)
+ }
+ defer tmpFile.Close()
+
+ content := make([]byte, bytesPerMiB+1)
+ if _, writeErr := tmpFile.Write(content); writeErr != nil {
+ t.Fatalf("Write() error = %v", writeErr)
+ }
+
+ ref, err := store.Store(tmpFile.Name(), media.MediaMeta{
+ Filename: "large.bin",
+ ContentType: "application/octet-stream",
+ }, "qq:test")
+ if err != nil {
+ t.Fatalf("Store() error = %v", err)
+ }
+
+ api := &fakeQQAPI{}
+ ch := &QQChannel{
+ BaseChannel: channels.NewBaseChannel("qq", nil, messageBus, nil),
+ config: config.QQConfig{
+ MaxBase64FileSizeMiB: 1,
+ },
+ api: api,
+ dedup: make(map[string]time.Time),
+ done: make(chan struct{}),
+ ctx: context.Background(),
+ }
+ ch.SetRunning(true)
+ ch.SetMediaStore(store)
+ ch.chatType.Store("group-1", "group")
+
+ err = ch.SendMedia(context.Background(), bus.OutboundMediaMessage{
+ ChatID: "group-1",
+ Parts: []bus.MediaPart{{
+ Type: "file",
+ Ref: ref,
+ }},
+ })
+ if !errors.Is(err, channels.ErrSendFailed) {
+ t.Fatalf("SendMedia() error = %v, want ErrSendFailed", err)
+ }
+ if len(api.transportCalls) != 0 {
+ t.Fatalf("transportCalls = %d, want 0", len(api.transportCalls))
+ }
+}
+
+type fakeQQAPI struct {
+ transportResp []byte
+ transportErr error
+ groupErr error
+ c2cErr error
+ transportCalls []fakeTransportCall
+ groupMessages []dto.APIMessage
+ c2cMessages []dto.APIMessage
+}
+
+type fakeTransportCall struct {
+ method string
+ url string
+ body qqMediaUpload
+}
+
+func (f *fakeQQAPI) WS(
+ context.Context,
+ map[string]string,
+ string,
+) (*dto.WebsocketAP, error) {
+ return nil, nil
+}
+
+func (f *fakeQQAPI) PostGroupMessage(
+ _ context.Context,
+ _ string,
+ msg dto.APIMessage,
+ _ ...options.Option,
+) (*dto.Message, error) {
+ f.groupMessages = append(f.groupMessages, msg)
+ return &dto.Message{}, f.groupErr
+}
+
+func (f *fakeQQAPI) PostC2CMessage(
+ _ context.Context,
+ _ string,
+ msg dto.APIMessage,
+ _ ...options.Option,
+) (*dto.Message, error) {
+ f.c2cMessages = append(f.c2cMessages, msg)
+ return &dto.Message{}, f.c2cErr
+}
+
+func (f *fakeQQAPI) Transport(_ context.Context, method, url string, body any) ([]byte, error) {
+ upload, ok := body.(*qqMediaUpload)
+ if !ok {
+ return nil, errors.New("unexpected transport body type")
+ }
+ f.transportCalls = append(f.transportCalls, fakeTransportCall{
+ method: method,
+ url: url,
+ body: *upload,
+ })
+ return f.transportResp, f.transportErr
+}
+
+func mustJSON(t *testing.T, v any) []byte {
+ t.Helper()
+
+ b, err := json.Marshal(v)
+ if err != nil {
+ t.Fatalf("json.Marshal() error = %v", err)
+ }
+ return b
+}
+
+func waitInboundMessage(t *testing.T, messageBus *bus.MessageBus) bus.InboundMessage {
+ t.Helper()
+
+ ctx, cancel := context.WithTimeout(context.Background(), time.Second)
+ defer cancel()
+
+ for {
+ select {
+ case <-ctx.Done():
+ t.Fatal("timeout waiting for inbound message")
+ case inbound, ok := <-messageBus.InboundChan():
+ if !ok {
+ t.Fatal("expected inbound message")
+ }
+ return inbound
+ }
+ }
+}
+
+func writeTempFile(t *testing.T, dir, name string, content []byte) string {
+ t.Helper()
+
+ path := dir + "/" + name
+ if err := os.WriteFile(path, content, 0o600); err != nil {
+ t.Fatalf("WriteFile() error = %v", err)
+ }
+ return path
+}
diff --git a/pkg/channels/telegram/parse_markdown_to_md_v2.go b/pkg/channels/telegram/parse_markdown_to_md_v2.go
new file mode 100644
index 000000000..8cae312c5
--- /dev/null
+++ b/pkg/channels/telegram/parse_markdown_to_md_v2.go
@@ -0,0 +1,197 @@
+package telegram
+
+import (
+ "regexp"
+ "strings"
+)
+
+// mdV2SpecialChars are all characters that must be escaped in Telegram MarkdownV2
+var mdV2SpecialChars = map[rune]bool{
+ '*': true,
+ '_': true,
+ '[': true,
+ ']': true,
+ '(': true,
+ ')': true,
+ '~': true,
+ '`': true,
+ '>': true,
+ '<': true,
+ '#': true,
+ '+': true,
+ '-': true,
+ '=': true,
+ '|': true,
+ '{': true,
+ '}': true,
+ '.': true,
+ '!': true,
+ '\\': true,
+}
+
+// entityPattern describes one Telegram MarkdownV2 inline entity type.
+type entityPattern struct {
+ re *regexp.Regexp
+ open string
+ close string
+}
+
+// allEntityPatterns lists every recognized entity in priority order
+// (longer / more-specific delimiters first so they win over shorter ones).
+// Each entry's regex is anchored to find the first occurrence in a string.
+var allEntityPatterns = []entityPattern{
+ // fenced code block — content is completely verbatim
+ {re: regexp.MustCompile("(?s)```(?:[\\w]*\\n)?[\\s\\S]*?```"), open: "```", close: "```"},
+ // inline code — content is completely verbatim
+ {re: regexp.MustCompile("`(?:[^`\\\n]|\\\\.)*`"), open: "`", close: "`"},
+ // expandable block-quote opener **>…
+ {re: regexp.MustCompile(`(?m)\*\*>(?:[^\n]*)`), open: "**>", close: ""},
+ // block-quote line >…
+ {re: regexp.MustCompile(`(?m)^>(?:[^\n]*)`), open: ">", close: ""},
+ // custom emoji / timestamp  — must come before plain link
+ {re: regexp.MustCompile(`!\[[^\]]*\]\([^)]*\)`), open: "!", close: ""},
+ // inline URL / user mention […](…)
+ {re: regexp.MustCompile(`\[[^\]]*\]\([^)]*\)`), open: "[", close: ""},
+ // spoiler ||…|| — before single | so it wins
+ {re: regexp.MustCompile(`\|\|(?:[^|\\\n]|\\.)*\|\|`), open: "||", close: "||"},
+ // underline __…__ — before single _ so it wins
+ {re: regexp.MustCompile(`__(?:[^_\\\n]|\\.)*__`), open: "__", close: "__"},
+ // bold *…*
+ {re: regexp.MustCompile(`\*(?:[^*\\\n]|\\.)*\*`), open: "*", close: "*"},
+ // italic _…_
+ {re: regexp.MustCompile(`_(?:[^_\\\n]|\\.)*_`), open: "_", close: "_"},
+ // strikethrough ~…~
+ {re: regexp.MustCompile(`~(?:[^~\\\n]|\\.)*~`), open: "~", close: "~"},
+}
+
+// verbatimEntities are entity types whose inner content must never be
+// touched (code blocks, URLs, quotes, custom emoji).
+// Their content is passed through completely unchanged.
+var verbatimEntities = map[string]bool{
+ "```": true,
+ "`": true,
+ "**>": true,
+ ">": true,
+ "!": true,
+ "[": true,
+}
+
+// markdownToTelegramMarkdownV2 converts a Markdown string into a string safe
+// for sending with Telegram's MarkdownV2 parse mode.
+//
+// Rules:
+// - Markdown headings (# … ######) are converted to *bold*.
+// - **bold** Markdown syntax is converted to *bold*.
+// - Recognized Telegram MarkdownV2 entity spans are preserved; their inner
+// content is processed recursively so that nested valid entities are kept
+// intact while stray special characters are escaped.
+// - All plain-text segments have their MarkdownV2 special characters escaped.
+//
+// Reference: https://core.telegram.org/bots/api#formatting-options
+func markdownToTelegramMarkdownV2(text string) string {
+ // 1. Convert Markdown headings → *escaped heading text*
+ text = reHeading.ReplaceAllStringFunc(text, func(match string) string {
+ sub := reHeading.FindStringSubmatch(match)
+ if len(sub) < 2 {
+ return match
+ }
+ // The heading content is fresh plain text — escape everything
+ // including * so the resulting *…* bold span stays valid.
+ return "*" + escapeMarkdownV2(sub[1]) + "*"
+ })
+
+ // 2. Convert **bold** → *bold*
+ text = reBoldStar.ReplaceAllString(text, "*$1*")
+
+ // 3. Recursively escape the full string.
+ return processText(text)
+}
+
+// processText walks `text`, finds the leftmost / longest matching entity,
+// escapes the gap before it, processes the entity (recursing into its inner
+// content when appropriate), then continues with the remainder.
+func processText(text string) string {
+ if text == "" {
+ return ""
+ }
+
+ // Find the leftmost match among all entity patterns.
+ bestStart := -1
+ bestEnd := -1
+ var bestPat *entityPattern
+
+ for i := range allEntityPatterns {
+ p := &allEntityPatterns[i]
+ loc := p.re.FindStringIndex(text)
+ if loc == nil {
+ continue
+ }
+ if bestStart == -1 || loc[0] < bestStart ||
+ (loc[0] == bestStart && (loc[1]-loc[0]) > (bestEnd-bestStart)) {
+ bestStart = loc[0]
+ bestEnd = loc[1]
+ bestPat = p
+ }
+ }
+
+ if bestPat == nil {
+ // No entity found — escape everything.
+ return escapeMarkdownV2(text)
+ }
+
+ var b strings.Builder
+
+ // Plain text before the entity.
+ if bestStart > 0 {
+ b.WriteString(escapeMarkdownV2(text[:bestStart]))
+ }
+
+ // The matched entity span.
+ matched := text[bestStart:bestEnd]
+
+ if verbatimEntities[bestPat.open] {
+ // Code blocks, URLs, quotes: pass through completely untouched.
+ b.WriteString(matched)
+ } else {
+ // Inline formatting (bold, italic, underline, strikethrough, spoiler):
+ // keep the delimiters and recursively process the inner content so that
+ // nested entities survive but stray specials get escaped.
+ openLen := len(bestPat.open)
+ closeLen := len(bestPat.close)
+ inner := matched[openLen : len(matched)-closeLen]
+
+ b.WriteString(bestPat.open)
+ b.WriteString(processText(inner))
+ b.WriteString(bestPat.close)
+ }
+
+ // Continue with the remainder of the string.
+ b.WriteString(processText(text[bestEnd:]))
+
+ return b.String()
+}
+
+// escapeMarkdownV2 escapes every MarkdownV2 special character in a plain-text
+// segment (i.e. a segment that is not part of any recognized entity).
+// Already-escaped sequences (backslash + char) are forwarded verbatim to avoid
+// double-escaping.
+func escapeMarkdownV2(s string) string {
+ var b strings.Builder
+ b.Grow(len(s) + 8)
+ runes := []rune(s)
+ for i := 0; i < len(runes); i++ {
+ ch := runes[i]
+ // Forward an existing escape sequence verbatim.
+ if ch == '\\' && i+1 < len(runes) {
+ b.WriteRune(ch)
+ b.WriteRune(runes[i+1])
+ i++
+ continue
+ }
+ if mdV2SpecialChars[ch] {
+ b.WriteByte('\\')
+ }
+ b.WriteRune(ch)
+ }
+ return b.String()
+}
diff --git a/pkg/channels/telegram/parse_markdown_to_md_v2_test.go b/pkg/channels/telegram/parse_markdown_to_md_v2_test.go
new file mode 100644
index 000000000..fd68a9b83
--- /dev/null
+++ b/pkg/channels/telegram/parse_markdown_to_md_v2_test.go
@@ -0,0 +1,68 @@
+package telegram
+
+import (
+ _ "embed"
+ "testing"
+
+ "github.com/stretchr/testify/require"
+)
+
+//go:embed testdata/md2_all_formats.txt
+var md2AllFormats string
+
+func Test_markdownToTelegramMarkdownV2(t *testing.T) {
+ cases := []struct {
+ name string
+ input string
+ expected string
+ }{
+ {
+ name: "heading -> bolding",
+ input: `## HeadingH2 #`,
+ expected: "*HeadingH2 \\#*",
+ },
+ {
+ name: "strikethrough",
+ input: "~strikethroughMD~",
+ expected: "~strikethroughMD~",
+ },
+ {
+ name: "inline URL",
+ input: "[inline URL](http://www.example.com/)",
+ expected: "[inline URL](http://www.example.com/)",
+ },
+ {
+ name: "all telegram formats",
+ input: md2AllFormats,
+ expected: md2AllFormats,
+ },
+ {
+ name: "empty",
+ input: "",
+ expected: "",
+ },
+ {
+ name: "one letter",
+ input: "o",
+ expected: "o",
+ },
+ {
+ name: "",
+ input: "*Last update: ~10 24h*",
+ expected: "*Last update: \\~10 24h*",
+ },
+ {
+ name: "",
+ input: "",
+ expected: "\\",
+ },
+ }
+
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ actual := markdownToTelegramMarkdownV2(tc.input)
+
+ require.EqualValues(t, tc.expected, actual)
+ })
+ }
+}
diff --git a/pkg/channels/telegram/parser_markdown_to_html.go b/pkg/channels/telegram/parser_markdown_to_html.go
new file mode 100644
index 000000000..bdaa51807
--- /dev/null
+++ b/pkg/channels/telegram/parser_markdown_to_html.go
@@ -0,0 +1,111 @@
+package telegram
+
+import (
+ "fmt"
+ "strings"
+)
+
+func markdownToTelegramHTML(text string) string {
+ if text == "" {
+ return ""
+ }
+
+ codeBlocks := extractCodeBlocks(text)
+ text = codeBlocks.text
+
+ inlineCodes := extractInlineCodes(text)
+ text = inlineCodes.text
+
+ text = reHeading.ReplaceAllString(text, "$1")
+
+ text = reBlockquote.ReplaceAllString(text, "$1")
+
+ text = escapeHTML(text)
+
+ text = reLink.ReplaceAllString(text, `$1`)
+
+ text = reBoldStar.ReplaceAllString(text, "$1")
+
+ text = reBoldUnder.ReplaceAllString(text, "$1")
+
+ text = reItalic.ReplaceAllStringFunc(text, func(s string) string {
+ match := reItalic.FindStringSubmatch(s)
+ if len(match) < 2 {
+ return s
+ }
+ return "" + match[1] + ""
+ })
+
+ text = reStrike.ReplaceAllString(text, "$1")
+
+ text = reListItem.ReplaceAllString(text, "• ")
+
+ for i, code := range inlineCodes.codes {
+ escaped := escapeHTML(code)
+ text = strings.ReplaceAll(text, fmt.Sprintf("\x00IC%d\x00", i), fmt.Sprintf("%s", escaped))
+ }
+
+ for i, code := range codeBlocks.codes {
+ escaped := escapeHTML(code)
+ text = strings.ReplaceAll(
+ text,
+ fmt.Sprintf("\x00CB%d\x00", i),
+ fmt.Sprintf("
%s
", escaped),
+ )
+ }
+
+ return text
+}
+
+type codeBlockMatch struct {
+ text string
+ codes []string
+}
+
+func extractCodeBlocks(text string) codeBlockMatch {
+ matches := reCodeBlock.FindAllStringSubmatch(text, -1)
+
+ codes := make([]string, 0, len(matches))
+ for _, match := range matches {
+ codes = append(codes, match[1])
+ }
+
+ i := 0
+ text = reCodeBlock.ReplaceAllStringFunc(text, func(m string) string {
+ placeholder := fmt.Sprintf("\x00CB%d\x00", i)
+ i++
+ return placeholder
+ })
+
+ return codeBlockMatch{text: text, codes: codes}
+}
+
+type inlineCodeMatch struct {
+ text string
+ codes []string
+}
+
+func extractInlineCodes(text string) inlineCodeMatch {
+ matches := reInlineCode.FindAllStringSubmatch(text, -1)
+
+ codes := make([]string, 0, len(matches))
+ for _, match := range matches {
+ codes = append(codes, match[1])
+ }
+
+ i := 0
+ text = reInlineCode.ReplaceAllStringFunc(text, func(m string) string {
+ placeholder := fmt.Sprintf("\x00IC%d\x00", i)
+ i++
+ return placeholder
+ })
+
+ return inlineCodeMatch{text: text, codes: codes}
+}
+
+func escapeHTML(text string) string {
+ text = strings.ReplaceAll(text, "&", "&")
+ text = strings.ReplaceAll(text, "<", "<")
+ text = strings.ReplaceAll(text, ">", ">")
+ return text
+}
diff --git a/pkg/channels/telegram/telegram.go b/pkg/channels/telegram/telegram.go
index 34ee46b7b..3eb89c636 100644
--- a/pkg/channels/telegram/telegram.go
+++ b/pkg/channels/telegram/telegram.go
@@ -2,13 +2,17 @@ package telegram
import (
"context"
+ "crypto/rand"
+ "encoding/binary"
"fmt"
+ "io"
"net/http"
"net/url"
"os"
"regexp"
"strconv"
"strings"
+ "sync"
"time"
"github.com/mymmrac/telego"
@@ -26,7 +30,7 @@ import (
)
var (
- reHeading = regexp.MustCompile(`^#{1,6}\s+(.+)$`)
+ reHeading = regexp.MustCompile(`(?m)^#{1,6}\s+([^\n]+)`)
reBlockquote = regexp.MustCompile(`^>\s*(.*)$`)
reLink = regexp.MustCompile(`\[([^\]]+)\]\(([^)]+)\)`)
reBoldStar = regexp.MustCompile(`\*\*(.+?)\*\*`)
@@ -169,6 +173,8 @@ func (c *TelegramChannel) Send(ctx context.Context, msg bus.OutboundMessage) err
return channels.ErrNotRunning
}
+ useMarkdownV2 := c.config.Channels.Telegram.UseMarkdownV2
+
chatID, threadID, err := parseTelegramChatID(msg.ChatID)
if err != nil {
return fmt.Errorf("invalid chat ID %s: %w", msg.ChatID, channels.ErrSendFailed)
@@ -187,22 +193,65 @@ func (c *TelegramChannel) Send(ctx context.Context, msg bus.OutboundMessage) err
chunk := queue[0]
queue = queue[1:]
- htmlContent := markdownToTelegramHTML(chunk)
+ content := parseContent(chunk, useMarkdownV2)
- if len([]rune(htmlContent)) > 4096 {
- ratio := float64(len([]rune(chunk))) / float64(len([]rune(htmlContent)))
+ if len([]rune(content)) > 4096 {
+ runeChunk := []rune(chunk)
+ ratio := float64(len(runeChunk)) / float64(len([]rune(content)))
smallerLen := int(float64(4096) * ratio * 0.95) // 5% safety margin
- if smallerLen < 100 {
- smallerLen = 100
+
+ // Guarantee progress: if estimated length is >= chunk length, force it smaller
+ if smallerLen >= len(runeChunk) {
+ smallerLen = len(runeChunk) - 1
}
- // Push sub-chunks back to the front of the queue for
- // re-validation instead of sending them blindly.
+
+ if smallerLen <= 0 {
+ if err := c.sendChunk(ctx, sendChunkParams{
+ chatID: chatID,
+ threadID: threadID,
+ content: content,
+ replyToID: replyToID,
+ mdFallback: chunk,
+ useMarkdownV2: useMarkdownV2,
+ }); err != nil {
+ return err
+ }
+ replyToID = ""
+ continue
+ }
+
+ // Use the estimated smaller length as a guide for SplitMessage.
+ // SplitMessage will find natural break points (newlines/spaces) and respect code blocks.
subChunks := channels.SplitMessage(chunk, smallerLen)
- queue = append(subChunks, queue...)
+
+ // Safety fallback: If SplitMessage failed to shorten the chunk, force a manual hard split.
+ if len(subChunks) == 1 && subChunks[0] == chunk {
+ part1 := string(runeChunk[:smallerLen])
+ part2 := string(runeChunk[smallerLen:])
+ subChunks = []string{part1, part2}
+ }
+
+ // Filter out empty chunks to avoid sending empty messages to Telegram.
+ nonEmpty := make([]string, 0, len(subChunks))
+ for _, s := range subChunks {
+ if s != "" {
+ nonEmpty = append(nonEmpty, s)
+ }
+ }
+
+ // Push sub-chunks back to the front of the queue
+ queue = append(nonEmpty, queue...)
continue
}
- if err := c.sendHTMLChunk(ctx, chatID, threadID, htmlContent, chunk, replyToID); err != nil {
+ if err := c.sendChunk(ctx, sendChunkParams{
+ chatID: chatID,
+ threadID: threadID,
+ content: content,
+ replyToID: replyToID,
+ mdFallback: chunk,
+ useMarkdownV2: useMarkdownV2,
+ }); err != nil {
return err
}
// Only the first chunk should be a reply; subsequent chunks are normal messages.
@@ -212,17 +261,31 @@ func (c *TelegramChannel) Send(ctx context.Context, msg bus.OutboundMessage) err
return nil
}
-// sendHTMLChunk sends a single HTML message, falling back to the original
-// markdown as plain text on parse failure so users never see raw HTML tags.
-func (c *TelegramChannel) sendHTMLChunk(
- ctx context.Context, chatID int64, threadID int, htmlContent, mdFallback string, replyToID string,
-) error {
- tgMsg := tu.Message(tu.ID(chatID), htmlContent)
- tgMsg.ParseMode = telego.ModeHTML
- tgMsg.MessageThreadID = threadID
+type sendChunkParams struct {
+ chatID int64
+ threadID int
+ content string
+ replyToID string
+ mdFallback string
+ useMarkdownV2 bool
+}
- if replyToID != "" {
- if mid, parseErr := strconv.Atoi(replyToID); parseErr == nil {
+// sendChunk sends a single HTML/MarkdownV2 message, falling back to the original
+// markdown as plain text on parse failure so users never see raw HTML/MarkdownV2 tags.
+func (c *TelegramChannel) sendChunk(
+ ctx context.Context,
+ params sendChunkParams,
+) error {
+ tgMsg := tu.Message(tu.ID(params.chatID), params.content)
+ tgMsg.MessageThreadID = params.threadID
+ if params.useMarkdownV2 {
+ tgMsg.WithParseMode(telego.ModeMarkdownV2)
+ } else {
+ tgMsg.WithParseMode(telego.ModeHTML)
+ }
+
+ if params.replyToID != "" {
+ if mid, parseErr := strconv.Atoi(params.replyToID); parseErr == nil {
tgMsg.ReplyParameters = &telego.ReplyParameters{
MessageID: mid,
}
@@ -230,22 +293,29 @@ func (c *TelegramChannel) sendHTMLChunk(
}
if _, err := c.bot.SendMessage(ctx, tgMsg); err != nil {
- logger.ErrorCF("telegram", "HTML parse failed, falling back to plain text", map[string]any{
- "error": err.Error(),
- })
- tgMsg.Text = mdFallback
+ logParseFailed(err, params.useMarkdownV2)
+
+ tgMsg.Text = params.mdFallback
tgMsg.ParseMode = ""
if _, err = c.bot.SendMessage(ctx, tgMsg); err != nil {
return fmt.Errorf("telegram send: %w", channels.ErrTemporary)
}
}
+
return nil
}
+// maxTypingDuration limits how long the typing indicator can run.
+// Prevents endless typing when the LLM fails/hangs and preSend never invokes cancel.
+// Matches channels.Manager's typingStopTTL (5 min) so behavior is consistent.
+const maxTypingDuration = 5 * time.Minute
+
// StartTyping implements channels.TypingCapable.
// It sends ChatAction(typing) immediately and then repeats every 4 seconds
// (Telegram's typing indicator expires after ~5s) in a background goroutine.
// The returned stop function is idempotent and cancels the goroutine.
+// The goroutine also exits automatically after maxTypingDuration if cancel is
+// never called (e.g. when the LLM fails or times out without publishing).
func (c *TelegramChannel) StartTyping(ctx context.Context, chatID string) (func(), error) {
cid, threadID, err := parseTelegramChatID(chatID)
if err != nil {
@@ -259,12 +329,15 @@ func (c *TelegramChannel) StartTyping(ctx context.Context, chatID string) (func(
_ = c.bot.SendChatAction(ctx, action)
typingCtx, cancel := context.WithCancel(ctx)
+ // Cap lifetime so the goroutine cannot run indefinitely if cancel is never called
+ maxCtx, maxCancel := context.WithTimeout(typingCtx, maxTypingDuration)
go func() {
+ defer maxCancel()
ticker := time.NewTicker(4 * time.Second)
defer ticker.Stop()
for {
select {
- case <-typingCtx.Done():
+ case <-maxCtx.Done():
return
case <-ticker.C:
a := tu.ChatAction(tu.ID(cid), telego.ChatActionTyping)
@@ -279,6 +352,7 @@ func (c *TelegramChannel) StartTyping(ctx context.Context, chatID string) (func(
// EditMessage implements channels.MessageEditor.
func (c *TelegramChannel) EditMessage(ctx context.Context, chatID string, messageID string, content string) error {
+ useMarkdownV2 := c.config.Channels.Telegram.UseMarkdownV2
cid, _, err := parseTelegramChatID(chatID)
if err != nil {
return err
@@ -287,13 +361,38 @@ func (c *TelegramChannel) EditMessage(ctx context.Context, chatID string, messag
if err != nil {
return err
}
- htmlContent := markdownToTelegramHTML(content)
- editMsg := tu.EditMessageText(tu.ID(cid), mid, htmlContent)
- editMsg.ParseMode = telego.ModeHTML
+ parsedContent := parseContent(content, useMarkdownV2)
+ editMsg := tu.EditMessageText(tu.ID(cid), mid, parsedContent)
+ if useMarkdownV2 {
+ editMsg.WithParseMode(telego.ModeMarkdownV2)
+ } else {
+ editMsg.WithParseMode(telego.ModeHTML)
+ }
_, err = c.bot.EditMessageText(ctx, editMsg)
+ if err != nil {
+ logParseFailed(err, useMarkdownV2)
+ _, err = c.bot.EditMessageText(ctx, tu.EditMessageText(tu.ID(cid), mid, content))
+ }
+
return err
}
+// DeleteMessage implements channels.MessageDeleter.
+func (c *TelegramChannel) DeleteMessage(ctx context.Context, chatID string, messageID string) error {
+ cid, _, err := parseTelegramChatID(chatID)
+ if err != nil {
+ return err
+ }
+ mid, err := strconv.Atoi(messageID)
+ if err != nil {
+ return err
+ }
+ return c.bot.DeleteMessage(ctx, &telego.DeleteMessageParams{
+ ChatID: tu.ID(cid),
+ MessageID: mid,
+ })
+}
+
// SendPlaceholder implements channels.PlaceholderCapable.
// It sends a placeholder message (e.g. "Thinking... 💭") that will later be
// edited to the actual response via EditMessage (channels.MessageEditor).
@@ -367,6 +466,20 @@ func (c *TelegramChannel) SendMedia(ctx context.Context, msg bus.OutboundMediaMe
Caption: part.Caption,
}
_, err = c.bot.SendPhoto(ctx, params)
+ if err != nil && strings.Contains(err.Error(), "PHOTO_INVALID_DIMENSIONS") {
+ if _, seekErr := file.Seek(0, io.SeekStart); seekErr != nil {
+ file.Close()
+ return fmt.Errorf("telegram rewind media after photo failure: %w", channels.ErrTemporary)
+ }
+
+ docParams := &telego.SendDocumentParams{
+ ChatID: tu.ID(chatID),
+ MessageThreadID: threadID,
+ Document: telego.InputFile{File: file},
+ Caption: part.Caption,
+ }
+ _, err = c.bot.SendDocument(ctx, docParams)
+ }
case "audio":
params := &telego.SendAudioParams{
ChatID: tu.ID(chatID),
@@ -624,6 +737,14 @@ func (c *TelegramChannel) downloadFile(ctx context.Context, fileID, ext string)
return c.downloadFileWithInfo(file, ext)
}
+func parseContent(text string, useMarkdownV2 bool) string {
+ if useMarkdownV2 {
+ return markdownToTelegramMarkdownV2(text)
+ }
+
+ return markdownToTelegramHTML(text)
+}
+
// parseTelegramChatID splits "chatID/threadID" into its components.
// Returns threadID=0 when no "/" is present (non-forum messages).
func parseTelegramChatID(chatID string) (int64, int, error) {
@@ -643,109 +764,18 @@ func parseTelegramChatID(chatID string) (int64, int, error) {
return cid, tid, nil
}
-func markdownToTelegramHTML(text string) string {
- if text == "" {
- return ""
+func logParseFailed(err error, useMarkdownV2 bool) {
+ parsingName := "HTML"
+ if useMarkdownV2 {
+ parsingName = "MarkdownV2"
}
- codeBlocks := extractCodeBlocks(text)
- text = codeBlocks.text
-
- inlineCodes := extractInlineCodes(text)
- text = inlineCodes.text
-
- text = reHeading.ReplaceAllString(text, "$1")
-
- text = reBlockquote.ReplaceAllString(text, "$1")
-
- text = escapeHTML(text)
-
- text = reLink.ReplaceAllString(text, `$1`)
-
- text = reBoldStar.ReplaceAllString(text, "$1")
-
- text = reBoldUnder.ReplaceAllString(text, "$1")
-
- text = reItalic.ReplaceAllStringFunc(text, func(s string) string {
- match := reItalic.FindStringSubmatch(s)
- if len(match) < 2 {
- return s
- }
- return "" + match[1] + ""
- })
-
- text = reStrike.ReplaceAllString(text, "$1")
-
- text = reListItem.ReplaceAllString(text, "• ")
-
- for i, code := range inlineCodes.codes {
- escaped := escapeHTML(code)
- text = strings.ReplaceAll(text, fmt.Sprintf("\x00IC%d\x00", i), fmt.Sprintf("%s", escaped))
- }
-
- for i, code := range codeBlocks.codes {
- escaped := escapeHTML(code)
- text = strings.ReplaceAll(
- text,
- fmt.Sprintf("\x00CB%d\x00", i),
- fmt.Sprintf("
%s
", escaped),
- )
- }
-
- return text
-}
-
-type codeBlockMatch struct {
- text string
- codes []string
-}
-
-func extractCodeBlocks(text string) codeBlockMatch {
- matches := reCodeBlock.FindAllStringSubmatch(text, -1)
-
- codes := make([]string, 0, len(matches))
- for _, match := range matches {
- codes = append(codes, match[1])
- }
-
- i := 0
- text = reCodeBlock.ReplaceAllStringFunc(text, func(m string) string {
- placeholder := fmt.Sprintf("\x00CB%d\x00", i)
- i++
- return placeholder
- })
-
- return codeBlockMatch{text: text, codes: codes}
-}
-
-type inlineCodeMatch struct {
- text string
- codes []string
-}
-
-func extractInlineCodes(text string) inlineCodeMatch {
- matches := reInlineCode.FindAllStringSubmatch(text, -1)
-
- codes := make([]string, 0, len(matches))
- for _, match := range matches {
- codes = append(codes, match[1])
- }
-
- i := 0
- text = reInlineCode.ReplaceAllStringFunc(text, func(m string) string {
- placeholder := fmt.Sprintf("\x00IC%d\x00", i)
- i++
- return placeholder
- })
-
- return inlineCodeMatch{text: text, codes: codes}
-}
-
-func escapeHTML(text string) string {
- text = strings.ReplaceAll(text, "&", "&")
- text = strings.ReplaceAll(text, "<", "<")
- text = strings.ReplaceAll(text, ">", ">")
- return text
+ logger.ErrorCF("telegram",
+ fmt.Sprintf("%s parse failed, falling back to plain text", parsingName),
+ map[string]any{
+ "error": err.Error(),
+ },
+ )
}
// isBotMentioned checks if the bot is mentioned in the message via entities.
@@ -836,3 +866,107 @@ func (c *TelegramChannel) stripBotMention(content string) string {
content = re.ReplaceAllString(content, "")
return strings.TrimSpace(content)
}
+
+// BeginStream implements channels.StreamingCapable.
+func (c *TelegramChannel) BeginStream(ctx context.Context, chatID string) (channels.Streamer, error) {
+ if !c.config.Channels.Telegram.Streaming.Enabled {
+ return nil, fmt.Errorf("streaming disabled in config")
+ }
+
+ cid, _, err := parseTelegramChatID(chatID)
+ if err != nil {
+ return nil, err
+ }
+
+ streamCfg := c.config.Channels.Telegram.Streaming
+ return &telegramStreamer{
+ bot: c.bot,
+ chatID: cid,
+ draftID: cryptoRandInt(),
+ throttleInterval: time.Duration(streamCfg.ThrottleSeconds) * time.Second,
+ minGrowth: streamCfg.MinGrowthChars,
+ }, nil
+}
+
+// telegramStreamer streams partial LLM output via Telegram's sendMessageDraft API.
+// On first API error (e.g. bot lacks forum mode), it silently degrades: Update
+// becomes a no-op, while Finalize still delivers the final message.
+type telegramStreamer struct {
+ bot *telego.Bot
+ chatID int64
+ draftID int
+ throttleInterval time.Duration
+ minGrowth int
+ lastLen int
+ lastAt time.Time
+ failed bool
+ mu sync.Mutex
+}
+
+func (s *telegramStreamer) Update(ctx context.Context, content string) error {
+ s.mu.Lock()
+ defer s.mu.Unlock()
+
+ if s.failed {
+ return nil
+ }
+
+ // Throttle: skip if not enough time or content has passed
+ now := time.Now()
+ growth := len(content) - s.lastLen
+ if s.lastLen > 0 && now.Sub(s.lastAt) < s.throttleInterval && growth < s.minGrowth {
+ return nil
+ }
+
+ htmlContent := markdownToTelegramHTML(content)
+
+ err := s.bot.SendMessageDraft(ctx, &telego.SendMessageDraftParams{
+ ChatID: s.chatID,
+ DraftID: s.draftID,
+ Text: htmlContent,
+ ParseMode: telego.ModeHTML,
+ })
+ if err != nil {
+ // First error → degrade silently (e.g. no forum mode)
+ logger.WarnCF("telegram", "sendMessageDraft failed, disabling streaming", map[string]any{
+ "error": err.Error(),
+ })
+ s.failed = true
+ return nil // don't propagate — Finalize will still deliver
+ }
+
+ s.lastLen = len(content)
+ s.lastAt = now
+ return nil
+}
+
+func (s *telegramStreamer) Finalize(ctx context.Context, content string) error {
+ htmlContent := markdownToTelegramHTML(content)
+ tgMsg := tu.Message(tu.ID(s.chatID), htmlContent)
+ tgMsg.ParseMode = telego.ModeHTML
+
+ if _, err := s.bot.SendMessage(ctx, tgMsg); err != nil {
+ // Fallback to plain text
+ tgMsg.ParseMode = ""
+ if _, err = s.bot.SendMessage(ctx, tgMsg); err != nil {
+ logger.ErrorCF("telegram", "Finalize failed after HTML and plain-text attempts", map[string]any{
+ "chat_id": s.chatID,
+ "error": err.Error(),
+ "len": len(content),
+ })
+ return fmt.Errorf("telegram finalize: %w", err)
+ }
+ }
+ return nil
+}
+
+func (s *telegramStreamer) Cancel(ctx context.Context) {
+ // Draft auto-expires on Telegram's side; nothing to clean up.
+}
+
+// cryptoRandInt returns a non-zero random int using crypto/rand.
+func cryptoRandInt() int {
+ var b [4]byte
+ _, _ = rand.Read(b[:])
+ return int(binary.BigEndian.Uint32(b[:])) | 1 // ensure non-zero
+}
diff --git a/pkg/channels/telegram/telegram_dispatch_test.go b/pkg/channels/telegram/telegram_dispatch_test.go
index 1ea4a4824..0eb1de5ea 100644
--- a/pkg/channels/telegram/telegram_dispatch_test.go
+++ b/pkg/channels/telegram/telegram_dispatch_test.go
@@ -3,7 +3,6 @@ package telegram
import (
"context"
"testing"
- "time"
"github.com/mymmrac/telego"
@@ -36,10 +35,7 @@ func TestHandleMessage_DoesNotConsumeGenericCommandsLocally(t *testing.T) {
t.Fatalf("handleMessage error: %v", err)
}
- ctx, cancel := context.WithTimeout(context.Background(), time.Second)
- defer cancel()
-
- inbound, ok := messageBus.ConsumeInbound(ctx)
+ inbound, ok := <-messageBus.InboundChan()
if !ok {
t.Fatal("expected inbound message to be forwarded")
}
diff --git a/pkg/channels/telegram/telegram_group_command_filter_test.go b/pkg/channels/telegram/telegram_group_command_filter_test.go
index 0d5b985fe..614b2ca7f 100644
--- a/pkg/channels/telegram/telegram_group_command_filter_test.go
+++ b/pkg/channels/telegram/telegram_group_command_filter_test.go
@@ -108,22 +108,24 @@ func TestHandleMessage_GroupMentionOnly_BotCommandEntity(t *testing.T) {
t.Fatalf("handleMessage error: %v", err)
}
- ctx, cancel := context.WithTimeout(context.Background(), 150*time.Millisecond)
+ ctx, cancel := context.WithTimeout(context.Background(), 200*time.Microsecond)
defer cancel()
-
- inbound, ok := messageBus.ConsumeInbound(ctx)
- if tc.wantForwarded {
- if !ok {
- t.Fatal("expected inbound message to be forwarded")
+ select {
+ case <-ctx.Done():
+ if tc.wantForwarded {
+ t.Fatal("timeout waiting for message to be forwarded")
+ return
}
- if inbound.Content != tc.wantContent {
- t.Fatalf("content=%q want=%q", inbound.Content, tc.wantContent)
+ case inbound, ok := <-messageBus.InboundChan():
+ if tc.wantForwarded {
+ if !ok {
+ t.Fatal("expected inbound message to be forwarded")
+ }
+ if inbound.Content != tc.wantContent {
+ t.Fatalf("content=%q want=%q", inbound.Content, tc.wantContent)
+ }
+ return
}
- return
- }
-
- if ok {
- t.Fatalf("expected message to be filtered, got content=%q", inbound.Content)
}
})
}
diff --git a/pkg/channels/telegram/telegram_test.go b/pkg/channels/telegram/telegram_test.go
index c2186d0a3..6bf1077af 100644
--- a/pkg/channels/telegram/telegram_test.go
+++ b/pkg/channels/telegram/telegram_test.go
@@ -4,9 +4,11 @@ import (
"context"
"encoding/json"
"errors"
+ "io"
+ "os"
+ "path/filepath"
"strings"
"testing"
- "time"
"github.com/mymmrac/telego"
ta "github.com/mymmrac/telego/telegoapi"
@@ -15,6 +17,8 @@ import (
"github.com/sipeed/picoclaw/pkg/bus"
"github.com/sipeed/picoclaw/pkg/channels"
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/media"
)
const testToken = "1234567890:aaaabbbbaaaabbbbaaaabbbbaaaabbbbccc"
@@ -38,8 +42,20 @@ func (s *stubCaller) Call(ctx context.Context, url string, data *ta.RequestData)
// stubConstructor implements ta.RequestConstructor for testing.
type stubConstructor struct{}
+type multipartCall struct {
+ Parameters map[string]string
+ FileSizes map[string]int
+}
+
func (s *stubConstructor) JSONRequest(parameters any) (*ta.RequestData, error) {
- return &ta.RequestData{}, nil
+ b, err := json.Marshal(parameters)
+ if err != nil {
+ return nil, err
+ }
+ return &ta.RequestData{
+ ContentType: "application/json",
+ BodyRaw: b,
+ }, nil
}
func (s *stubConstructor) MultipartRequest(
@@ -49,6 +65,36 @@ func (s *stubConstructor) MultipartRequest(
return &ta.RequestData{}, nil
}
+type multipartRecordingConstructor struct {
+ stubConstructor
+ calls []multipartCall
+}
+
+func (s *multipartRecordingConstructor) MultipartRequest(
+ parameters map[string]string,
+ files map[string]ta.NamedReader,
+) (*ta.RequestData, error) {
+ call := multipartCall{
+ Parameters: make(map[string]string, len(parameters)),
+ FileSizes: make(map[string]int, len(files)),
+ }
+ for k, v := range parameters {
+ call.Parameters[k] = v
+ }
+ for field, file := range files {
+ if file == nil {
+ continue
+ }
+ data, err := io.ReadAll(file)
+ if err != nil {
+ return nil, err
+ }
+ call.FileSizes[field] = len(data)
+ }
+ s.calls = append(s.calls, call)
+ return &ta.RequestData{}, nil
+}
+
// successResponse returns a ta.Response that telego will treat as a successful SendMessage.
func successResponse(t *testing.T) *ta.Response {
t.Helper()
@@ -60,11 +106,19 @@ func successResponse(t *testing.T) *ta.Response {
// newTestChannel creates a TelegramChannel with a mocked bot for unit testing.
func newTestChannel(t *testing.T, caller *stubCaller) *TelegramChannel {
+ return newTestChannelWithConstructor(t, caller, &stubConstructor{})
+}
+
+func newTestChannelWithConstructor(
+ t *testing.T,
+ caller *stubCaller,
+ constructor ta.RequestConstructor,
+) *TelegramChannel {
t.Helper()
bot, err := telego.NewBot(testToken,
telego.WithAPICaller(caller),
- telego.WithRequestConstructor(&stubConstructor{}),
+ telego.WithRequestConstructor(constructor),
telego.WithDiscardLogger(),
)
require.NoError(t, err)
@@ -78,9 +132,96 @@ func newTestChannel(t *testing.T, caller *stubCaller) *TelegramChannel {
BaseChannel: base,
bot: bot,
chatIDs: make(map[string]int64),
+ config: config.DefaultConfig(),
}
}
+func TestSendMedia_ImageFallbacksToDocumentOnInvalidDimensions(t *testing.T) {
+ constructor := &multipartRecordingConstructor{}
+ caller := &stubCaller{
+ callFn: func(ctx context.Context, url string, data *ta.RequestData) (*ta.Response, error) {
+ switch {
+ case strings.Contains(url, "sendPhoto"):
+ return nil, errors.New(`api: 400 "Bad Request: PHOTO_INVALID_DIMENSIONS"`)
+ case strings.Contains(url, "sendDocument"):
+ return successResponse(t), nil
+ default:
+ t.Fatalf("unexpected API call: %s", url)
+ return nil, nil
+ }
+ },
+ }
+ ch := newTestChannelWithConstructor(t, caller, constructor)
+
+ store := media.NewFileMediaStore()
+ ch.SetMediaStore(store)
+
+ tmpDir := t.TempDir()
+ localPath := filepath.Join(tmpDir, "woodstock-en-10s.png")
+ content := []byte("fake-png-content")
+ require.NoError(t, os.WriteFile(localPath, content, 0o644))
+
+ ref, err := store.Store(
+ localPath,
+ media.MediaMeta{Filename: "woodstock-en-10s.png", ContentType: "image/png"},
+ "scope-1",
+ )
+ require.NoError(t, err)
+
+ err = ch.SendMedia(context.Background(), bus.OutboundMediaMessage{
+ ChatID: "12345",
+ Parts: []bus.MediaPart{{
+ Type: "image",
+ Ref: ref,
+ Caption: "caption",
+ }},
+ })
+
+ require.NoError(t, err)
+ require.Len(t, caller.calls, 2)
+ assert.Contains(t, caller.calls[0].URL, "sendPhoto")
+ assert.Contains(t, caller.calls[1].URL, "sendDocument")
+ require.Len(t, constructor.calls, 2)
+ assert.Equal(t, len(content), constructor.calls[0].FileSizes["photo"])
+ assert.Equal(t, len(content), constructor.calls[1].FileSizes["document"])
+ assert.Equal(t, "caption", constructor.calls[1].Parameters["caption"])
+}
+
+func TestSendMedia_ImageNonDimensionErrorDoesNotFallback(t *testing.T) {
+ constructor := &multipartRecordingConstructor{}
+ caller := &stubCaller{
+ callFn: func(ctx context.Context, url string, data *ta.RequestData) (*ta.Response, error) {
+ return nil, errors.New("api: 500 \"server exploded\"")
+ },
+ }
+ ch := newTestChannelWithConstructor(t, caller, constructor)
+
+ store := media.NewFileMediaStore()
+ ch.SetMediaStore(store)
+
+ tmpDir := t.TempDir()
+ localPath := filepath.Join(tmpDir, "image.png")
+ require.NoError(t, os.WriteFile(localPath, []byte("fake-png-content"), 0o644))
+
+ ref, err := store.Store(localPath, media.MediaMeta{Filename: "image.png", ContentType: "image/png"}, "scope-1")
+ require.NoError(t, err)
+
+ err = ch.SendMedia(context.Background(), bus.OutboundMediaMessage{
+ ChatID: "12345",
+ Parts: []bus.MediaPart{{
+ Type: "image",
+ Ref: ref,
+ }},
+ })
+
+ require.Error(t, err)
+ assert.ErrorIs(t, err, channels.ErrTemporary)
+ require.Len(t, caller.calls, 1)
+ assert.Contains(t, caller.calls[0].URL, "sendPhoto")
+ require.Len(t, constructor.calls, 1)
+ assert.NotContains(t, caller.calls[0].URL, "sendDocument")
+}
+
func TestSend_EmptyContent(t *testing.T) {
caller := &stubCaller{
callFn: func(ctx context.Context, url string, data *ta.RequestData) (*ta.Response, error) {
@@ -235,6 +376,55 @@ func TestSend_MarkdownShortButHTMLLong_MultipleCalls(t *testing.T) {
)
}
+func TestSend_HTMLOverflow_WordBoundary(t *testing.T) {
+ caller := &stubCaller{
+ callFn: func(ctx context.Context, url string, data *ta.RequestData) (*ta.Response, error) {
+ return successResponse(t), nil
+ },
+ }
+ ch := newTestChannel(t, caller)
+
+ // We want to force a split near index ~2600 while keeping markdown length <= 4000.
+ // Prefix of 430 bold units (6 chars each) = 2580 chars.
+ // Expansion per unit is +3 chars when converted to HTML, so 2580 + 430*3 = 3870.
+ prefix := strings.Repeat("**a** ", 430)
+ targetWord := "TARGETWORDTHATSTAYSTOGETHER"
+ // Suffix of 230 bold units (6 chars each) = 1380 chars.
+ // Total markdown length: 2580 (prefix) + 27 (target word) + 1380 (suffix) = 3987 <= 4000.
+ // HTML expansion adds ~3 chars per bold unit: (430 + 230)*3 = 1980 extra chars,
+ // so total HTML length comfortably exceeds 4096.
+ suffix := strings.Repeat(" **b**", 230)
+ content := prefix + targetWord + suffix
+
+ // Ensure the test content matches the intended boundary conditions.
+ assert.LessOrEqual(t, len([]rune(content)), 4000, "markdown content must not exceed chunk size for this test")
+
+ err := ch.Send(context.Background(), bus.OutboundMessage{
+ ChatID: "123456",
+ Content: content,
+ })
+
+ assert.NoError(t, err)
+
+ foundFullWord := false
+ for i, call := range caller.calls {
+ var params map[string]any
+ err := json.Unmarshal(call.Data.BodyRaw, ¶ms)
+ require.NoError(t, err)
+ text, _ := params["text"].(string)
+
+ hasWord := strings.Contains(text, targetWord)
+ t.Logf("Chunk %d length: %d, contains target word: %v", i, len(text), hasWord)
+
+ if hasWord {
+ foundFullWord = true
+ break
+ }
+ }
+
+ assert.True(t, foundFullWord, "The target word should not be split between chunks")
+}
+
func TestSend_NotRunning(t *testing.T) {
caller := &stubCaller{
callFn: func(ctx context.Context, url string, data *ta.RequestData) (*ta.Response, error) {
@@ -355,10 +545,7 @@ func TestHandleMessage_ForumTopic_SetsMetadata(t *testing.T) {
err := ch.handleMessage(context.Background(), msg)
require.NoError(t, err)
- ctx, cancel := context.WithTimeout(context.Background(), time.Second)
- defer cancel()
-
- inbound, ok := messageBus.ConsumeInbound(ctx)
+ inbound, ok := <-messageBus.InboundChan()
require.True(t, ok, "expected inbound message")
// Composite chatID should include thread ID
@@ -397,10 +584,7 @@ func TestHandleMessage_NoForum_NoThreadMetadata(t *testing.T) {
err := ch.handleMessage(context.Background(), msg)
require.NoError(t, err)
- ctx, cancel := context.WithTimeout(context.Background(), time.Second)
- defer cancel()
-
- inbound, ok := messageBus.ConsumeInbound(ctx)
+ inbound, ok := <-messageBus.InboundChan()
require.True(t, ok)
// Plain chatID without thread suffix
@@ -443,10 +627,7 @@ func TestHandleMessage_ReplyThread_NonForum_NoIsolation(t *testing.T) {
err := ch.handleMessage(context.Background(), msg)
require.NoError(t, err)
- ctx, cancel := context.WithTimeout(context.Background(), time.Second)
- defer cancel()
-
- inbound, ok := messageBus.ConsumeInbound(ctx)
+ inbound, ok := <-messageBus.InboundChan()
require.True(t, ok)
// chatID should NOT include thread suffix for non-forum groups
diff --git a/pkg/channels/telegram/testdata/md2_all_formats.txt b/pkg/channels/telegram/testdata/md2_all_formats.txt
new file mode 100644
index 000000000..f78fcc72f
--- /dev/null
+++ b/pkg/channels/telegram/testdata/md2_all_formats.txt
@@ -0,0 +1,31 @@
+*bold \*text*
+_italic \*text_
+__underline__
+~strikethrough~
+||spoiler||
+*bold _italic bold ~italic bold strikethrough ||italic bold strikethrough spoiler||~ __underline italic bold___ bold*
+[inline URL](http://www.example.com/)
+[inline mention of a user](tg://user?id=123456789)
+
+
+
+
+
+`inline fixed-width code`
+```
+pre-formatted fixed-width code block
+```
+```python
+pre-formatted fixed-width code block written in the Python programming language
+```
+>Block quotation started
+>Block quotation continued
+>Block quotation continued
+>Block quotation continued
+>The last line of the block quotation
+**>The expandable block quotation started right after the previous block quotation
+>It is separated from the previous block quotation by an empty bold entity
+>Expandable block quotation continued
+>Hidden by default part of the expandable block quotation started
+>Expandable block quotation continued
+>The last line of the expandable block quotation with the expandability mark||
diff --git a/pkg/channels/wecom/aibot.go b/pkg/channels/wecom/aibot.go
index 93fe8c36d..2264b8492 100644
--- a/pkg/channels/wecom/aibot.go
+++ b/pkg/channels/wecom/aibot.go
@@ -22,6 +22,10 @@ import (
"github.com/sipeed/picoclaw/pkg/utils"
)
+// responseURLHTTPClient is a shared HTTP client for posting to WeCom response_url.
+// Reusing it enables connection pooling across replies.
+var responseURLHTTPClient = &http.Client{Timeout: 15 * time.Second}
+
// WeComAIBotChannel implements the Channel interface for WeCom AI Bot (企业微信智能机器人)
type WeComAIBotChannel struct {
*channels.BaseChannel
@@ -134,13 +138,28 @@ type WeComAIBotEncryptedResponse struct {
Nonce string `json:"nonce"`
}
-// NewWeComAIBotChannel creates a new WeCom AI Bot channel instance
+// NewWeComAIBotChannel creates a WeCom AI Bot channel instance.
+// If cfg.BotID and cfg.Secret are both set, it returns a WeComAIBotWSChannel
+// using the WebSocket long-connection API.
+// Otherwise it returns the webhook-mode WeComAIBotChannel (requires Token +
+// EncodingAESKey).
func NewWeComAIBotChannel(
cfg config.WeComAIBotConfig,
messageBus *bus.MessageBus,
-) (*WeComAIBotChannel, error) {
+) (channels.Channel, error) {
+ // WebSocket long-connection mode takes priority when BotID + Secret are set.
+ if cfg.BotID != "" && cfg.Secret != "" {
+ logger.InfoC("wecom_aibot", "BotID and Secret provided, using WebSocket mode")
+ return newWeComAIBotWSChannel(cfg, messageBus)
+ }
+ // Webhook (short-connection) mode.
if cfg.Token == "" || cfg.EncodingAESKey == "" {
- return nil, fmt.Errorf("token and encoding_aes_key are required for WeCom AI Bot")
+ return nil, fmt.Errorf(
+ "WeCom AI Bot requires either (bot_id + secret) for WebSocket mode " +
+ "or (token + encoding_aes_key) for webhook mode")
+ }
+ if cfg.ProcessingMessage == "" {
+ cfg.ProcessingMessage = config.DefaultWeComAIBotProcessingMessage
}
base := channels.NewBaseChannel("wecom_aibot", cfg, messageBus, cfg.AllowFrom,
@@ -693,7 +712,7 @@ func (c *WeComAIBotChannel) getStreamResponse(task *streamTask, timestamp, nonce
default:
if time.Now().After(task.Deadline) {
// Deadline reached: close the stream with a notice, then wait for agent via response_url.
- content = "⏳ Processing, please wait. The results will be sent shortly."
+ content = c.config.ProcessingMessage
finish = true
closeStreamOnly = true
logger.InfoCF(
@@ -782,8 +801,7 @@ func (c *WeComAIBotChannel) sendViaResponseURL(responseURL, content string) erro
}
req.Header.Set("Content-Type", "application/json; charset=utf-8")
- client := &http.Client{Timeout: 15 * time.Second}
- resp, err := client.Do(req)
+ resp, err := responseURLHTTPClient.Do(req)
if err != nil {
return fmt.Errorf("post to response_url failed: %w: %w", channels.ErrTemporary, err)
}
@@ -793,7 +811,8 @@ func (c *WeComAIBotChannel) sendViaResponseURL(responseURL, content string) erro
return nil
}
- respBody, err := io.ReadAll(resp.Body)
+ const maxErrBody = 64 << 10 // 64 KB is more than enough for any error response
+ respBody, err := io.ReadAll(io.LimitReader(resp.Body, maxErrBody))
if err != nil {
return fmt.Errorf("reading response_url body: %w: %w", channels.ErrTemporary, err)
}
@@ -895,17 +914,80 @@ func (c *WeComAIBotChannel) encryptMessage(plaintext, receiveid string) (string,
return base64.StdEncoding.EncodeToString(ciphertext), nil
}
-// generateStreamID generates a random stream ID
-func (c *WeComAIBotChannel) generateStreamID() string {
+// func (c *WeComAIBotChannel) downloadAndDecryptImage(
+// ctx context.Context,
+// imageURL string,
+// ) ([]byte, error) {
+// // Download image
+// req, err := http.NewRequestWithContext(ctx, http.MethodGet, imageURL, nil)
+// if err != nil {
+// return nil, fmt.Errorf("failed to create request: %w", err)
+// }
+
+// client := &http.Client{
+// Timeout: 15 * time.Second,
+// }
+
+// resp, err := client.Do(req)
+// if err != nil {
+// return nil, fmt.Errorf("failed to download image: %w", err)
+// }
+// defer resp.Body.Close()
+
+// if resp.StatusCode != http.StatusOK {
+// return nil, fmt.Errorf("download failed with status: %d", resp.StatusCode)
+// }
+
+// // Limit image download to 20 MB to prevent memory exhaustion
+// const maxImageSize = 20 << 20 // 20 MB
+// encryptedData, err := io.ReadAll(io.LimitReader(resp.Body, maxImageSize+1))
+// if err != nil {
+// return nil, fmt.Errorf("failed to read image data: %w", err)
+// }
+// if len(encryptedData) > maxImageSize {
+// return nil, fmt.Errorf("image too large (exceeds %d MB)", maxImageSize>>20)
+// }
+
+// logger.DebugCF("wecom_aibot", "Image downloaded", map[string]any{
+// "size": len(encryptedData),
+// })
+
+// // Decode AES key
+// aesKey, err := decodeWeComAESKey(c.config.EncodingAESKey)
+// if err != nil {
+// return nil, err
+// }
+
+// // Decrypt image (AES-CBC with IV = first 16 bytes of key, PKCS7 padding stripped)
+// decryptedData, err := decryptAESCBC(aesKey, encryptedData)
+// if err != nil {
+// return nil, fmt.Errorf("failed to decrypt image: %w", err)
+// }
+
+// logger.DebugCF("wecom_aibot", "Image decrypted", map[string]any{
+// "size": len(decryptedData),
+// })
+
+// return decryptedData, nil
+// }
+
+// generateRandomID generates a cryptographically random alphanumeric ID of
+// length n. Used for stream IDs and WebSocket request IDs.
+func generateRandomID(n int) string {
const letters = "abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789"
- b := make([]byte, 10)
+ b := make([]byte, n)
for i := range b {
- n, _ := rand.Int(rand.Reader, big.NewInt(int64(len(letters))))
- b[i] = letters[n.Int64()]
+ num, _ := rand.Int(rand.Reader, big.NewInt(int64(len(letters))))
+ b[i] = letters[num.Int64()]
}
return string(b)
}
+// generateStreamID generates a random 10-character stream ID (webhook mode).
+func (c *WeComAIBotChannel) generateStreamID() string {
+ return generateRandomID(10)
+}
+
// cleanupLoop periodically cleans up old streaming tasks
func (c *WeComAIBotChannel) cleanupLoop() {
ticker := time.NewTicker(5 * time.Minute)
diff --git a/pkg/channels/wecom/aibot_test.go b/pkg/channels/wecom/aibot_test.go
index 6f0664187..957b51c38 100644
--- a/pkg/channels/wecom/aibot_test.go
+++ b/pkg/channels/wecom/aibot_test.go
@@ -2,13 +2,18 @@ package wecom
import (
"context"
+ "encoding/json"
"testing"
+ "time"
"github.com/sipeed/picoclaw/pkg/bus"
+ "github.com/sipeed/picoclaw/pkg/channels"
"github.com/sipeed/picoclaw/pkg/config"
)
-func TestNewWeComAIBotChannel(t *testing.T) {
+// ---- Webhook mode tests ----
+
+func TestNewWeComAIBotChannel_WebhookMode(t *testing.T) {
t.Run("success with valid config", func(t *testing.T) {
cfg := config.WeComAIBotConfig{
Enabled: true,
@@ -22,14 +27,16 @@ func TestNewWeComAIBotChannel(t *testing.T) {
if err != nil {
t.Fatalf("Expected no error, got %v", err)
}
-
if ch == nil {
t.Fatal("Expected channel to be created")
}
-
if ch.Name() != "wecom_aibot" {
t.Errorf("Expected name 'wecom_aibot', got '%s'", ch.Name())
}
+ // Webhook mode must implement WebhookHandler.
+ if _, ok := ch.(channels.WebhookHandler); !ok {
+ t.Error("Webhook mode channel should implement WebhookHandler")
+ }
})
t.Run("error with missing token", func(t *testing.T) {
@@ -37,10 +44,8 @@ func TestNewWeComAIBotChannel(t *testing.T) {
Enabled: true,
EncodingAESKey: "testkey1234567890123456789012345678901234567",
}
-
messageBus := bus.NewMessageBus()
_, err := NewWeComAIBotChannel(cfg, messageBus)
-
if err == nil {
t.Fatal("Expected error for missing token, got nil")
}
@@ -51,17 +56,15 @@ func TestNewWeComAIBotChannel(t *testing.T) {
Enabled: true,
Token: "test_token",
}
-
messageBus := bus.NewMessageBus()
_, err := NewWeComAIBotChannel(cfg, messageBus)
-
if err == nil {
t.Fatal("Expected error for missing encoding key, got nil")
}
})
}
-func TestWeComAIBotChannelStartStop(t *testing.T) {
+func TestWeComAIBotWebhookChannelStartStop(t *testing.T) {
cfg := config.WeComAIBotConfig{
Enabled: true,
Token: "test_token",
@@ -76,22 +79,18 @@ func TestWeComAIBotChannelStartStop(t *testing.T) {
ctx := context.Background()
- // Test Start
if err := ch.Start(ctx); err != nil {
t.Fatalf("Failed to start channel: %v", err)
}
-
if !ch.IsRunning() {
- t.Error("Expected channel to be running")
+ t.Error("Expected channel to be running after Start")
}
- // Test Stop
if err := ch.Stop(ctx); err != nil {
t.Fatalf("Failed to stop channel: %v", err)
}
-
if ch.IsRunning() {
- t.Error("Expected channel to be stopped")
+ t.Error("Expected channel to be stopped after Stop")
}
}
@@ -102,13 +101,16 @@ func TestWeComAIBotChannelWebhookPath(t *testing.T) {
Token: "test_token",
EncodingAESKey: "testkey1234567890123456789012345678901234567",
}
-
messageBus := bus.NewMessageBus()
ch, _ := NewWeComAIBotChannel(cfg, messageBus)
+ wh, ok := ch.(channels.WebhookHandler)
+ if !ok {
+ t.Fatal("Expected channel to implement WebhookHandler")
+ }
expectedPath := "/webhook/wecom-aibot"
- if ch.WebhookPath() != expectedPath {
- t.Errorf("Expected webhook path '%s', got '%s'", expectedPath, ch.WebhookPath())
+ if wh.WebhookPath() != expectedPath {
+ t.Errorf("Expected webhook path '%s', got '%s'", expectedPath, wh.WebhookPath())
}
})
@@ -120,12 +122,96 @@ func TestWeComAIBotChannelWebhookPath(t *testing.T) {
EncodingAESKey: "testkey1234567890123456789012345678901234567",
WebhookPath: customPath,
}
-
messageBus := bus.NewMessageBus()
ch, _ := NewWeComAIBotChannel(cfg, messageBus)
- if ch.WebhookPath() != customPath {
- t.Errorf("Expected webhook path '%s', got '%s'", customPath, ch.WebhookPath())
+ wh, ok := ch.(channels.WebhookHandler)
+ if !ok {
+ t.Fatal("Expected channel to implement WebhookHandler")
+ }
+ if wh.WebhookPath() != customPath {
+ t.Errorf("Expected webhook path '%s', got '%s'", customPath, wh.WebhookPath())
+ }
+ })
+}
+
+func TestWeComAIBotChannelGetStreamResponseProcessingMessage(t *testing.T) {
+ validAESKey := "abcdefghijklmnopqrstuvwxyz0123456789ABCDEFG"
+
+ t.Run("uses default processing message", func(t *testing.T) {
+ cfg := config.WeComAIBotConfig{
+ Enabled: true,
+ Token: "test_token",
+ EncodingAESKey: validAESKey,
+ }
+
+ messageBus := bus.NewMessageBus()
+ channel, err := NewWeComAIBotChannel(cfg, messageBus)
+ if err != nil {
+ t.Fatalf("Failed to create channel: %v", err)
+ }
+ ch, ok := channel.(*WeComAIBotChannel)
+ if !ok {
+ t.Fatal("Expected webhook mode channel")
+ }
+
+ task := &streamTask{
+ StreamID: "stream-default",
+ ChatID: "chat-default",
+ Deadline: time.Now().Add(-time.Second),
+ }
+ ch.streamTasks[task.StreamID] = task
+ ch.chatTasks[task.ChatID] = []*streamTask{task}
+
+ resp := decodeStreamResponse(t, ch, ch.getStreamResponse(task, "1234567890", "nonce"))
+
+ if !resp.Stream.Finish {
+ t.Fatal("Expected finished stream response after deadline")
+ }
+ if resp.Stream.Content != config.DefaultWeComAIBotProcessingMessage {
+ t.Fatalf("Expected default processing message %q, got %q",
+ config.DefaultWeComAIBotProcessingMessage, resp.Stream.Content)
+ }
+ if !task.StreamClosed {
+ t.Fatal("Expected task stream to be marked closed")
+ }
+ if _, ok := ch.streamTasks[task.StreamID]; ok {
+ t.Fatal("Expected closed stream task to be removed from streamTasks")
+ }
+ if len(ch.chatTasks[task.ChatID]) != 1 {
+ t.Fatalf("Expected task to remain queued for response_url delivery, got %d entries",
+ len(ch.chatTasks[task.ChatID]))
+ }
+ })
+
+ t.Run("uses custom processing message", func(t *testing.T) {
+ cfg := config.WeComAIBotConfig{
+ Enabled: true,
+ Token: "test_token",
+ EncodingAESKey: validAESKey,
+ ProcessingMessage: "Please wait a moment. The result will be delivered in a follow-up message.",
+ }
+
+ messageBus := bus.NewMessageBus()
+ channel, err := NewWeComAIBotChannel(cfg, messageBus)
+ if err != nil {
+ t.Fatalf("Failed to create channel: %v", err)
+ }
+ ch, ok := channel.(*WeComAIBotChannel)
+ if !ok {
+ t.Fatal("Expected webhook mode channel")
+ }
+
+ task := &streamTask{
+ StreamID: "stream-custom",
+ ChatID: "chat-custom",
+ Deadline: time.Now().Add(-time.Second),
+ }
+
+ resp := decodeStreamResponse(t, ch, ch.getStreamResponse(task, "1234567890", "nonce"))
+
+ if resp.Stream.Content != cfg.ProcessingMessage {
+ t.Fatalf("Expected custom processing message %q, got %q", cfg.ProcessingMessage, resp.Stream.Content)
}
})
}
@@ -136,19 +222,19 @@ func TestGenerateStreamID(t *testing.T) {
Token: "test_token",
EncodingAESKey: "testkey1234567890123456789012345678901234567",
}
-
messageBus := bus.NewMessageBus()
ch, _ := NewWeComAIBotChannel(cfg, messageBus)
+ webhookCh, ok := ch.(*WeComAIBotChannel)
+ if !ok {
+ t.Fatal("Expected webhook mode channel")
+ }
- // Generate multiple IDs and check they are unique
ids := make(map[string]bool)
for i := 0; i < 100; i++ {
- id := ch.generateStreamID()
-
+ id := webhookCh.generateStreamID()
if len(id) != 10 {
t.Errorf("Expected stream ID length 10, got %d", len(id))
}
-
if ids[id] {
t.Errorf("Duplicate stream ID generated: %s", id)
}
@@ -157,35 +243,33 @@ func TestGenerateStreamID(t *testing.T) {
}
func TestEncryptDecrypt(t *testing.T) {
- // Use a valid 43-character base64 key (企业微信标准格式)
cfg := config.WeComAIBotConfig{
Enabled: true,
Token: "test_token",
EncodingAESKey: "abcdefghijklmnopqrstuvwxyz0123456789ABCDEFG", // 43 characters
}
-
messageBus := bus.NewMessageBus()
ch, _ := NewWeComAIBotChannel(cfg, messageBus)
+ webhookCh, ok := ch.(*WeComAIBotChannel)
+ if !ok {
+ t.Fatal("Expected webhook mode channel")
+ }
plaintext := "Hello, World!"
receiveid := ""
- // Encrypt
- encrypted, err := ch.encryptMessage(plaintext, receiveid)
+ encrypted, err := webhookCh.encryptMessage(plaintext, receiveid)
if err != nil {
t.Fatalf("Failed to encrypt message: %v", err)
}
-
if encrypted == "" {
t.Fatal("Encrypted message is empty")
}
- // Decrypt
decrypted, err := decryptMessageWithVerify(encrypted, cfg.EncodingAESKey, receiveid)
if err != nil {
t.Fatalf("Failed to decrypt message: %v", err)
}
-
if decrypted != plaintext {
t.Errorf("Expected decrypted message '%s', got '%s'", plaintext, decrypted)
}
@@ -198,13 +282,277 @@ func TestGenerateSignature(t *testing.T) {
encrypt := "encrypted_msg"
signature := computeSignature(token, timestamp, nonce, encrypt)
-
if signature == "" {
t.Error("Generated signature is empty")
}
-
- // Verify signature using verifySignature function
if !verifySignature(token, signature, timestamp, nonce, encrypt) {
t.Error("Generated signature does not verify correctly")
}
}
+
+func decodeStreamResponse(t *testing.T, ch *WeComAIBotChannel, encryptedResponse string) WeComAIBotStreamResponse {
+ t.Helper()
+
+ var wrapped WeComAIBotEncryptedResponse
+ if err := json.Unmarshal([]byte(encryptedResponse), &wrapped); err != nil {
+ t.Fatalf("Failed to unmarshal encrypted response: %v", err)
+ }
+
+ plaintext, err := decryptMessageWithVerify(wrapped.Encrypt, ch.config.EncodingAESKey, "")
+ if err != nil {
+ t.Fatalf("Failed to decrypt response: %v", err)
+ }
+
+ var resp WeComAIBotStreamResponse
+ if err := json.Unmarshal([]byte(plaintext), &resp); err != nil {
+ t.Fatalf("Failed to unmarshal decrypted response: %v", err)
+ }
+
+ return resp
+}
+
+// ---- WebSocket long-connection mode tests ----
+
+func TestNewWeComAIBotChannel_WSMode(t *testing.T) {
+ t.Run("success with bot_id and secret", func(t *testing.T) {
+ cfg := config.WeComAIBotConfig{
+ Enabled: true,
+ BotID: "test_bot_id",
+ Secret: "test_secret",
+ }
+ messageBus := bus.NewMessageBus()
+ ch, err := NewWeComAIBotChannel(cfg, messageBus)
+ if err != nil {
+ t.Fatalf("Expected no error, got %v", err)
+ }
+ if ch == nil {
+ t.Fatal("Expected channel to be created")
+ }
+ if ch.Name() != "wecom_aibot" {
+ t.Errorf("Expected name 'wecom_aibot', got '%s'", ch.Name())
+ }
+ // WebSocket mode must NOT implement WebhookHandler.
+ if _, ok := ch.(channels.WebhookHandler); ok {
+ t.Error("WebSocket mode channel should NOT implement WebhookHandler")
+ }
+ })
+
+ t.Run("ws mode takes priority over webhook fields", func(t *testing.T) {
+ cfg := config.WeComAIBotConfig{
+ Enabled: true,
+ BotID: "test_bot_id",
+ Secret: "test_secret",
+ Token: "also_set",
+ EncodingAESKey: "testkey1234567890123456789012345678901234567",
+ }
+ messageBus := bus.NewMessageBus()
+ ch, err := NewWeComAIBotChannel(cfg, messageBus)
+ if err != nil {
+ t.Fatalf("Expected no error, got %v", err)
+ }
+ if _, ok := ch.(*WeComAIBotWSChannel); !ok {
+ t.Error("Expected WebSocket mode channel when both BotID+Secret and Token+Key are set")
+ }
+ })
+
+ t.Run("error with missing bot_id", func(t *testing.T) {
+ cfg := config.WeComAIBotConfig{
+ Enabled: true,
+ Secret: "test_secret",
+ }
+ messageBus := bus.NewMessageBus()
+ _, err := NewWeComAIBotChannel(cfg, messageBus)
+ // Missing bot_id alone means neither WS mode nor webhook mode is fully configured.
+ if err == nil {
+ t.Fatal("Expected error for missing bot_id, got nil")
+ }
+ })
+
+ t.Run("error with missing secret", func(t *testing.T) {
+ cfg := config.WeComAIBotConfig{
+ Enabled: true,
+ BotID: "test_bot_id",
+ }
+ messageBus := bus.NewMessageBus()
+ _, err := NewWeComAIBotChannel(cfg, messageBus)
+ if err == nil {
+ t.Fatal("Expected error for missing secret, got nil")
+ }
+ })
+}
+
+func TestWeComAIBotWSChannelStartStop(t *testing.T) {
+ cfg := config.WeComAIBotConfig{
+ Enabled: true,
+ BotID: "test_bot_id",
+ Secret: "test_secret",
+ }
+ messageBus := bus.NewMessageBus()
+ ch, err := NewWeComAIBotChannel(cfg, messageBus)
+ if err != nil {
+ t.Fatalf("Failed to create channel: %v", err)
+ }
+
+ ctx := context.Background()
+
+ // Start launches a background goroutine; it should not block or return an error.
+ if err := ch.Start(ctx); err != nil {
+ t.Fatalf("Failed to start channel: %v", err)
+ }
+ if !ch.IsRunning() {
+ t.Error("Expected channel to be running after Start")
+ }
+
+ // Stop should work regardless of whether the WebSocket actually connected.
+ if err := ch.Stop(ctx); err != nil {
+ t.Fatalf("Failed to stop channel: %v", err)
+ }
+ if ch.IsRunning() {
+ t.Error("Expected channel to be stopped after Stop")
+ }
+}
+
+func TestGenerateRandomID(t *testing.T) {
+ ids := make(map[string]bool)
+ for i := 0; i < 200; i++ {
+ id := generateRandomID(10)
+ if len(id) != 10 {
+ t.Errorf("Expected ID length 10, got %d", len(id))
+ }
+ if ids[id] {
+ t.Errorf("Duplicate ID generated: %s", id)
+ }
+ ids[id] = true
+ }
+}
+
+func TestWSGenerateID(t *testing.T) {
+ ids := make(map[string]bool)
+ for i := 0; i < 200; i++ {
+ id := wsGenerateID()
+ if len(id) != 10 {
+ t.Errorf("Expected ID length 10, got %d", len(id))
+ }
+ if ids[id] {
+ t.Errorf("Duplicate wsGenerateID result: %s", id)
+ }
+ ids[id] = true
+ }
+}
+
+// ---- Webhook streaming fallback tests ----
+
+// makeWebhookChannel creates a started WeComAIBotChannel for testing.
+func makeWebhookChannel(t *testing.T) *WeComAIBotChannel {
+ t.Helper()
+ cfg := config.WeComAIBotConfig{
+ Enabled: true,
+ Token: "test_token",
+ EncodingAESKey: "abcdefghijklmnopqrstuvwxyz0123456789ABCDEFG",
+ }
+ ch, err := NewWeComAIBotChannel(cfg, bus.NewMessageBus())
+ if err != nil {
+ t.Fatalf("create channel: %v", err)
+ }
+ wc := ch.(*WeComAIBotChannel)
+ wc.ctx, wc.cancel = context.WithCancel(context.Background())
+ return wc
+}
+
+// makeStreamTask creates and registers a streamTask for testing.
+func makeStreamTask(t *testing.T, ch *WeComAIBotChannel, streamID, chatID string, deadline time.Time) *streamTask {
+ t.Helper()
+ task := &streamTask{
+ StreamID: streamID,
+ ChatID: chatID,
+ Deadline: deadline,
+ answerCh: make(chan string, 1),
+ }
+ task.ctx, task.cancel = context.WithCancel(ch.ctx)
+ ch.taskMu.Lock()
+ ch.streamTasks[streamID] = task
+ ch.chatTasks[chatID] = append(ch.chatTasks[chatID], task)
+ ch.taskMu.Unlock()
+ return task
+}
+
+// TestGetStreamResponse_ImmediateAnswer verifies that when the agent has already
+// placed its answer in answerCh, getStreamResponse returns a finish=true response
+// and fully removes the task.
+func TestGetStreamResponse_ImmediateAnswer(t *testing.T) {
+ ch := makeWebhookChannel(t)
+ defer ch.cancel()
+
+ task := makeStreamTask(t, ch, "stream-1", "chat-1", time.Now().Add(30*time.Second))
+ task.answerCh <- "hello from agent"
+
+ result := ch.getStreamResponse(task, "ts123", "nonce123")
+ if result == "" {
+ t.Fatal("expected non-empty encrypted response")
+ }
+
+ ch.taskMu.RLock()
+ _, exists := ch.streamTasks["stream-1"]
+ ch.taskMu.RUnlock()
+ if exists {
+ t.Error("task should have been removed from streamTasks after normal finish")
+ }
+ if !task.Finished {
+ t.Error("task.Finished should be true after normal finish")
+ }
+}
+
+// TestGetStreamResponse_DeadlinePassed verifies that when the stream deadline has
+// elapsed (no agent reply yet), getStreamResponse closes the stream but keeps the
+// task alive so the response_url fallback can still deliver the answer.
+func TestGetStreamResponse_DeadlinePassed(t *testing.T) {
+ ch := makeWebhookChannel(t)
+ defer ch.cancel()
+
+ task := makeStreamTask(t, ch, "stream-2", "chat-2", time.Now().Add(-time.Millisecond))
+
+ result := ch.getStreamResponse(task, "ts456", "nonce456")
+ if result == "" {
+ t.Fatal("expected non-empty encrypted response")
+ }
+
+ ch.taskMu.RLock()
+ _, stillStreaming := ch.streamTasks["stream-2"]
+ ch.taskMu.RUnlock()
+ if stillStreaming {
+ t.Error("task should have been removed from streamTasks after deadline")
+ }
+ if !task.StreamClosed {
+ t.Error("task.StreamClosed should be true after deadline")
+ }
+ if task.Finished {
+ t.Error("task.Finished must remain false: agent reply still expected via response_url")
+ }
+}
+
+// TestGetStreamResponse_StillPending verifies that when neither the agent has
+// replied nor the deadline has passed, getStreamResponse returns without altering
+// task state (client should poll again).
+func TestGetStreamResponse_StillPending(t *testing.T) {
+ ch := makeWebhookChannel(t)
+ defer ch.cancel()
+
+ task := makeStreamTask(t, ch, "stream-3", "chat-3", time.Now().Add(30*time.Second))
+
+ result := ch.getStreamResponse(task, "ts789", "nonce789")
+ if result == "" {
+ t.Fatal("expected non-empty encrypted response")
+ }
+
+ ch.taskMu.RLock()
+ _, exists := ch.streamTasks["stream-3"]
+ ch.taskMu.RUnlock()
+ if !exists {
+ t.Error("pending task should still be in streamTasks")
+ }
+ if task.Finished || task.StreamClosed {
+ t.Error("pending task should not be finished or stream-closed")
+ }
+ // Cleanup.
+ ch.removeTask(task)
+}
diff --git a/pkg/channels/wecom/aibot_ws.go b/pkg/channels/wecom/aibot_ws.go
new file mode 100644
index 000000000..830e763b9
--- /dev/null
+++ b/pkg/channels/wecom/aibot_ws.go
@@ -0,0 +1,1346 @@
+package wecom
+
+import (
+ "context"
+ "encoding/base64"
+ "encoding/json"
+ "fmt"
+ "io"
+ "net/http"
+ "os"
+ "path/filepath"
+ "strings"
+ "sync"
+ "time"
+
+ "github.com/gorilla/websocket"
+
+ "github.com/sipeed/picoclaw/pkg/bus"
+ "github.com/sipeed/picoclaw/pkg/channels"
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/identity"
+ "github.com/sipeed/picoclaw/pkg/logger"
+ "github.com/sipeed/picoclaw/pkg/media"
+ "github.com/sipeed/picoclaw/pkg/utils"
+)
+
+// Long-connection WebSocket endpoint.
+// Ref: https://developer.work.weixin.qq.com/document/path/101463
+const (
+ wsEndpoint = "wss://openws.work.weixin.qq.com"
+ wsHeartbeatInterval = 30 * time.Second
+ wsConnectTimeout = 15 * time.Second
+ wsSubscribeTimeout = 10 * time.Second
+ wsSendMsgTimeout = 10 * time.Second
+ wsRespondMsgTimeout = 10 * time.Second
+ wsWelcomeMsgTimeout = 5 * time.Second // WeCom requires welcome reply within 5 seconds
+ wsMaxReconnectWait = 60 * time.Second
+ wsInitialReconnect = time.Second
+
+ // WeCom requires finish=true within 6 minutes of the first stream frame.
+ // wsStreamTickInterval controls how often we send an in-progress hint.
+ // wsStreamMaxDuration is a safety margin below the 6-minute hard limit.
+ wsStreamTickInterval = 30 * time.Second
+ wsStreamMaxDuration = 5*time.Minute + 30*time.Second
+
+ // wsImageDownloadTimeout caps the time we spend downloading an inbound image.
+ wsImageDownloadTimeout = 30 * time.Second
+
+ // Keep req_id -> chat route for late fallback pushes after stream window closes.
+ wsLateReplyRouteTTL = 30 * time.Minute
+
+ // wsStreamMaxContentBytes is the maximum UTF-8 byte length for the content field
+ // of a single WeCom AI Bot stream / text / markdown frame.
+ // Ref: https://developer.work.weixin.qq.com/document/path/101463
+ wsStreamMaxContentBytes = 20480
+)
+
+// wsImageHTTPClient is a shared HTTP client for downloading inbound images.
+// Reusing it enables connection pooling across multiple image downloads.
+var wsImageHTTPClient = &http.Client{Timeout: wsImageDownloadTimeout}
+
+// WeComAIBotWSChannel implements channels.Channel for WeCom AI Bot using the
+// WebSocket long-connection API.
+// Unlike the webhook counterpart it does NOT implement WebhookHandler, so the
+// HTTP manager will not register any callback URL for it.
+type WeComAIBotWSChannel struct {
+ *channels.BaseChannel
+ config config.WeComAIBotConfig
+ ctx context.Context
+ cancel context.CancelFunc
+
+ // conn is the active WebSocket connection; nil when disconnected.
+ // All writes are serialized through connMu.
+ conn *websocket.Conn
+ connMu sync.Mutex
+
+ // dedupe prevents duplicate message processing (WeCom may re-deliver).
+ dedupe *MessageDeduplicator
+
+ // reqStates holds per-req_id runtime state.
+ // It unifies active task state and late-reply fallback routing.
+ reqStates map[string]*wsReqState
+ reqStatesMu sync.Mutex
+
+ // reqPending correlates command req_ids with response channels.
+ // Used only for subscribe/ping command-response pairs.
+ reqPending map[string]chan wsEnvelope
+ reqPendingMu sync.Mutex
+}
+
+// wsTask tracks one in-progress agent reply for a single chat turn.
+type wsTask struct {
+ ReqID string // req_id echoed in all replies for this turn
+ ChatID string
+ ChatType uint32
+ StreamID string // our generated stream.id
+ answerCh chan string // agent delivers its reply here via Send()
+ ctx context.Context
+ cancel context.CancelFunc
+}
+
+type wsReqState struct {
+ Task *wsTask
+ Route wsLateReplyRoute
+}
+
+type wsLateReplyRoute struct {
+ ChatID string
+ ChatType uint32
+ ReadyAt time.Time
+ ExpiresAt time.Time
+}
+
+// ---- WebSocket protocol types ----
+
+// wsEnvelope is the generic JSON envelope for all WebSocket messages.
+type wsEnvelope struct {
+ Cmd string `json:"cmd,omitempty"`
+ Headers wsHeaders `json:"headers"`
+ Body json.RawMessage `json:"body,omitempty"`
+ ErrCode int `json:"errcode,omitempty"`
+ ErrMsg string `json:"errmsg,omitempty"`
+}
+
+type wsHeaders struct {
+ ReqID string `json:"req_id"`
+}
+
+// wsCommand is an outgoing request sent over the WebSocket.
+type wsCommand struct {
+ Cmd string `json:"cmd"`
+ Headers wsHeaders `json:"headers"`
+ Body any `json:"body,omitempty"`
+}
+
+type wsSendMsgBody struct {
+ ChatID string `json:"chatid"`
+ ChatType uint32 `json:"chat_type,omitempty"`
+ MsgType string `json:"msgtype"`
+ Markdown *wsMarkdownContent `json:"markdown,omitempty"`
+}
+
+// wsRespondMsgBody is the body for aibot_respond_msg / aibot_respond_welcome_msg.
+type wsRespondMsgBody struct {
+ MsgType string `json:"msgtype"`
+ Stream *wsStreamContent `json:"stream,omitempty"`
+ Text *wsTextContent `json:"text,omitempty"`
+ Markdown *wsMarkdownContent `json:"markdown,omitempty"`
+ Image *wsImageContent `json:"image,omitempty"`
+}
+
+type wsStreamContent struct {
+ ID string `json:"id"`
+ Finish bool `json:"finish"`
+ Content string `json:"content,omitempty"`
+}
+
+// wsImageContent carries a base64-encoded image payload for outbound messages.
+type wsImageContent struct {
+ Base64 string `json:"base64"`
+ MD5 string `json:"md5"`
+}
+
+type wsTextContent struct {
+ Content string `json:"content"`
+}
+
+type wsMarkdownContent struct {
+ Content string `json:"content"`
+}
+
+// WeComAIBotWSMessage is the decoded body of aibot_msg_callback /
+// aibot_event_callback in WebSocket long-connection mode.
+// The structure mirrors WeComAIBotMessage but includes extra fields
+// that only appear in long-connection callbacks (Voice, AESKey on Image/File).
+type WeComAIBotWSMessage struct {
+ MsgID string `json:"msgid"`
+ CreateTime int64 `json:"create_time,omitempty"`
+ AIBotID string `json:"aibotid"`
+ ChatID string `json:"chatid,omitempty"`
+ ChatType string `json:"chattype,omitempty"` // "single" | "group"
+ From struct {
+ UserID string `json:"userid"`
+ } `json:"from"`
+ MsgType string `json:"msgtype"`
+ Text *struct {
+ Content string `json:"content"`
+ } `json:"text,omitempty"`
+ Image *struct {
+ URL string `json:"url"`
+ AESKey string `json:"aeskey,omitempty"` // long-connection: per-resource decrypt key
+ } `json:"image,omitempty"`
+ Voice *struct {
+ Content string `json:"content"` // WeCom transcribes voice to text in callbacks
+ } `json:"voice,omitempty"`
+ Mixed *struct {
+ MsgItem []struct {
+ MsgType string `json:"msgtype"`
+ Text *struct {
+ Content string `json:"content"`
+ } `json:"text,omitempty"`
+ Image *struct {
+ URL string `json:"url"`
+ AESKey string `json:"aeskey,omitempty"`
+ } `json:"image,omitempty"`
+ } `json:"msg_item"`
+ } `json:"mixed,omitempty"`
+ Event *struct {
+ EventType string `json:"eventtype"`
+ } `json:"event,omitempty"`
+ File *struct {
+ URL string `json:"url"`
+ AESKey string `json:"aeskey,omitempty"`
+ } `json:"file,omitempty"`
+ Video *struct {
+ URL string `json:"url"`
+ AESKey string `json:"aeskey,omitempty"`
+ } `json:"video,omitempty"`
+}
+
+// ---- Constructor ----
+
+// newWeComAIBotWSChannel creates a WeComAIBotWSChannel for WebSocket mode.
+func newWeComAIBotWSChannel(
+ cfg config.WeComAIBotConfig,
+ messageBus *bus.MessageBus,
+) (*WeComAIBotWSChannel, error) {
+ if cfg.BotID == "" || cfg.Secret == "" {
+ return nil, fmt.Errorf("bot_id and secret are required for WeCom AI Bot WebSocket mode")
+ }
+
+ base := channels.NewBaseChannel("wecom_aibot", cfg, messageBus, cfg.AllowFrom,
+ channels.WithReasoningChannelID(cfg.ReasoningChannelID),
+ )
+
+ return &WeComAIBotWSChannel{
+ BaseChannel: base,
+ config: cfg,
+ dedupe: NewMessageDeduplicator(wecomMaxProcessedMessages),
+ reqStates: make(map[string]*wsReqState),
+ reqPending: make(map[string]chan wsEnvelope),
+ }, nil
+}
+
+// ---- Channel interface ----
+
+// Name implements channels.Channel.
+func (c *WeComAIBotWSChannel) Name() string { return "wecom_aibot" }
+
+// Start connects to the WeCom WebSocket endpoint and begins message processing.
+func (c *WeComAIBotWSChannel) Start(ctx context.Context) error {
+ logger.InfoC("wecom_aibot", "Starting WeCom AI Bot channel (WebSocket long-connection mode)...")
+ c.ctx, c.cancel = context.WithCancel(ctx)
+ c.SetRunning(true)
+ go c.connectLoop()
+ logger.InfoC("wecom_aibot", "WeCom AI Bot channel started (WebSocket mode)")
+ return nil
+}
+
+// Stop shuts down the channel and closes the WebSocket connection.
+func (c *WeComAIBotWSChannel) Stop(_ context.Context) error {
+ logger.InfoC("wecom_aibot", "Stopping WeCom AI Bot channel (WebSocket mode)...")
+ if c.cancel != nil {
+ c.cancel()
+ }
+ c.connMu.Lock()
+ if c.conn != nil {
+ c.conn.Close()
+ c.conn = nil
+ }
+ c.connMu.Unlock()
+ c.SetRunning(false)
+ logger.InfoC("wecom_aibot", "WeCom AI Bot channel stopped")
+ return nil
+}
+
+// Send delivers the agent reply for msg.ChatID.
+// The waiting task goroutine picks it up and writes the final stream response.
+func (c *WeComAIBotWSChannel) Send(ctx context.Context, msg bus.OutboundMessage) error {
+ if !c.IsRunning() {
+ return channels.ErrNotRunning
+ }
+
+ // msg.ChatID carries the inbound req_id (set by dispatchWSAgentTask).
+ // For cron-triggered messages, msg.ChatID is the real WeCom chat/user ID
+ // and there will be no matching entry in reqStates; fall through to proactive push.
+ task, route, ok := c.getReqState(msg.ChatID)
+ if !ok {
+ // No req_id record found — this is a cron/scheduler-originated message.
+ // Send it as a proactive markdown push using the chat ID directly.
+ logger.InfoCF("wecom_aibot", "Send: no req_id state, delivering via proactive push (cron/scheduler)",
+ map[string]any{"chat_id": msg.ChatID})
+ if err := c.wsSendActivePush(msg.ChatID, 0, msg.Content); err != nil {
+ logger.WarnCF("wecom_aibot", "Proactive push failed",
+ map[string]any{"chat_id": msg.ChatID, "error": err.Error()})
+ return fmt.Errorf("websocket delivery failed: %w", channels.ErrSendFailed)
+ }
+ return nil
+ }
+
+ if task == nil {
+ if time.Now().Before(route.ReadyAt) {
+ // Keep using aibot_respond_msg within stream window; do not proactively
+ // push unless wsStreamMaxDuration has elapsed.
+ logger.WarnCF("wecom_aibot", "Send: stream window still open, skip proactive push",
+ map[string]any{"req_id": msg.ChatID, "ready_at": route.ReadyAt.Format(time.RFC3339)})
+ return nil
+ }
+
+ if err := c.wsSendActivePush(route.ChatID, route.ChatType, msg.Content); err != nil {
+ logger.WarnCF("wecom_aibot", "Late reply proactive push failed",
+ map[string]any{"req_id": msg.ChatID, "chat_id": route.ChatID, "error": err.Error()})
+ return fmt.Errorf("websocket delivery failed: %w", channels.ErrSendFailed)
+ }
+ logger.InfoCF("wecom_aibot", "Late reply delivered via proactive push",
+ map[string]any{"req_id": msg.ChatID, "chat_id": route.ChatID, "chat_type": route.ChatType})
+ c.deleteReqState(msg.ChatID)
+ return nil
+ }
+
+ // Non-blocking fast path: when answerCh has space, deliver without racing
+ // against task.ctx.Done() (which fires when the task is canceled by a new
+ // incoming message, but the response must still be sent).
+ select {
+ case task.answerCh <- msg.Content:
+ return nil
+ default:
+ }
+ // answerCh was full; block with cancellation guards.
+ select {
+ case task.answerCh <- msg.Content:
+ case <-task.ctx.Done():
+ return nil
+ case <-ctx.Done():
+ return ctx.Err()
+ }
+ return nil
+}
+
+// ---- Connection management ----
+
+// wsBackoffResetDuration is the minimum duration a WebSocket connection must
+// stay up before we reset the reconnect backoff to its initial value. This
+// prevents a short burst of failures from causing long waits after later,
+// stable connection periods.
+const wsBackoffResetDuration = time.Minute
+
+// connectLoop maintains the WebSocket connection, reconnecting on failure with
+// exponential backoff.
+func (c *WeComAIBotWSChannel) connectLoop() {
+ backoff := wsInitialReconnect
+ for {
+ select {
+ case <-c.ctx.Done():
+ return
+ default:
+ }
+
+ logger.InfoC("wecom_aibot", "Connecting to WeCom WebSocket endpoint...")
+ start := time.Now()
+ if err := c.runConnection(); err != nil {
+ elapsed := time.Since(start)
+ // If the connection was stable for long enough, reset backoff so that
+ // a previous burst of failures does not keep us at the maximum delay.
+ if elapsed >= wsBackoffResetDuration {
+ backoff = wsInitialReconnect
+ }
+ select {
+ case <-c.ctx.Done():
+ return
+ default:
+ logger.WarnCF("wecom_aibot", "WebSocket connection lost, reconnecting",
+ map[string]any{"error": err.Error(), "backoff": backoff.String()})
+ select {
+ case <-time.After(backoff):
+ case <-c.ctx.Done():
+ return
+ }
+ if backoff < wsMaxReconnectWait {
+ backoff *= 2
+ if backoff > wsMaxReconnectWait {
+ backoff = wsMaxReconnectWait
+ }
+ }
+ }
+ } else {
+ // Clean exit (context canceled); stop reconnecting.
+ return
+ }
+ }
+}
+
+// runConnection dials, subscribes, and runs the read/heartbeat loops until the
+// connection closes or the channel context is canceled.
+func (c *WeComAIBotWSChannel) runConnection() error {
+ dialCtx, dialCancel := context.WithTimeout(c.ctx, wsConnectTimeout)
+ conn, httpResp, err := websocket.DefaultDialer.DialContext(dialCtx, wsEndpoint, nil)
+ dialCancel()
+ if httpResp != nil {
+ httpResp.Body.Close()
+ }
+ if err != nil {
+ return fmt.Errorf("dial failed: %w", err)
+ }
+
+ c.connMu.Lock()
+ c.conn = conn
+ c.connMu.Unlock()
+
+ defer func() {
+ c.connMu.Lock()
+ if c.conn == conn {
+ c.conn = nil
+ }
+ c.connMu.Unlock()
+ // Cancel any tasks that were started over this connection so their
+ // agent goroutines do not keep running after the connection is gone.
+ c.cancelAllTasks()
+ }()
+
+ // ---- Read loop (must start BEFORE subscribing) ----
+ // sendAndWait blocks waiting for the subscribe response on reqPending;
+ // readLoop is the only goroutine that delivers messages to reqPending.
+ // Starting readLoop first avoids a deadlock where sendAndWait times out
+ // because no one reads the server's reply.
+ readErrCh := make(chan error, 1)
+ go func() { readErrCh <- c.readLoop(conn) }()
+
+ // ---- Subscribe ----
+ reqID := wsGenerateID()
+ resp, err := c.sendAndWait(conn, reqID, wsCommand{
+ Cmd: "aibot_subscribe",
+ Headers: wsHeaders{ReqID: reqID},
+ Body: map[string]string{
+ "bot_id": c.config.BotID,
+ "secret": c.config.Secret,
+ },
+ }, wsSubscribeTimeout)
+ if err != nil {
+ conn.Close() // stop readLoop
+ <-readErrCh
+ return fmt.Errorf("subscribe failed: %w", err)
+ }
+ if resp.ErrCode != 0 {
+ conn.Close()
+ <-readErrCh
+ return fmt.Errorf("subscribe rejected (errcode=%d): %s", resp.ErrCode, resp.ErrMsg)
+ }
+
+ logger.InfoC("wecom_aibot", "WebSocket subscription successful")
+
+ // ---- Heartbeat goroutine ----
+ hbDone := make(chan struct{})
+ go func() {
+ defer close(hbDone)
+ c.heartbeatLoop(conn)
+ }()
+
+ // Wait for the read loop to exit, then tear down the heartbeat.
+ readErr := <-readErrCh
+ conn.Close() // signal heartbeat to stop (idempotent)
+ <-hbDone
+ return readErr
+}
+
+// sendAndWait registers a pending-response slot, sends cmd, and blocks until
+// the matching response arrives or the timeout/context fires.
+func (c *WeComAIBotWSChannel) sendAndWait(
+ conn *websocket.Conn,
+ reqID string,
+ cmd wsCommand,
+ timeout time.Duration,
+) (wsEnvelope, error) {
+ ch := make(chan wsEnvelope, 1)
+ c.reqPendingMu.Lock()
+ c.reqPending[reqID] = ch
+ c.reqPendingMu.Unlock()
+
+ cleanup := func() {
+ c.reqPendingMu.Lock()
+ delete(c.reqPending, reqID)
+ c.reqPendingMu.Unlock()
+ }
+
+ data, err := json.Marshal(cmd)
+ if err != nil {
+ cleanup()
+ return wsEnvelope{}, fmt.Errorf("marshal command: %w", err)
+ }
+ c.connMu.Lock()
+ err = conn.WriteMessage(websocket.TextMessage, data)
+ c.connMu.Unlock()
+ if err != nil {
+ cleanup()
+ return wsEnvelope{}, fmt.Errorf("write command: %w", err)
+ }
+
+ timer := time.NewTimer(timeout)
+ defer timer.Stop()
+ select {
+ case env := <-ch:
+ return env, nil
+ case <-timer.C:
+ cleanup()
+ return wsEnvelope{}, fmt.Errorf("timeout waiting for response (req_id=%s)", reqID)
+ case <-c.ctx.Done():
+ cleanup()
+ return wsEnvelope{}, c.ctx.Err()
+ }
+}
+
+// heartbeatLoop sends a ping every wsHeartbeatInterval until conn is closed.
+// It validates the server's pong response via sendAndWait; a failed pong
+// triggers a reconnection by closing the connection.
+func (c *WeComAIBotWSChannel) heartbeatLoop(conn *websocket.Conn) {
+ ticker := time.NewTicker(wsHeartbeatInterval)
+ defer ticker.Stop()
+ for {
+ select {
+ case <-ticker.C:
+ reqID := wsGenerateID()
+ resp, err := c.sendAndWait(conn, reqID, wsCommand{
+ Cmd: "ping",
+ Headers: wsHeaders{ReqID: reqID},
+ }, wsHeartbeatInterval)
+ if err != nil {
+ logger.WarnCF("wecom_aibot", "Heartbeat failed, closing connection",
+ map[string]any{"error": err.Error()})
+ conn.Close()
+ return
+ }
+ if resp.ErrCode != 0 {
+ logger.WarnCF("wecom_aibot", "Heartbeat rejected",
+ map[string]any{"errcode": resp.ErrCode, "errmsg": resp.ErrMsg})
+ conn.Close()
+ return
+ }
+ logger.DebugCF("wecom_aibot", "Heartbeat pong received", map[string]any{"req_id": reqID})
+ case <-c.ctx.Done():
+ return
+ }
+ }
+}
+
+// readLoop reads WebSocket messages and dispatches them until the connection
+// closes or the channel is stopped.
+func (c *WeComAIBotWSChannel) readLoop(conn *websocket.Conn) error {
+ for {
+ _, raw, err := conn.ReadMessage()
+ if err != nil {
+ select {
+ case <-c.ctx.Done():
+ return nil // clean shutdown
+ default:
+ return fmt.Errorf("read error: %w", err)
+ }
+ }
+
+ var env wsEnvelope
+ if err := json.Unmarshal(raw, &env); err != nil {
+ logger.WarnCF("wecom_aibot", "Failed to parse WebSocket message",
+ map[string]any{"error": err.Error(), "raw": string(raw)})
+ continue
+ }
+
+ // Command responses have an empty Cmd field; forward to any waiting
+ // sendAndWait() call, or silently drop if no one is waiting (e.g.
+ // late responses after timeout).
+ if env.Cmd == "" && env.Headers.ReqID != "" {
+ c.reqPendingMu.Lock()
+ ch, ok := c.reqPending[env.Headers.ReqID]
+ if ok {
+ delete(c.reqPending, env.Headers.ReqID)
+ }
+ c.reqPendingMu.Unlock()
+ if ok {
+ ch <- env
+ }
+ continue
+ }
+
+ // Dispatch to appropriate handler in a separate goroutine so the
+ // read loop is never blocked by a slow agent.
+ go c.handleEnvelope(env)
+ }
+}
+
+// ---- Message / event handlers ----
+
+// handleEnvelope routes a WebSocket envelope to the right handler.
+func (c *WeComAIBotWSChannel) handleEnvelope(env wsEnvelope) {
+ switch env.Cmd {
+ case "aibot_msg_callback":
+ c.handleMsgCallback(env)
+ case "aibot_event_callback":
+ c.handleEventCallback(env)
+ default:
+ logger.DebugCF("wecom_aibot", "Unhandled WebSocket command",
+ map[string]any{"cmd": env.Cmd})
+ }
+}
+
+// handleMsgCallback processes aibot_msg_callback.
+func (c *WeComAIBotWSChannel) handleMsgCallback(env wsEnvelope) {
+ var msg WeComAIBotWSMessage
+ if err := json.Unmarshal(env.Body, &msg); err != nil {
+ logger.WarnCF("wecom_aibot", "Failed to parse msg callback body",
+ map[string]any{"error": err.Error()})
+ return
+ }
+
+ // Deduplicate by msgid (WeCom may re-deliver on network issues).
+ if msg.MsgID != "" && !c.dedupe.MarkMessageProcessed(msg.MsgID) {
+ logger.DebugCF("wecom_aibot", "Duplicate message ignored",
+ map[string]any{"msgid": msg.MsgID})
+ return
+ }
+
+ reqID := env.Headers.ReqID
+ switch msg.MsgType {
+ case "text":
+ c.handleWSTextMessage(reqID, msg)
+ case "image":
+ c.handleWSImageMessage(reqID, msg)
+ case "voice":
+ c.handleWSVoiceMessage(reqID, msg)
+ case "mixed":
+ c.handleWSMixedMessage(reqID, msg)
+ case "file":
+ c.handleWSFileMessage(reqID, msg)
+ case "video":
+ c.handleWSVideoMessage(reqID, msg)
+ default:
+ logger.WarnCF("wecom_aibot", "Unsupported message type",
+ map[string]any{"msgtype": msg.MsgType})
+ c.wsSendStreamFinish(reqID, wsGenerateID(),
+ "Unsupported message type: "+msg.MsgType)
+ }
+}
+
+// handleEventCallback processes aibot_event_callback.
+func (c *WeComAIBotWSChannel) handleEventCallback(env wsEnvelope) {
+ var msg WeComAIBotWSMessage
+ if err := json.Unmarshal(env.Body, &msg); err != nil {
+ logger.WarnCF("wecom_aibot", "Failed to parse event callback body",
+ map[string]any{"error": err.Error()})
+ return
+ }
+
+ // Deduplicate by msgid.
+ if msg.MsgID != "" && !c.dedupe.MarkMessageProcessed(msg.MsgID) {
+ logger.DebugCF("wecom_aibot", "Duplicate event ignored",
+ map[string]any{"msgid": msg.MsgID})
+ return
+ }
+
+ var eventType string
+ if msg.Event != nil {
+ eventType = msg.Event.EventType
+ }
+ logger.DebugCF("wecom_aibot", "Received event callback",
+ map[string]any{"event_type": eventType})
+
+ switch eventType {
+ case "enter_chat":
+ if c.config.WelcomeMessage != "" {
+ c.wsSendWelcomeMsg(env.Headers.ReqID, c.config.WelcomeMessage)
+ }
+ case "disconnected_event":
+ // The server will close this connection after sending this event.
+ // connectLoop will detect the closure and reconnect automatically.
+ logger.WarnC("wecom_aibot",
+ "Received disconnected_event: this connection is being replaced by a newer one")
+ default:
+ logger.DebugCF("wecom_aibot", "Unhandled event type",
+ map[string]any{"event_type": eventType})
+ }
+}
+
+// handleWSTextMessage dispatches a plain-text message to the agent and streams
+// the reply back over the WebSocket connection.
+func (c *WeComAIBotWSChannel) handleWSTextMessage(reqID string, msg WeComAIBotWSMessage) {
+ if msg.Text == nil {
+ logger.ErrorC("wecom_aibot", "text message missing text field")
+ return
+ }
+ c.dispatchWSAgentTask(reqID, msg, msg.Text.Content, nil)
+}
+
+// handleWSImageMessage downloads and stores the inbound image, then dispatches
+// it to the agent as a media-tagged message.
+func (c *WeComAIBotWSChannel) handleWSImageMessage(reqID string, msg WeComAIBotWSMessage) {
+ if msg.Image == nil {
+ logger.WarnC("wecom_aibot", "Image message missing image field")
+ c.wsSendStreamFinish(reqID, wsGenerateID(), "Image message could not be processed.")
+ return
+ }
+ c.wsHandleMediaMessage(reqID, msg, msg.Image.URL, msg.Image.AESKey, "image")
+}
+
+// wsHandleMediaMessage is a shared helper for image, file and video messages.
+// It downloads the resource, stores it in MediaStore, and dispatches to the agent.
+func (c *WeComAIBotWSChannel) wsHandleMediaMessage(
+ reqID string, msg WeComAIBotWSMessage,
+ resourceURL, aesKey, label string,
+) {
+ chatID := wsChatID(msg)
+
+ ctx, cancel := context.WithTimeout(c.ctx, wsImageDownloadTimeout)
+ defer cancel()
+
+ ref, err := c.storeWSMedia(ctx, chatID, msg.MsgID, resourceURL, aesKey, wsLabelToDefaultExt(label))
+ if err != nil {
+ logger.WarnCF("wecom_aibot", "Failed to download/store WS "+label,
+ map[string]any{"error": err.Error(), "url": resourceURL})
+ c.wsSendStreamFinish(reqID, wsGenerateID(),
+ strings.ToUpper(label[:1])+label[1:]+" message could not be processed.")
+ return
+ }
+
+ c.dispatchWSAgentTask(reqID, msg, "["+label+"]", []string{ref})
+}
+
+// handleWSMixedMessage handles mixed text+image messages.
+// All text parts are collected into the content string; all image parts are
+// downloaded and stored in MediaStore before dispatching to the agent.
+func (c *WeComAIBotWSChannel) handleWSMixedMessage(reqID string, msg WeComAIBotWSMessage) {
+ if msg.Mixed == nil {
+ logger.WarnC("wecom_aibot", "Mixed message has no content")
+ c.wsSendStreamFinish(reqID, wsGenerateID(), "Mixed message type is not yet fully supported.")
+ return
+ }
+
+ chatID := wsChatID(msg)
+
+ ctx, cancel := context.WithTimeout(c.ctx, wsImageDownloadTimeout)
+ defer cancel()
+
+ var textParts []string
+ var mediaRefs []string
+ for _, item := range msg.Mixed.MsgItem {
+ switch item.MsgType {
+ case "text":
+ if item.Text != nil && item.Text.Content != "" {
+ textParts = append(textParts, item.Text.Content)
+ }
+ case "image":
+ if item.Image != nil {
+ ref, err := c.storeWSMedia(ctx, chatID,
+ msg.MsgID+"-"+wsGenerateID(), item.Image.URL, item.Image.AESKey, ".jpg")
+ if err != nil {
+ logger.WarnCF("wecom_aibot", "Failed to download/store mixed image",
+ map[string]any{"error": err.Error()})
+ } else {
+ mediaRefs = append(mediaRefs, ref)
+ }
+ }
+ default:
+ logger.WarnCF("wecom_aibot", "Unsupported item type in mixed message",
+ map[string]any{"msgtype": item.MsgType})
+ }
+ }
+
+ if len(textParts) == 0 && len(mediaRefs) == 0 {
+ logger.WarnC("wecom_aibot", "Mixed message has no usable content")
+ c.wsSendStreamFinish(reqID, wsGenerateID(), "Mixed message type is not yet fully supported.")
+ return
+ }
+
+ content := strings.Join(textParts, "\n")
+ if content == "" {
+ content = "[images]"
+ }
+ c.dispatchWSAgentTask(reqID, msg, content, mediaRefs)
+}
+
+// dispatchWSAgentTask registers a new agent task, sends the opening stream frame,
+// and starts a goroutine that runs the agent and streams the reply back.
+// content is the text forwarded to the agent; mediaRefs are optional media
+// store references attached to the inbound message.
+func (c *WeComAIBotWSChannel) dispatchWSAgentTask(
+ reqID string,
+ msg WeComAIBotWSMessage,
+ content string,
+ mediaRefs []string,
+) {
+ userID := msg.From.UserID
+ if userID == "" {
+ userID = "unknown"
+ }
+ // actualChatID is the real WeCom chat/user ID used for peer identification.
+ // reqID is used as the routing chatID so each turn is independently addressable.
+ actualChatID := wsChatID(msg)
+
+ streamID := wsGenerateID()
+ chatType := wsChatTypeValue(msg.ChatType)
+ taskCtx, taskCancel := context.WithCancel(c.ctx)
+
+ task := &wsTask{
+ ReqID: reqID,
+ ChatID: actualChatID,
+ ChatType: chatType,
+ StreamID: streamID,
+ answerCh: make(chan string, 1),
+ ctx: taskCtx,
+ cancel: taskCancel,
+ }
+ // Each req_id is unique per WeCom turn; tasks run concurrently, no cancellation.
+ c.setReqState(reqID, &wsReqState{
+ Task: task,
+ Route: wsLateReplyRoute{
+ ChatID: actualChatID,
+ ChatType: chatType,
+ ReadyAt: time.Now().Add(wsStreamMaxDuration),
+ ExpiresAt: time.Now().Add(wsLateReplyRouteTTL),
+ },
+ })
+
+ logger.DebugCF("wecom_aibot", "Registered new agent task",
+ map[string]any{"chat_id": actualChatID, "req_id": reqID, "stream_id": streamID})
+
+ // Send an empty stream opening frame (finish=false) immediately.
+ c.wsSendStreamChunk(reqID, streamID, false, "")
+
+ go func() {
+ defer func() {
+ taskCancel()
+ c.clearReqTask(reqID, task)
+ }()
+
+ sender := bus.SenderInfo{
+ Platform: "wecom_aibot",
+ PlatformID: userID,
+ CanonicalID: identity.BuildCanonicalID("wecom_aibot", userID),
+ DisplayName: userID,
+ }
+ peerKind := "direct"
+ if msg.ChatType == "group" {
+ peerKind = "group"
+ }
+ peer := bus.Peer{Kind: peerKind, ID: actualChatID}
+ metadata := map[string]string{
+ "channel": "wecom_aibot",
+ "chat_id": actualChatID,
+ "chat_type": msg.ChatType,
+ "msg_type": msg.MsgType,
+ "msgid": msg.MsgID,
+ "aibotid": msg.AIBotID,
+ "stream_id": streamID,
+ }
+ // Pass reqID as chatID: OutboundMessage.ChatID = reqID → Send() finds tasks[reqID].
+ c.HandleMessage(taskCtx, peer, reqID, userID, reqID,
+ content, mediaRefs, metadata, sender)
+
+ // Wait for the agent reply. While waiting, send periodic finish=false
+ // hints so the user knows processing is still in progress.
+ // WeCom requires finish=true within 6 minutes of the first stream frame;
+ // wsStreamMaxDuration enforces that limit with a safety margin.
+ waitHints := []string{
+ "⏳ Processing, please wait...",
+ "⏳ Still processing, please wait...",
+ "⏳ Almost there, please wait...",
+ }
+ ticker := time.NewTicker(wsStreamTickInterval)
+ defer ticker.Stop()
+ deadlineTimer := time.NewTimer(wsStreamMaxDuration)
+ defer deadlineTimer.Stop()
+ tickCount := 0
+ for {
+ select {
+ case answer := <-task.answerCh:
+ // Split the answer into byte-bounded chunks and send as stream frames.
+ // All but the last carry finish=false; the final frame closes the stream.
+ chunks := splitWSContent(answer, wsStreamMaxContentBytes)
+ for i, chunk := range chunks {
+ c.wsSendStreamChunk(reqID, streamID, i == len(chunks)-1, chunk)
+ }
+ c.deleteReqState(reqID)
+ return
+ case <-ticker.C:
+ hint := waitHints[tickCount%len(waitHints)]
+ tickCount++
+ logger.DebugCF("wecom_aibot", "Sending stream progress hint",
+ map[string]any{"chat_id": actualChatID, "tick": tickCount})
+ c.wsSendStreamChunk(reqID, streamID, false, hint)
+ case <-deadlineTimer.C:
+ logger.WarnCF("wecom_aibot",
+ "Stream response deadline reached, closing stream; late reply will be pushed",
+ map[string]any{"chat_id": actualChatID})
+ c.wsSendStreamFinish(reqID, streamID,
+ "⏳ Processing is taking longer than expected, the response will be sent as a follow-up message.")
+ return
+ case <-taskCtx.Done():
+ // Give a short grace period so that a response queued in the bus
+ // just before cancellation can still be delivered. This closes a
+ // race where a rapid second message cancels this task after the
+ // agent already published but before Send() wrote to answerCh.
+ //
+ // The connection is gone at this point, so we cannot use
+ // wsSendStreamFinish. Try wsSendActivePush on the (possibly
+ // already-restored) connection; if that also fails, leave the
+ // route intact so Send() can push the reply once reconnected.
+ select {
+ case answer := <-task.answerCh:
+ if err := c.wsSendActivePush(task.ChatID, task.ChatType, answer); err != nil {
+ logger.WarnCF("wecom_aibot",
+ "Grace-period push failed after task cancellation; reply may be lost",
+ map[string]any{"req_id": reqID, "chat_id": task.ChatID, "error": err.Error()})
+ } else {
+ c.deleteReqState(reqID)
+ }
+ case <-time.After(100 * time.Millisecond):
+ }
+ return
+ }
+ }
+ }()
+}
+
+// handleWSVoiceMessage handles voice messages.
+// WeCom transcribes voice to text in the callback; if the transcription is
+// present it is dispatched as plain text to the agent.
+func (c *WeComAIBotWSChannel) handleWSVoiceMessage(reqID string, msg WeComAIBotWSMessage) {
+ if msg.Voice != nil && msg.Voice.Content != "" {
+ c.dispatchWSAgentTask(reqID, msg, msg.Voice.Content, nil)
+ return
+ }
+ c.wsSendStreamFinish(reqID, wsGenerateID(), "Voice messages are not yet supported.")
+}
+
+// handleWSFileMessage handles file messages.
+func (c *WeComAIBotWSChannel) handleWSFileMessage(reqID string, msg WeComAIBotWSMessage) {
+ if msg.File == nil {
+ logger.WarnC("wecom_aibot", "File message missing file field")
+ c.wsSendStreamFinish(reqID, wsGenerateID(), "File message could not be processed.")
+ return
+ }
+ c.wsHandleMediaMessage(reqID, msg, msg.File.URL, msg.File.AESKey, "file")
+}
+
+// handleWSVideoMessage handles video messages.
+func (c *WeComAIBotWSChannel) handleWSVideoMessage(reqID string, msg WeComAIBotWSMessage) {
+ if msg.Video == nil {
+ logger.WarnC("wecom_aibot", "Video message missing video field")
+ c.wsSendStreamFinish(reqID, wsGenerateID(), "Video message could not be processed.")
+ return
+ }
+ c.wsHandleMediaMessage(reqID, msg, msg.Video.URL, msg.Video.AESKey, "video")
+}
+
+// ---- WebSocket write helpers ----
+
+// wsSendStreamChunk sends an aibot_respond_msg stream frame.
+func (c *WeComAIBotWSChannel) wsSendStreamChunk(reqID, streamID string, finish bool, content string) {
+ logger.DebugCF("wecom_aibot", "Sending stream chunk", map[string]any{
+ "stream_id": streamID,
+ "finish": finish,
+ "preview": utils.Truncate(content, 100),
+ })
+ cmd := wsCommand{
+ Cmd: "aibot_respond_msg",
+ Headers: wsHeaders{ReqID: reqID},
+ Body: wsRespondMsgBody{
+ MsgType: "stream",
+ Stream: &wsStreamContent{
+ ID: streamID,
+ Finish: finish,
+ Content: content,
+ },
+ },
+ }
+ if err := c.writeWSAndWait(cmd, wsRespondMsgTimeout); err != nil {
+ logger.WarnCF("wecom_aibot", "Stream chunk ack failed", map[string]any{
+ "req_id": reqID,
+ "stream_id": streamID,
+ "finish": finish,
+ "error": err,
+ })
+ }
+}
+
+// wsSendStreamFinish sends the final aibot_respond_msg frame (finish=true, no images).
+func (c *WeComAIBotWSChannel) wsSendStreamFinish(reqID, streamID, content string) {
+ c.wsSendStreamChunk(reqID, streamID, true, content)
+}
+
+// wsSendWelcomeMsg sends a text welcome message via aibot_respond_welcome_msg.
+func (c *WeComAIBotWSChannel) wsSendWelcomeMsg(reqID, content string) {
+ logger.DebugCF("wecom_aibot", "Sending welcome message", map[string]any{"req_id": reqID})
+ cmd := wsCommand{
+ Cmd: "aibot_respond_welcome_msg",
+ Headers: wsHeaders{ReqID: reqID},
+ Body: wsRespondMsgBody{
+ MsgType: "text",
+ Text: &wsTextContent{Content: content},
+ },
+ }
+ if err := c.writeWSAndWait(cmd, wsWelcomeMsgTimeout); err != nil {
+ logger.WarnCF("wecom_aibot", "Welcome message ack failed",
+ map[string]any{"req_id": reqID, "error": err.Error()})
+ }
+}
+
+// wsSendActivePush sends a proactive markdown message using aibot_send_msg.
+// Long content is automatically split into byte-bounded chunks (≤ wsStreamMaxContentBytes
+// each) and delivered as consecutive messages.
+// It is used as a fallback for late replies after stream response window expires.
+func (c *WeComAIBotWSChannel) wsSendActivePush(chatID string, chatType uint32, content string) error {
+ if chatID == "" {
+ return fmt.Errorf("chatid is empty")
+ }
+ for _, chunk := range splitWSContent(content, wsStreamMaxContentBytes) {
+ reqID := wsGenerateID()
+ if err := c.writeWSAndWait(wsCommand{
+ Cmd: "aibot_send_msg",
+ Headers: wsHeaders{ReqID: reqID},
+ Body: wsSendMsgBody{
+ ChatID: chatID,
+ ChatType: chatType,
+ MsgType: "markdown",
+ Markdown: &wsMarkdownContent{Content: chunk},
+ },
+ }, wsSendMsgTimeout); err != nil {
+ return err
+ }
+ }
+ return nil
+}
+
+// writeWSAndWait writes cmd to the active connection and validates the command response.
+func (c *WeComAIBotWSChannel) writeWSAndWait(cmd wsCommand, timeout time.Duration) error {
+ if cmd.Headers.ReqID == "" {
+ return fmt.Errorf("req_id is empty")
+ }
+
+ c.connMu.Lock()
+ conn := c.conn
+ c.connMu.Unlock()
+ if conn == nil {
+ return fmt.Errorf("websocket not connected")
+ }
+
+ resp, err := c.sendAndWait(conn, cmd.Headers.ReqID, cmd, timeout)
+ if err != nil {
+ return err
+ }
+ if resp.ErrCode != 0 {
+ return fmt.Errorf("%s rejected (errcode=%d): %s", cmd.Cmd, resp.ErrCode, resp.ErrMsg)
+ }
+ return nil
+}
+
+// cancelAllTasks cancels every pending agent task; called when the connection drops.
+// It also expires each task's stream window (ReadyAt = now) so that when the agent
+// eventually delivers its reply via Send(), the message is forwarded via
+// wsSendActivePush on the restored connection instead of being silently discarded.
+func (c *WeComAIBotWSChannel) cancelAllTasks() {
+ c.reqStatesMu.Lock()
+ defer c.reqStatesMu.Unlock()
+ now := time.Now()
+ for _, state := range c.reqStates {
+ if state != nil && state.Task != nil {
+ state.Task.cancel()
+ state.Task = nil
+ // Expire the stream window immediately so Send() uses wsSendActivePush.
+ state.Route.ReadyAt = now
+ }
+ }
+}
+
+func (c *WeComAIBotWSChannel) setReqState(reqID string, state *wsReqState) {
+ c.reqStatesMu.Lock()
+ defer c.reqStatesMu.Unlock()
+ now := time.Now()
+ for k, v := range c.reqStates {
+ if v == nil || now.After(v.Route.ExpiresAt) {
+ delete(c.reqStates, k)
+ }
+ }
+ c.reqStates[reqID] = state
+}
+
+func (c *WeComAIBotWSChannel) getReqState(reqID string) (*wsTask, wsLateReplyRoute, bool) {
+ c.reqStatesMu.Lock()
+ defer c.reqStatesMu.Unlock()
+ state, ok := c.reqStates[reqID]
+ if !ok || state == nil {
+ return nil, wsLateReplyRoute{}, false
+ }
+ if time.Now().After(state.Route.ExpiresAt) {
+ delete(c.reqStates, reqID)
+ return nil, wsLateReplyRoute{}, false
+ }
+ return state.Task, state.Route, true
+}
+
+func (c *WeComAIBotWSChannel) deleteReqState(reqID string) {
+ c.reqStatesMu.Lock()
+ delete(c.reqStates, reqID)
+ c.reqStatesMu.Unlock()
+}
+
+func (c *WeComAIBotWSChannel) clearReqTask(reqID string, task *wsTask) {
+ c.reqStatesMu.Lock()
+ defer c.reqStatesMu.Unlock()
+ state, ok := c.reqStates[reqID]
+ if !ok || state == nil {
+ return
+ }
+ if state.Task == task {
+ state.Task = nil
+ }
+}
+
+func wsChatTypeValue(chatType string) uint32 {
+ if chatType == "group" {
+ return 2
+ }
+ return 1
+}
+
+// wsChatID returns the effective chat ID from a WS message.
+// For group messages it is msg.ChatID; for single chats it falls back to the sender's UserID.
+func wsChatID(msg WeComAIBotWSMessage) string {
+ if msg.ChatID != "" {
+ return msg.ChatID
+ }
+ return msg.From.UserID
+}
+
+// wsGenerateID generates a random 10-character alphanumeric ID.
+// It is package-level (not a method) so it can be shared by both channel modes.
+func wsGenerateID() string {
+ return generateRandomID(10)
+}
+
+// ---- Inbound media download helpers ----
+
+// storeWSMedia downloads the resource at resourceURL (with optional AES-CBC
+// decryption) and stores it in the MediaStore. The file extension is inferred
+// from the HTTP Content-Type response header; defaultExt is used as a fallback
+// when the content type is absent or unrecognized.
+func (c *WeComAIBotWSChannel) storeWSMedia(
+ ctx context.Context,
+ chatID, msgID, resourceURL, aesKey, defaultExt string,
+) (string, error) {
+ store := c.GetMediaStore()
+ if store == nil {
+ return "", fmt.Errorf("no media store available")
+ }
+
+ const maxSize = 20 << 20 // 20 MB
+
+ req, err := http.NewRequestWithContext(ctx, http.MethodGet, resourceURL, nil)
+ if err != nil {
+ return "", fmt.Errorf("create request: %w", err)
+ }
+ resp, err := wsImageHTTPClient.Do(req)
+ if err != nil {
+ return "", fmt.Errorf("download: %w", err)
+ }
+ defer resp.Body.Close()
+ if resp.StatusCode != http.StatusOK {
+ return "", fmt.Errorf("download HTTP %d", resp.StatusCode)
+ }
+
+ // Infer file extension from the Content-Type response header.
+ ext := wsMediaExtFromContentType(resp.Header.Get("Content-Type"))
+ if ext == "" {
+ ext = defaultExt
+ }
+
+ // Buffer the media in memory, bounded to maxSize.
+ data, err := io.ReadAll(io.LimitReader(resp.Body, int64(maxSize)+1))
+ if err != nil {
+ return "", fmt.Errorf("read media: %w", err)
+ }
+ if len(data) > maxSize {
+ return "", fmt.Errorf("media too large (> %d MB)", maxSize>>20)
+ }
+
+ // AES-CBC decryption if a key is present.
+ if aesKey != "" {
+ key, decErr := base64.StdEncoding.DecodeString(aesKey)
+ if decErr != nil || len(key) != 32 {
+ key, decErr = decodeWeComAESKey(aesKey)
+ if decErr != nil {
+ return "", fmt.Errorf("decode media AES key: %w", decErr)
+ }
+ }
+ data, err = decryptAESCBC(key, data)
+ if err != nil {
+ return "", fmt.Errorf("decrypt media: %w", err)
+ }
+ }
+
+ // Write to a temp file. The file is owned by the MediaStore and deleted by
+ // store.ReleaseAll — no caller-side cleanup needed.
+ mediaDir := filepath.Join(os.TempDir(), "picoclaw_media")
+ if err = os.MkdirAll(mediaDir, 0o700); err != nil {
+ return "", fmt.Errorf("mkdir: %w", err)
+ }
+ tmpFile, err := os.CreateTemp(mediaDir, msgID+"-*"+ext)
+ if err != nil {
+ return "", fmt.Errorf("create temp file: %w", err)
+ }
+ tmpPath := tmpFile.Name()
+ _, writeErr := tmpFile.Write(data)
+ closeErr := tmpFile.Close()
+ if writeErr != nil {
+ os.Remove(tmpPath)
+ return "", fmt.Errorf("write media: %w", writeErr)
+ }
+ if closeErr != nil {
+ os.Remove(tmpPath)
+ return "", fmt.Errorf("close media: %w", closeErr)
+ }
+
+ scope := channels.BuildMediaScope("wecom_aibot", chatID, msgID)
+ ref, err := store.Store(tmpPath, media.MediaMeta{
+ Filename: msgID + ext,
+ Source: "wecom_aibot",
+ }, scope)
+ if err != nil {
+ os.Remove(tmpPath)
+ return "", fmt.Errorf("store: %w", err)
+ }
+ return ref, nil
+}
+
+// wsMediaExtFromContentType returns the lowercase file extension (with leading
+// dot) for the given Content-Type value, or "" when the type is unrecognized.
+func wsMediaExtFromContentType(contentType string) string {
+ if contentType == "" {
+ return ""
+ }
+ // Strip parameters (e.g. "image/jpeg; charset=utf-8" → "image/jpeg").
+ mt := strings.ToLower(strings.TrimSpace(strings.SplitN(contentType, ";", 2)[0]))
+ switch mt {
+ case "image/jpeg", "image/jpg":
+ return ".jpg"
+ case "image/png":
+ return ".png"
+ case "image/gif":
+ return ".gif"
+ case "image/webp":
+ return ".webp"
+ case "video/mp4":
+ return ".mp4"
+ case "video/mpeg", "video/x-mpeg":
+ return ".mpeg"
+ case "video/quicktime":
+ return ".mov"
+ case "video/webm":
+ return ".webm"
+ case "audio/mpeg", "audio/mp3":
+ return ".mp3"
+ case "audio/ogg":
+ return ".ogg"
+ case "audio/wav":
+ return ".wav"
+ case "application/pdf":
+ return ".pdf"
+ case "application/zip":
+ return ".zip"
+ case "application/x-rar-compressed", "application/vnd.rar":
+ return ".rar"
+ case "text/plain":
+ return ".txt"
+ case "application/msword":
+ return ".doc"
+ case "application/vnd.openxmlformats-officedocument.wordprocessingml.document":
+ return ".docx"
+ case "application/vnd.ms-excel":
+ return ".xls"
+ case "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet":
+ return ".xlsx"
+ case "application/vnd.ms-powerpoint":
+ return ".ppt"
+ case "application/vnd.openxmlformats-officedocument.presentationml.presentation":
+ return ".pptx"
+ }
+ return ""
+}
+
+// wsLabelToDefaultExt returns the default file extension for the given media label
+// used in wsHandleMediaMessage. It is the fallback when Content-Type detection fails.
+func wsLabelToDefaultExt(label string) string {
+ switch label {
+ case "image":
+ return ".jpg"
+ case "video":
+ return ".mp4"
+ default: // "file" and any future labels
+ return ".bin"
+ }
+}
+
+// ---- Content length helpers ----
+
+// splitWSContent splits content into chunks each fitting within maxBytes UTF-8
+// bytes, preserving code block integrity via channels.SplitMessage.
+// When SplitMessage still produces an oversized chunk (e.g. dense CJK content),
+// splitAtByteBoundary is applied as a last-resort byte-level fallback.
+func splitWSContent(content string, maxBytes int) []string {
+ if len(content) <= maxBytes {
+ return []string{content}
+ }
+ // SplitMessage works in runes. Use maxBytes as the rune limit: for pure ASCII
+ // this is exact; for multibyte content the byte verification below catches
+ // any chunk that still overflows.
+ chunks := channels.SplitMessage(content, maxBytes)
+ var result []string
+ for _, chunk := range chunks {
+ if len(chunk) <= maxBytes {
+ result = append(result, chunk)
+ } else {
+ // Still too large in bytes (e.g. dense CJK); force-split at UTF-8 boundaries.
+ result = append(result, splitAtByteBoundary(chunk, maxBytes)...)
+ }
+ }
+ return result
+}
+
+// splitAtByteBoundary splits s into parts each ≤ maxBytes bytes by walking back
+// from the hard byte limit to find a valid UTF-8 rune start boundary.
+// This is a last-resort fallback; it does not try to preserve code blocks.
+func splitAtByteBoundary(s string, maxBytes int) []string {
+ var parts []string
+ for len(s) > maxBytes {
+ end := maxBytes
+ // Walk back past any UTF-8 continuation bytes (high two bits == 10).
+ for end > 0 && s[end]>>6 == 0b10 {
+ end--
+ }
+ if end == 0 {
+ end = maxBytes // shouldn't happen with valid UTF-8
+ }
+ parts = append(parts, s[:end])
+ s = strings.TrimLeft(s[end:], " \t\n\r")
+ }
+ if s != "" {
+ parts = append(parts, s)
+ }
+ return parts
+}
diff --git a/pkg/channels/wecom/aibot_ws_test.go b/pkg/channels/wecom/aibot_ws_test.go
new file mode 100644
index 000000000..0a533da5d
--- /dev/null
+++ b/pkg/channels/wecom/aibot_ws_test.go
@@ -0,0 +1,295 @@
+package wecom
+
+import (
+ "bytes"
+ "context"
+ "net/http"
+ "net/http/httptest"
+ "os"
+ "strings"
+ "testing"
+
+ "github.com/sipeed/picoclaw/pkg/bus"
+ "github.com/sipeed/picoclaw/pkg/channels"
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/media"
+)
+
+// newTestWSChannel creates a WeComAIBotWSChannel ready for unit testing.
+func newTestWSChannel(t *testing.T) *WeComAIBotWSChannel {
+ t.Helper()
+ cfg := config.WeComAIBotConfig{
+ Enabled: true,
+ BotID: "test_bot_id",
+ Secret: "test_secret",
+ }
+ ch, err := newWeComAIBotWSChannel(cfg, bus.NewMessageBus())
+ if err != nil {
+ t.Fatalf("create WS channel: %v", err)
+ }
+ return ch
+}
+
+// TestStoreWSMedia_NilStore verifies that storeWSMedia returns an error when no
+// MediaStore has been injected.
+func TestStoreWSMedia_NilStore(t *testing.T) {
+ ch := newTestWSChannel(t)
+ _, err := ch.storeWSMedia(context.Background(), "chat1", "msg1", "http://any", "", ".jpg")
+ if err == nil {
+ t.Fatal("expected error when no MediaStore is set")
+ }
+}
+
+// TestStoreWSMedia_HTTPError verifies that storeWSMedia propagates HTTP errors
+// from the media server.
+func TestStoreWSMedia_HTTPError(t *testing.T) {
+ srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
+ http.Error(w, "not found", http.StatusNotFound)
+ }))
+ defer srv.Close()
+
+ ch := newTestWSChannel(t)
+ ch.SetMediaStore(media.NewFileMediaStore())
+
+ _, err := ch.storeWSMedia(context.Background(), "chat1", "msg1", srv.URL, "", ".jpg")
+ if err == nil {
+ t.Fatal("expected error for HTTP 404")
+ }
+}
+
+// TestStoreWSMedia_ServerUnavailable verifies that storeWSMedia returns a clear
+// error when the media server cannot be reached.
+func TestStoreWSMedia_ServerUnavailable(t *testing.T) {
+ ch := newTestWSChannel(t)
+ ch.SetMediaStore(media.NewFileMediaStore())
+
+ // Port 1 is reserved and will refuse the connection immediately.
+ _, err := ch.storeWSMedia(context.Background(), "chat1", "msg1", "http://127.0.0.1:1", "", ".jpg")
+ if err == nil {
+ t.Fatal("expected error for unreachable server")
+ }
+}
+
+// TestStoreWSMedia_Success_NoAES verifies the happy path: the media is downloaded,
+// a media ref is returned, and the file persists and is readable via Resolve until
+// ReleaseAll is called. The server returns no Content-Type, so the defaultExt is used.
+func TestStoreWSMedia_Success_NoAES(t *testing.T) {
+ imageData := bytes.Repeat([]byte("x"), 256)
+ srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
+ w.WriteHeader(http.StatusOK)
+ _, _ = w.Write(imageData)
+ }))
+ defer srv.Close()
+
+ ch := newTestWSChannel(t)
+ store := media.NewFileMediaStore()
+ ch.SetMediaStore(store)
+
+ ref, err := ch.storeWSMedia(context.Background(), "chat1", "msg1", srv.URL, "", ".jpg")
+ if err != nil {
+ t.Fatalf("expected no error, got %v", err)
+ }
+ if ref == "" {
+ t.Fatal("expected non-empty ref")
+ }
+
+ // File must be accessible after storeWSMedia returns (no premature deletion).
+ path, err := store.Resolve(ref)
+ if err != nil {
+ t.Fatalf("ref should resolve: %v", err)
+ }
+ got, err := os.ReadFile(path)
+ if err != nil {
+ t.Fatalf("file should exist at %s: %v", path, err)
+ }
+ if !bytes.Equal(got, imageData) {
+ t.Errorf("content mismatch: got len=%d, want len=%d", len(got), len(imageData))
+ }
+
+ // ReleaseAll must delete the file (store owns lifecycle).
+ scope := channels.BuildMediaScope("wecom_aibot", "chat1", "msg1")
+ if err := store.ReleaseAll(scope); err != nil {
+ t.Fatalf("ReleaseAll failed: %v", err)
+ }
+ if _, err := os.Stat(path); !os.IsNotExist(err) {
+ t.Errorf("file should have been deleted by ReleaseAll, stat err: %v", err)
+ }
+}
+
+// TestStoreWSMedia_MultipleMessages verifies that concurrent media messages with
+// different msgIDs do not collide and each resolve to distinct files.
+func TestStoreWSMedia_MultipleMessages(t *testing.T) {
+ imageA := bytes.Repeat([]byte("a"), 64)
+ imageB := bytes.Repeat([]byte("b"), 64)
+
+ srvA := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
+ w.WriteHeader(http.StatusOK)
+ _, _ = w.Write(imageA)
+ }))
+ defer srvA.Close()
+ srvB := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
+ w.WriteHeader(http.StatusOK)
+ _, _ = w.Write(imageB)
+ }))
+ defer srvB.Close()
+
+ ch := newTestWSChannel(t)
+ store := media.NewFileMediaStore()
+ ch.SetMediaStore(store)
+
+ refA, err := ch.storeWSMedia(context.Background(), "chat1", "msgA", srvA.URL, "", ".jpg")
+ if err != nil {
+ t.Fatalf("storeWSMedia A: %v", err)
+ }
+ refB, err := ch.storeWSMedia(context.Background(), "chat1", "msgB", srvB.URL, "", ".jpg")
+ if err != nil {
+ t.Fatalf("storeWSMedia B: %v", err)
+ }
+ if refA == refB {
+ t.Fatal("distinct messages must produce distinct refs")
+ }
+
+ pathA, _ := store.Resolve(refA)
+ pathB, _ := store.Resolve(refB)
+ if pathA == pathB {
+ t.Fatal("distinct messages must be stored at distinct paths")
+ }
+
+ gotA, _ := os.ReadFile(pathA)
+ gotB, _ := os.ReadFile(pathB)
+ if !bytes.Equal(gotA, imageA) {
+ t.Errorf("content mismatch for message A")
+ }
+ if !bytes.Equal(gotB, imageB) {
+ t.Errorf("content mismatch for message B")
+ }
+}
+
+// TestStoreWSMedia_ContentTypeExt verifies that the file extension is inferred
+// from the HTTP Content-Type header and the defaultExt fallback is used when the
+// type is absent or unrecognized.
+func TestStoreWSMedia_ContentTypeExt(t *testing.T) {
+ tests := []struct {
+ contentType string
+ wantExt string
+ }{
+ {"image/jpeg", ".jpg"},
+ {"image/png", ".png"},
+ {"video/mp4", ".mp4"},
+ {"application/pdf", ".pdf"},
+ {"application/zip", ".zip"},
+ // With parameters stripped.
+ {"video/mp4; codecs=avc1", ".mp4"},
+ // Unknown type → falls back to defaultExt.
+ {"", ""},
+ {"application/octet-stream", ""},
+ }
+ for _, tc := range tests {
+ got := wsMediaExtFromContentType(tc.contentType)
+ if got != tc.wantExt {
+ t.Errorf("wsMediaExtFromContentType(%q) = %q, want %q", tc.contentType, got, tc.wantExt)
+ }
+ }
+
+ // End-to-end: server returns Content-Type: video/mp4, defaultExt is .bin.
+ // The stored file should carry the .mp4 extension, not .bin.
+ payload := bytes.Repeat([]byte("v"), 128)
+ srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
+ w.Header().Set("Content-Type", "video/mp4")
+ w.WriteHeader(http.StatusOK)
+ _, _ = w.Write(payload)
+ }))
+ defer srv.Close()
+
+ ch := newTestWSChannel(t)
+ store := media.NewFileMediaStore()
+ ch.SetMediaStore(store)
+
+ ref, err := ch.storeWSMedia(context.Background(), "chat1", "vid1", srv.URL, "", ".bin")
+ if err != nil {
+ t.Fatalf("storeWSMedia: %v", err)
+ }
+ path, err := store.Resolve(ref)
+ if err != nil {
+ t.Fatalf("resolve: %v", err)
+ }
+ if ext := path[len(path)-4:]; ext != ".mp4" {
+ t.Errorf("expected .mp4 extension from Content-Type, got %q", ext)
+ }
+}
+
+// TestSplitWSContent verifies byte-aware splitting of stream content.
+func TestSplitWSContent(t *testing.T) {
+ t.Run("short content is not split", func(t *testing.T) {
+ chunks := splitWSContent("hello", 20480)
+ if len(chunks) != 1 || chunks[0] != "hello" {
+ t.Fatalf("unexpected chunks: %v", chunks)
+ }
+ })
+
+ t.Run("ASCII content split at byte boundary", func(t *testing.T) {
+ // Build a string just over the limit.
+ content := strings.Repeat("a", 20481)
+ chunks := splitWSContent(content, 20480)
+ if len(chunks) < 2 {
+ t.Fatalf("expected >= 2 chunks, got %d", len(chunks))
+ }
+ for i, c := range chunks {
+ if len(c) > 20480 {
+ t.Errorf("chunk %d has %d bytes, want <= 20480", i, len(c))
+ }
+ }
+ // Reassembled content must equal the original (possibly without leading
+ // whitespace that splitWSContent trims between chunks).
+ joined := strings.Join(chunks, "")
+ if len(joined) < len(content)-len(chunks) {
+ t.Errorf("joined length %d too short (original %d)", len(joined), len(content))
+ }
+ })
+
+ t.Run("CJK content split within byte limit", func(t *testing.T) {
+ // Each CJK rune is 3 bytes in UTF-8.
+ // 7000 CJK chars = 21000 bytes, which exceeds 20480.
+ content := strings.Repeat("\u4e2d", 7000)
+ chunks := splitWSContent(content, 20480)
+ if len(chunks) < 2 {
+ t.Fatalf("expected >= 2 chunks for 21000-byte CJK content, got %d", len(chunks))
+ }
+ for i, c := range chunks {
+ if len(c) > 20480 {
+ t.Errorf("chunk %d has %d bytes, want <= 20480", i, len(c))
+ }
+ // Every chunk must be valid UTF-8.
+ if !strings.ContainsRune(c, '\u4e2d') && len(c) > 0 {
+ // quick plausibility check — content was pure CJK
+ }
+ }
+ })
+}
+
+// TestSplitAtByteBoundary verifies the last-resort byte-boundary splitter.
+func TestSplitAtByteBoundary(t *testing.T) {
+ t.Run("ASCII fits in one chunk", func(t *testing.T) {
+ parts := splitAtByteBoundary("hello world", 100)
+ if len(parts) != 1 {
+ t.Fatalf("expected 1 part, got %d", len(parts))
+ }
+ })
+
+ t.Run("splits at byte boundary, never mid-rune", func(t *testing.T) {
+ // 10 CJK characters = 30 bytes; split at 20 bytes.
+ s := strings.Repeat("\u6587", 10) // 10 × 3 bytes = 30 bytes
+ parts := splitAtByteBoundary(s, 20)
+ for i, p := range parts {
+ if len(p) > 20 {
+ t.Errorf("part %d has %d bytes, want <= 20", i, len(p))
+ }
+ // Must be valid UTF-8 (no torn multi-byte sequences).
+ for j, r := range p {
+ if r == '\uFFFD' {
+ t.Errorf("part %d has replacement rune at position %d: torn UTF-8", i, j)
+ }
+ }
+ }
+ })
+}
diff --git a/pkg/channels/whatsapp/whatsapp_command_test.go b/pkg/channels/whatsapp/whatsapp_command_test.go
index ee8aa4a52..2d85d74f8 100644
--- a/pkg/channels/whatsapp/whatsapp_command_test.go
+++ b/pkg/channels/whatsapp/whatsapp_command_test.go
@@ -3,7 +3,6 @@ package whatsapp
import (
"context"
"testing"
- "time"
"github.com/sipeed/picoclaw/pkg/bus"
"github.com/sipeed/picoclaw/pkg/channels"
@@ -25,10 +24,7 @@ func TestHandleIncomingMessage_DoesNotConsumeGenericCommandsLocally(t *testing.T
"content": "/help",
})
- ctx, cancel := context.WithTimeout(context.Background(), time.Second)
- defer cancel()
-
- inbound, ok := messageBus.ConsumeInbound(ctx)
+ inbound, ok := <-messageBus.InboundChan()
if !ok {
t.Fatal("expected inbound message to be forwarded")
}
diff --git a/pkg/channels/whatsapp_native/whatsapp_command_test.go b/pkg/channels/whatsapp_native/whatsapp_command_test.go
index cc2dcb619..e51bec392 100644
--- a/pkg/channels/whatsapp_native/whatsapp_command_test.go
+++ b/pkg/channels/whatsapp_native/whatsapp_command_test.go
@@ -43,14 +43,19 @@ func TestHandleIncoming_DoesNotConsumeGenericCommandsLocally(t *testing.T) {
ctx, cancel := context.WithTimeout(context.Background(), time.Second)
defer cancel()
- inbound, ok := messageBus.ConsumeInbound(ctx)
- if !ok {
- t.Fatal("expected inbound message to be forwarded")
- }
- if inbound.Channel != "whatsapp_native" {
- t.Fatalf("channel=%q", inbound.Channel)
- }
- if inbound.Content != "/new" {
- t.Fatalf("content=%q", inbound.Content)
+ select {
+ case <-ctx.Done():
+ t.Fatal("timeout waiting for message to be forwarded")
+ return
+ case inbound, ok := <-messageBus.InboundChan():
+ if !ok {
+ t.Fatal("expected inbound message to be forwarded")
+ }
+ if inbound.Channel != "whatsapp_native" {
+ t.Fatalf("channel=%q", inbound.Channel)
+ }
+ if inbound.Content != "/new" {
+ t.Fatalf("content=%q", inbound.Content)
+ }
}
}
diff --git a/pkg/commands/builtin.go b/pkg/commands/builtin.go
index aed6a1874..7bd36b653 100644
--- a/pkg/commands/builtin.go
+++ b/pkg/commands/builtin.go
@@ -13,5 +13,7 @@ func BuiltinDefinitions() []Definition {
switchCommand(),
checkCommand(),
clearCommand(),
+ subagentsCommand(),
+ reloadCommand(),
}
}
diff --git a/pkg/commands/cmd_reload.go b/pkg/commands/cmd_reload.go
new file mode 100644
index 000000000..07ab44016
--- /dev/null
+++ b/pkg/commands/cmd_reload.go
@@ -0,0 +1,20 @@
+package commands
+
+import "context"
+
+func reloadCommand() Definition {
+ return Definition{
+ Name: "reload",
+ Description: "Reload the configuration file",
+ Usage: "/reload",
+ Handler: func(_ context.Context, req Request, rt *Runtime) error {
+ if rt == nil || rt.ReloadConfig == nil {
+ return req.Reply(unavailableMsg)
+ }
+ if err := rt.ReloadConfig(); err != nil {
+ return req.Reply("Failed to reload configuration: " + err.Error())
+ }
+ return req.Reply("Config reload triggered!")
+ },
+ }
+}
diff --git a/pkg/commands/cmd_subagents.go b/pkg/commands/cmd_subagents.go
new file mode 100644
index 000000000..29321823c
--- /dev/null
+++ b/pkg/commands/cmd_subagents.go
@@ -0,0 +1,42 @@
+package commands
+
+import (
+ "context"
+ "fmt"
+)
+
+// TurnInfo is a mirrored struct from agent.TurnInfo to avoid circular dependencies.
+type TurnInfo struct {
+ TurnID string
+ ParentTurnID string
+ Depth int
+ ChildTurnIDs []string
+ IsFinished bool
+}
+
+func subagentsCommand() Definition {
+ return Definition{
+ Name: "subagents",
+ Description: "Show running subagents and task tree",
+ Handler: func(ctx context.Context, req Request, rt *Runtime) error {
+ getTurnFn := rt.GetActiveTurn
+ if getTurnFn == nil {
+ return req.Reply("Runtime does not support querying active turns.")
+ }
+
+ turnRaw := getTurnFn()
+ if turnRaw == nil {
+ return req.Reply("No active tasks running in this session.")
+ }
+
+ if treeStr, ok := turnRaw.(string); ok {
+ if treeStr == "" {
+ return req.Reply("No active tasks running in this session.")
+ }
+ return req.Reply(fmt.Sprintf("🤖 **Active Subagents Tree**\n```text\n%s\n```", treeStr))
+ }
+
+ return req.Reply(fmt.Sprintf("🤖 **Active Subagents List**\n```text\n%+v\n```", turnRaw))
+ },
+ }
+}
diff --git a/pkg/commands/runtime.go b/pkg/commands/runtime.go
index 037184686..f714e1ca4 100644
--- a/pkg/commands/runtime.go
+++ b/pkg/commands/runtime.go
@@ -11,7 +11,9 @@ type Runtime struct {
ListAgentIDs func() []string
ListDefinitions func() []Definition
GetEnabledChannels func() []string
+ GetActiveTurn func() any // Returning any to avoid circular dependency with agent package
SwitchModel func(value string) (oldModel string, err error)
SwitchChannel func(value string) error
ClearHistory func() error
+ ReloadConfig func() error
}
diff --git a/pkg/config/config.go b/pkg/config/config.go
index a7c44c825..89d89af04 100644
--- a/pkg/config/config.go
+++ b/pkg/config/config.go
@@ -4,11 +4,13 @@ import (
"encoding/json"
"fmt"
"os"
+ "path/filepath"
"strings"
"sync/atomic"
"github.com/caarlos0/env/v11"
+ "github.com/sipeed/picoclaw/pkg/credential"
"github.com/sipeed/picoclaw/pkg/fileutil"
)
@@ -248,28 +250,48 @@ type RoutingConfig struct {
Threshold float64 `json:"threshold"` // complexity score in [0,1]; score >= threshold → primary model
}
-type AgentDefaults struct {
- Workspace string `json:"workspace" env:"PICOCLAW_AGENTS_DEFAULTS_WORKSPACE"`
- RestrictToWorkspace bool `json:"restrict_to_workspace" env:"PICOCLAW_AGENTS_DEFAULTS_RESTRICT_TO_WORKSPACE"`
- AllowReadOutsideWorkspace bool `json:"allow_read_outside_workspace" env:"PICOCLAW_AGENTS_DEFAULTS_ALLOW_READ_OUTSIDE_WORKSPACE"`
- Provider string `json:"provider" env:"PICOCLAW_AGENTS_DEFAULTS_PROVIDER"`
- ModelName string `json:"model_name" env:"PICOCLAW_AGENTS_DEFAULTS_MODEL_NAME"`
- Model string `json:"model,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_MODEL"` // Deprecated: use model_name instead
- ModelFallbacks []string `json:"model_fallbacks,omitempty"`
- ImageModel string `json:"image_model,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_IMAGE_MODEL"`
- ImageModelFallbacks []string `json:"image_model_fallbacks,omitempty"`
- MaxTokens int `json:"max_tokens" env:"PICOCLAW_AGENTS_DEFAULTS_MAX_TOKENS"`
- ContextWindow int `json:"context_window,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_CONTEXT_WINDOW"`
- Temperature *float64 `json:"temperature,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_TEMPERATURE"`
- MaxToolIterations int `json:"max_tool_iterations" env:"PICOCLAW_AGENTS_DEFAULTS_MAX_TOOL_ITERATIONS"`
- SummarizeMessageThreshold int `json:"summarize_message_threshold" env:"PICOCLAW_AGENTS_DEFAULTS_SUMMARIZE_MESSAGE_THRESHOLD"`
- SummarizeTokenPercent int `json:"summarize_token_percent" env:"PICOCLAW_AGENTS_DEFAULTS_SUMMARIZE_TOKEN_PERCENT"`
- MaxMediaSize int `json:"max_media_size,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_MAX_MEDIA_SIZE"`
- Routing *RoutingConfig `json:"routing,omitempty"`
- SteeringMode string `json:"steering_mode,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_STEERING_MODE"` // "one-at-a-time" (default) or "all"
+// SubTurnConfig configures the SubTurn execution system.
+type SubTurnConfig struct {
+ MaxDepth int `json:"max_depth" env:"PICOCLAW_AGENTS_DEFAULTS_SUBTURN_MAX_DEPTH"`
+ MaxConcurrent int `json:"max_concurrent" env:"PICOCLAW_AGENTS_DEFAULTS_SUBTURN_MAX_CONCURRENT"`
+ DefaultTimeoutMinutes int `json:"default_timeout_minutes" env:"PICOCLAW_AGENTS_DEFAULTS_SUBTURN_DEFAULT_TIMEOUT_MINUTES"`
+ DefaultTokenBudget int `json:"default_token_budget" env:"PICOCLAW_AGENTS_DEFAULTS_SUBTURN_DEFAULT_TOKEN_BUDGET"`
+ ConcurrencyTimeoutSec int `json:"concurrency_timeout_sec" env:"PICOCLAW_AGENTS_DEFAULTS_SUBTURN_CONCURRENCY_TIMEOUT_SEC"`
}
-const DefaultMaxMediaSize = 20 * 1024 * 1024 // 20 MB
+type ToolFeedbackConfig struct {
+ Enabled bool `json:"enabled" env:"PICOCLAW_AGENTS_DEFAULTS_TOOL_FEEDBACK_ENABLED"`
+ MaxArgsLength int `json:"max_args_length" env:"PICOCLAW_AGENTS_DEFAULTS_TOOL_FEEDBACK_MAX_ARGS_LENGTH"`
+}
+
+type AgentDefaults struct {
+ Workspace string `json:"workspace" env:"PICOCLAW_AGENTS_DEFAULTS_WORKSPACE"`
+ RestrictToWorkspace bool `json:"restrict_to_workspace" env:"PICOCLAW_AGENTS_DEFAULTS_RESTRICT_TO_WORKSPACE"`
+ AllowReadOutsideWorkspace bool `json:"allow_read_outside_workspace" env:"PICOCLAW_AGENTS_DEFAULTS_ALLOW_READ_OUTSIDE_WORKSPACE"`
+ Provider string `json:"provider" env:"PICOCLAW_AGENTS_DEFAULTS_PROVIDER"`
+ ModelName string `json:"model_name" env:"PICOCLAW_AGENTS_DEFAULTS_MODEL_NAME"`
+ Model string `json:"model,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_MODEL"` // Deprecated: use model_name instead
+ ModelFallbacks []string `json:"model_fallbacks,omitempty"`
+ ImageModel string `json:"image_model,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_IMAGE_MODEL"`
+ ImageModelFallbacks []string `json:"image_model_fallbacks,omitempty"`
+ MaxTokens int `json:"max_tokens" env:"PICOCLAW_AGENTS_DEFAULTS_MAX_TOKENS"`
+ ContextWindow int `json:"context_window,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_CONTEXT_WINDOW"`
+ Temperature *float64 `json:"temperature,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_TEMPERATURE"`
+ MaxToolIterations int `json:"max_tool_iterations" env:"PICOCLAW_AGENTS_DEFAULTS_MAX_TOOL_ITERATIONS"`
+ SummarizeMessageThreshold int `json:"summarize_message_threshold" env:"PICOCLAW_AGENTS_DEFAULTS_SUMMARIZE_MESSAGE_THRESHOLD"`
+ SummarizeTokenPercent int `json:"summarize_token_percent" env:"PICOCLAW_AGENTS_DEFAULTS_SUMMARIZE_TOKEN_PERCENT"`
+ MaxMediaSize int `json:"max_media_size,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_MAX_MEDIA_SIZE"`
+ Routing *RoutingConfig `json:"routing,omitempty"`
+ SteeringMode string `json:"steering_mode,omitempty" env:"PICOCLAW_AGENTS_DEFAULTS_STEERING_MODE"` // "one-at-a-time" (default) or "all"
+ SubTurn SubTurnConfig `json:"subturn" envPrefix:"PICOCLAW_AGENTS_DEFAULTS_SUBTURN_"`
+ ToolFeedback ToolFeedbackConfig `json:"tool_feedback,omitempty"`
+ LogLevel string `json:"log_level,omitempty" env:"PICOCLAW_LOG_LEVEL"`
+}
+
+const (
+ DefaultMaxMediaSize = 20 * 1024 * 1024 // 20 MB
+ DefaultWeComAIBotProcessingMessage = "⏳ Processing, please wait. The results will be sent shortly."
+)
func (d *AgentDefaults) GetMaxMediaSize() int {
if d.MaxMediaSize > 0 {
@@ -278,6 +300,19 @@ func (d *AgentDefaults) GetMaxMediaSize() int {
return DefaultMaxMediaSize
}
+// GetToolFeedbackMaxArgsLength returns the max args preview length for tool feedback messages.
+func (d *AgentDefaults) GetToolFeedbackMaxArgsLength() int {
+ if d.ToolFeedback.MaxArgsLength > 0 {
+ return d.ToolFeedback.MaxArgsLength
+ }
+ return 300
+}
+
+// IsToolFeedbackEnabled returns true when tool feedback messages should be sent to the chat.
+func (d *AgentDefaults) IsToolFeedbackEnabled() bool {
+ return d.ToolFeedback.Enabled
+}
+
// GetModelName returns the effective model name for the agent defaults.
// It prefers the new "model_name" field but falls back to "model" for backward compatibility.
func (d *AgentDefaults) GetModelName() string {
@@ -303,6 +338,7 @@ type ChannelsConfig struct {
WeComApp WeComAppConfig `json:"wecom_app"`
WeComAIBot WeComAIBotConfig `json:"wecom_aibot"`
Pico PicoConfig `json:"pico"`
+ PicoClient PicoClientConfig `json:"pico_client"`
IRC IRCConfig `json:"irc"`
}
@@ -323,6 +359,12 @@ type PlaceholderConfig struct {
Text string `json:"text,omitempty"`
}
+type StreamingConfig struct {
+ Enabled bool `json:"enabled,omitempty" env:"PICOCLAW_CHANNELS_TELEGRAM_STREAMING_ENABLED"`
+ ThrottleSeconds int `json:"throttle_seconds,omitempty" env:"PICOCLAW_CHANNELS_TELEGRAM_STREAMING_THROTTLE_SECONDS"`
+ MinGrowthChars int `json:"min_growth_chars,omitempty" env:"PICOCLAW_CHANNELS_TELEGRAM_STREAMING_MIN_GROWTH_CHARS"`
+}
+
type WhatsAppConfig struct {
Enabled bool `json:"enabled" env:"PICOCLAW_CHANNELS_WHATSAPP_ENABLED"`
BridgeURL string `json:"bridge_url" env:"PICOCLAW_CHANNELS_WHATSAPP_BRIDGE_URL"`
@@ -341,7 +383,9 @@ type TelegramConfig struct {
GroupTrigger GroupTriggerConfig `json:"group_trigger,omitempty"`
Typing TypingConfig `json:"typing,omitempty"`
Placeholder PlaceholderConfig `json:"placeholder,omitempty"`
+ Streaming StreamingConfig `json:"streaming,omitempty"`
ReasoningChannelID string `json:"reasoning_channel_id" env:"PICOCLAW_CHANNELS_TELEGRAM_REASONING_CHANNEL_ID"`
+ UseMarkdownV2 bool `json:"use_markdown_v2" env:"PICOCLAW_CHANNELS_TELEGRAM_USE_MARKDOWN_V2"`
}
type FeishuConfig struct {
@@ -355,6 +399,7 @@ type FeishuConfig struct {
Placeholder PlaceholderConfig `json:"placeholder,omitempty"`
ReasoningChannelID string `json:"reasoning_channel_id" env:"PICOCLAW_CHANNELS_FEISHU_REASONING_CHANNEL_ID"`
RandomReactionEmoji FlexibleStringSlice `json:"random_reaction_emoji" env:"PICOCLAW_CHANNELS_FEISHU_RANDOM_REACTION_EMOJI"`
+ IsLark bool `json:"is_lark" env:"PICOCLAW_CHANNELS_FEISHU_IS_LARK"`
}
type DiscordConfig struct {
@@ -378,14 +423,15 @@ type MaixCamConfig struct {
}
type QQConfig struct {
- Enabled bool `json:"enabled" env:"PICOCLAW_CHANNELS_QQ_ENABLED"`
- AppID string `json:"app_id" env:"PICOCLAW_CHANNELS_QQ_APP_ID"`
- AppSecret string `json:"app_secret" env:"PICOCLAW_CHANNELS_QQ_APP_SECRET"`
- AllowFrom FlexibleStringSlice `json:"allow_from" env:"PICOCLAW_CHANNELS_QQ_ALLOW_FROM"`
- GroupTrigger GroupTriggerConfig `json:"group_trigger,omitempty"`
- MaxMessageLength int `json:"max_message_length" env:"PICOCLAW_CHANNELS_QQ_MAX_MESSAGE_LENGTH"`
- SendMarkdown bool `json:"send_markdown" env:"PICOCLAW_CHANNELS_QQ_SEND_MARKDOWN"`
- ReasoningChannelID string `json:"reasoning_channel_id" env:"PICOCLAW_CHANNELS_QQ_REASONING_CHANNEL_ID"`
+ Enabled bool `json:"enabled" env:"PICOCLAW_CHANNELS_QQ_ENABLED"`
+ AppID string `json:"app_id" env:"PICOCLAW_CHANNELS_QQ_APP_ID"`
+ AppSecret string `json:"app_secret" env:"PICOCLAW_CHANNELS_QQ_APP_SECRET"`
+ AllowFrom FlexibleStringSlice `json:"allow_from" env:"PICOCLAW_CHANNELS_QQ_ALLOW_FROM"`
+ GroupTrigger GroupTriggerConfig `json:"group_trigger,omitempty"`
+ MaxMessageLength int `json:"max_message_length" env:"PICOCLAW_CHANNELS_QQ_MAX_MESSAGE_LENGTH"`
+ MaxBase64FileSizeMiB int64 `json:"max_base64_file_size_mib" env:"PICOCLAW_CHANNELS_QQ_MAX_BASE64_FILE_SIZE_MIB"`
+ SendMarkdown bool `json:"send_markdown" env:"PICOCLAW_CHANNELS_QQ_SEND_MARKDOWN"`
+ ReasoningChannelID string `json:"reasoning_channel_id" env:"PICOCLAW_CHANNELS_QQ_REASONING_CHANNEL_ID"`
}
type DingTalkConfig struct {
@@ -480,15 +526,18 @@ type WeComAppConfig struct {
}
type WeComAIBotConfig struct {
- Enabled bool `json:"enabled" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_ENABLED"`
- Token string `json:"token" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_TOKEN"`
- EncodingAESKey string `json:"encoding_aes_key" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_ENCODING_AES_KEY"`
- WebhookPath string `json:"webhook_path" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_WEBHOOK_PATH"`
- AllowFrom FlexibleStringSlice `json:"allow_from" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_ALLOW_FROM"`
- ReplyTimeout int `json:"reply_timeout" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_REPLY_TIMEOUT"`
- MaxSteps int `json:"max_steps" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_MAX_STEPS"` // Maximum streaming steps
- WelcomeMessage string `json:"welcome_message" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_WELCOME_MESSAGE"` // Sent on enter_chat event; empty = no welcome
- ReasoningChannelID string `json:"reasoning_channel_id" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_REASONING_CHANNEL_ID"`
+ Enabled bool `json:"enabled" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_ENABLED"`
+ BotID string `json:"bot_id,omitempty" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_BOT_ID"`
+ Secret string `json:"secret,omitempty" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_SECRET"`
+ Token string `json:"token,omitempty" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_TOKEN"`
+ EncodingAESKey string `json:"encoding_aes_key,omitempty" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_ENCODING_AES_KEY"`
+ WebhookPath string `json:"webhook_path,omitempty" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_WEBHOOK_PATH"`
+ AllowFrom FlexibleStringSlice `json:"allow_from" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_ALLOW_FROM"`
+ ReplyTimeout int `json:"reply_timeout" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_REPLY_TIMEOUT"`
+ MaxSteps int `json:"max_steps" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_MAX_STEPS"` // Maximum streaming steps
+ WelcomeMessage string `json:"welcome_message" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_WELCOME_MESSAGE"` // Sent on enter_chat event; empty = no welcome
+ ProcessingMessage string `json:"processing_message,omitempty" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_PROCESSING_MESSAGE"`
+ ReasoningChannelID string `json:"reasoning_channel_id" env:"PICOCLAW_CHANNELS_WECOM_AIBOT_REASONING_CHANNEL_ID"`
}
type PicoConfig struct {
@@ -504,6 +553,16 @@ type PicoConfig struct {
Placeholder PlaceholderConfig `json:"placeholder,omitempty"`
}
+type PicoClientConfig struct {
+ Enabled bool `json:"enabled" env:"PICOCLAW_CHANNELS_PICO_CLIENT_ENABLED"`
+ URL string `json:"url" env:"PICOCLAW_CHANNELS_PICO_CLIENT_URL"`
+ Token string `json:"token" env:"PICOCLAW_CHANNELS_PICO_CLIENT_TOKEN"`
+ SessionID string `json:"session_id,omitempty"`
+ PingInterval int `json:"ping_interval,omitempty"`
+ ReadTimeout int `json:"read_timeout,omitempty"`
+ AllowFrom FlexibleStringSlice `json:"allow_from" env:"PICOCLAW_CHANNELS_PICO_CLIENT_ALLOW_FROM"`
+}
+
type IRCConfig struct {
Enabled bool `json:"enabled" env:"PICOCLAW_CHANNELS_IRC_ENABLED"`
Server string `json:"server" env:"PICOCLAW_CHANNELS_IRC_SERVER"`
@@ -562,6 +621,7 @@ type ProvidersConfig struct {
Minimax ProviderConfig `json:"minimax"`
LongCat ProviderConfig `json:"longcat"`
ModelScope ProviderConfig `json:"modelscope"`
+ Novita ProviderConfig `json:"novita"`
}
// IsEmpty checks if all provider configs are empty (no API keys or API bases set)
@@ -590,7 +650,8 @@ func (p ProvidersConfig) IsEmpty() bool {
p.Avian.APIKey == "" && p.Avian.APIBase == "" &&
p.Minimax.APIKey == "" && p.Minimax.APIBase == "" &&
p.LongCat.APIKey == "" && p.LongCat.APIBase == "" &&
- p.ModelScope.APIKey == "" && p.ModelScope.APIBase == ""
+ p.ModelScope.APIKey == "" && p.ModelScope.APIBase == "" &&
+ p.Novita.APIKey == "" && p.Novita.APIBase == ""
}
// MarshalJSON implements custom JSON marshaling for ProvidersConfig
@@ -620,7 +681,9 @@ type OpenAIProviderConfig struct {
// ModelConfig represents a model-centric provider configuration.
// It allows adding new providers (especially OpenAI-compatible ones) via configuration only.
// The model field uses protocol prefix format: [protocol/]model-identifier
-// Supported protocols: openai, anthropic, antigravity, claude-cli, codex-cli, github-copilot
+// Supported protocols include openai, anthropic, antigravity, claude-cli,
+// codex-cli, github-copilot, and named OpenAI-compatible protocols such as
+// groq, deepseek, modelscope, and novita.
// Default protocol is "openai" if no prefix is specified.
type ModelConfig struct {
// Required fields
@@ -628,9 +691,11 @@ type ModelConfig struct {
Model string `json:"model"` // Protocol/model-identifier (e.g., "openai/gpt-4o", "anthropic/claude-sonnet-4.6")
// HTTP-based providers
- APIBase string `json:"api_base,omitempty"` // API endpoint URL
- APIKey string `json:"api_key"` // API authentication key
- Proxy string `json:"proxy,omitempty"` // HTTP proxy URL
+ APIBase string `json:"api_base,omitempty"` // API endpoint URL
+ APIKey string `json:"api_key"` // API authentication key (single key)
+ APIKeys []string `json:"api_keys,omitempty"` // API authentication keys (multiple keys for failover)
+ Proxy string `json:"proxy,omitempty"` // HTTP proxy URL
+ Fallbacks []string `json:"fallbacks,omitempty"` // Fallback model names for failover
// Special providers (CLI-based, OAuth, etc.)
AuthMethod string `json:"auth_method,omitempty"` // Authentication method: oauth, token
@@ -656,8 +721,9 @@ func (c *ModelConfig) Validate() error {
}
type GatewayConfig struct {
- Host string `json:"host" env:"PICOCLAW_GATEWAY_HOST"`
- Port int `json:"port" env:"PICOCLAW_GATEWAY_PORT"`
+ Host string `json:"host" env:"PICOCLAW_GATEWAY_HOST"`
+ Port int `json:"port" env:"PICOCLAW_GATEWAY_PORT"`
+ HotReload bool `json:"hot_reload" env:"PICOCLAW_GATEWAY_HOT_RELOAD"`
}
type ToolDiscoveryConfig struct {
@@ -723,15 +789,24 @@ type WebToolsConfig struct {
Perplexity PerplexityConfig ` json:"perplexity"`
SearXNG SearXNGConfig ` json:"searxng"`
GLMSearch GLMSearchConfig ` json:"glm_search"`
+ // PreferNative controls whether to use provider-native web search when
+ // the active LLM supports it (e.g. OpenAI web_search_preview). When true,
+ // the client-side web_search tool is hidden to avoid duplicate search surfaces,
+ // and the provider's built-in search is used instead. Falls back to client-side
+ // search when the provider does not support native search.
+ PreferNative bool `json:"prefer_native" env:"PICOCLAW_TOOLS_WEB_PREFER_NATIVE"`
// Proxy is an optional proxy URL for web tools (http/https/socks5/socks5h).
// For authenticated proxies, prefer HTTP_PROXY/HTTPS_PROXY env vars instead of embedding credentials in config.
- Proxy string `json:"proxy,omitempty" env:"PICOCLAW_TOOLS_WEB_PROXY"`
- FetchLimitBytes int64 `json:"fetch_limit_bytes,omitempty" env:"PICOCLAW_TOOLS_WEB_FETCH_LIMIT_BYTES"`
+ Proxy string `json:"proxy,omitempty" env:"PICOCLAW_TOOLS_WEB_PROXY"`
+ FetchLimitBytes int64 `json:"fetch_limit_bytes,omitempty" env:"PICOCLAW_TOOLS_WEB_FETCH_LIMIT_BYTES"`
+ Format string `json:"format,omitempty" env:"PICOCLAW_TOOLS_WEB_FORMAT"`
+ PrivateHostWhitelist FlexibleStringSlice `json:"private_host_whitelist,omitempty" env:"PICOCLAW_TOOLS_WEB_PRIVATE_HOST_WHITELIST"`
}
type CronToolsConfig struct {
- ToolConfig ` envPrefix:"PICOCLAW_TOOLS_CRON_"`
- ExecTimeoutMinutes int ` env:"PICOCLAW_TOOLS_CRON_EXEC_TIMEOUT_MINUTES" json:"exec_timeout_minutes"` // 0 means no timeout
+ ToolConfig ` envPrefix:"PICOCLAW_TOOLS_CRON_"`
+ ExecTimeoutMinutes int ` env:"PICOCLAW_TOOLS_CRON_EXEC_TIMEOUT_MINUTES" json:"exec_timeout_minutes"` // 0 means no timeout
+ AllowCommand bool ` env:"PICOCLAW_TOOLS_CRON_ALLOW_COMMAND" json:"allow_command"`
}
type ExecConfig struct {
@@ -781,6 +856,7 @@ type ToolsConfig struct {
ReadFile ReadFileToolConfig `json:"read_file" envPrefix:"PICOCLAW_TOOLS_READ_FILE_"`
SendFile ToolConfig `json:"send_file" envPrefix:"PICOCLAW_TOOLS_SEND_FILE_"`
Spawn ToolConfig `json:"spawn" envPrefix:"PICOCLAW_TOOLS_SPAWN_"`
+ SpawnStatus ToolConfig `json:"spawn_status" envPrefix:"PICOCLAW_TOOLS_SPAWN_STATUS_"`
SPI ToolConfig `json:"spi" envPrefix:"PICOCLAW_TOOLS_SPI_"`
Subagent ToolConfig `json:"subagent" envPrefix:"PICOCLAW_TOOLS_SUBAGENT_"`
WebFetch ToolConfig `json:"web_fetch" envPrefix:"PICOCLAW_TOOLS_WEB_FETCH_"`
@@ -817,6 +893,10 @@ type ClawHubRegistryConfig struct {
type MCPServerConfig struct {
// Enabled indicates whether this MCP server is active
Enabled bool `json:"enabled"`
+ // Deferred controls whether this server's tools are registered as hidden (deferred/discovery mode).
+ // When nil, the global Discovery.Enabled setting applies.
+ // When explicitly set to true or false, it overrides the global setting for this server only.
+ Deferred *bool `json:"deferred,omitempty"`
// Command is the executable to run (e.g., "npx", "python", "/path/to/server")
Command string `json:"command"`
// Args are the arguments to pass to the command
@@ -870,10 +950,30 @@ func LoadConfig(path string) (*Config, error) {
return nil, err
}
+ if passphrase := credential.PassphraseProvider(); passphrase != "" {
+ for _, m := range cfg.ModelList {
+ if m.APIKey != "" && !strings.HasPrefix(m.APIKey, "enc://") &&
+ !strings.HasPrefix(m.APIKey, "file://") {
+ fmt.Fprintf(
+ os.Stderr,
+ "picoclaw: warning: model %q has a plaintext api_key; call SaveConfig to encrypt it\n",
+ m.ModelName,
+ )
+ }
+ }
+ }
+
if err := env.Parse(cfg); err != nil {
return nil, err
}
+ if err := resolveAPIKeys(cfg.ModelList, filepath.Dir(path)); err != nil {
+ return nil, err
+ }
+
+ // Expand multi-key configs into separate entries for key-level failover
+ cfg.ModelList = ExpandMultiKeyModels(cfg.ModelList)
+
// Migrate legacy channel config fields to new unified structures
cfg.migrateChannelConfigs()
@@ -882,6 +982,15 @@ func LoadConfig(path string) (*Config, error) {
cfg.ModelList = ConvertProvidersToModelList(cfg)
}
+ // Inherit credentials from providers to model_list entries (#1635).
+ // When both providers and model_list are present, model_list entries
+ // whose api_key/api_base are empty will inherit from the matching
+ // provider (matched by protocol prefix). Explicit model_list values
+ // always take precedence.
+ if cfg.HasProvidersConfig() {
+ InheritProviderCredentials(cfg.ModelList, cfg.Providers)
+ }
+
// Validate model_list for uniqueness and required fields
if err := cfg.ValidateModelList(); err != nil {
return nil, err
@@ -890,6 +999,66 @@ func LoadConfig(path string) (*Config, error) {
return cfg, nil
}
+// encryptPlaintextAPIKeys returns a copy of models with plaintext api_key values
+// encrypted. Returns (nil, nil) when nothing changed (all keys already sealed or
+// empty). Returns (nil, error) if any key fails to encrypt — callers must treat
+// this as a hard failure to prevent a mixed plaintext/ciphertext state on disk.
+// Symmetric counterpart of resolveAPIKeys: both operate purely on []ModelConfig
+// and leave JSON marshaling to the caller.
+func encryptPlaintextAPIKeys(models []ModelConfig, passphrase string) ([]ModelConfig, error) {
+ sealed := make([]ModelConfig, len(models))
+ copy(sealed, models)
+ changed := false
+ for i := range sealed {
+ m := &sealed[i]
+ if m.APIKey == "" || strings.HasPrefix(m.APIKey, "enc://") ||
+ strings.HasPrefix(m.APIKey, "file://") {
+ continue
+ }
+ encrypted, err := credential.Encrypt(passphrase, "", m.APIKey)
+ if err != nil {
+ return nil, fmt.Errorf("cannot seal api_key for model %q: %w", m.ModelName, err)
+ }
+ m.APIKey = encrypted
+ changed = true
+ }
+ if !changed {
+ return nil, nil
+ }
+ return sealed, nil
+}
+
+// resolveAPIKeys decrypts or dereferences each api_key in models in-place.
+// Supports plaintext (no-op), file:// (read from configDir), and enc:// (AES-GCM decrypt).
+// Also resolves api_keys array if present.
+func resolveAPIKeys(models []ModelConfig, configDir string) error {
+ cr := credential.NewResolver(configDir)
+ for i := range models {
+ // Resolve single APIKey
+ resolved, err := cr.Resolve(models[i].APIKey)
+ if err != nil {
+ return fmt.Errorf("model_list[%d] (%s): %w", i, models[i].ModelName, err)
+ }
+ models[i].APIKey = resolved
+
+ // Resolve APIKeys array
+ for j, key := range models[i].APIKeys {
+ resolved, err := cr.Resolve(key)
+ if err != nil {
+ return fmt.Errorf(
+ "model_list[%d] (%s): api_keys[%d]: %w",
+ i,
+ models[i].ModelName,
+ j,
+ err,
+ )
+ }
+ models[i].APIKeys[j] = resolved
+ }
+ }
+ return nil
+}
+
func (c *Config) migrateChannelConfigs() {
// Discord: mention_only -> group_trigger.mention_only
if c.Channels.Discord.MentionOnly && !c.Channels.Discord.GroupTrigger.MentionOnly {
@@ -904,12 +1073,22 @@ func (c *Config) migrateChannelConfigs() {
}
func SaveConfig(path string, cfg *Config) error {
+ if passphrase := credential.PassphraseProvider(); passphrase != "" {
+ sealed, err := encryptPlaintextAPIKeys(cfg.ModelList, passphrase)
+ if err != nil {
+ return err
+ }
+ if sealed != nil {
+ tmp := *cfg
+ tmp.ModelList = sealed
+ cfg = &tmp
+ }
+ }
+
data, err := json.MarshalIndent(cfg, "", " ")
if err != nil {
return err
}
-
- // Use unified atomic write utility with explicit sync for flash storage reliability.
return fileutil.WriteFileAtomic(path, data, 0o600)
}
@@ -991,7 +1170,7 @@ func (c *Config) GetModelConfig(modelName string) (*ModelConfig, error) {
}
// Multiple configs - use round-robin for load balancing
- idx := rrCounter.Add(1) % uint64(len(matches))
+ idx := (rrCounter.Add(1) - 1) % uint64(len(matches))
return &matches[idx], nil
}
@@ -1046,6 +1225,89 @@ func MergeAPIKeys(apiKey string, apiKeys []string) []string {
return all
}
+// ExpandMultiKeyModels expands ModelConfig entries with multiple API keys into
+// separate entries for key-level failover. Each key gets its own ModelConfig entry,
+// and the original entry's fallbacks are set up to chain through the expanded entries.
+//
+// Example: {"model_name": "gpt-4", "api_keys": ["k1", "k2", "k3"]}
+// Becomes:
+// - {"model_name": "gpt-4", "api_key": "k1", "fallbacks": ["gpt-4__key_1", "gpt-4__key_2"]}
+// - {"model_name": "gpt-4__key_1", "api_key": "k2"}
+// - {"model_name": "gpt-4__key_2", "api_key": "k3"}
+func ExpandMultiKeyModels(models []ModelConfig) []ModelConfig {
+ var expanded []ModelConfig
+
+ for _, m := range models {
+ keys := MergeAPIKeys(m.APIKey, m.APIKeys)
+
+ // Single key or no keys: keep as-is
+ if len(keys) <= 1 {
+ // Ensure APIKey is set from APIKeys if needed
+ if m.APIKey == "" && len(keys) == 1 {
+ m.APIKey = keys[0]
+ }
+ m.APIKeys = nil // Clear APIKeys to avoid confusion
+ expanded = append(expanded, m)
+ continue
+ }
+
+ // Multiple keys: expand
+ originalName := m.ModelName
+
+ // Create entries for additional keys (key_1, key_2, ...)
+ var fallbackNames []string
+ for i := 1; i < len(keys); i++ {
+ suffix := fmt.Sprintf("__key_%d", i)
+ expandedName := originalName + suffix
+
+ // Create a copy for the additional key
+ additionalEntry := ModelConfig{
+ ModelName: expandedName,
+ Model: m.Model,
+ APIBase: m.APIBase,
+ APIKey: keys[i],
+ Proxy: m.Proxy,
+ AuthMethod: m.AuthMethod,
+ ConnectMode: m.ConnectMode,
+ Workspace: m.Workspace,
+ RPM: m.RPM,
+ MaxTokensField: m.MaxTokensField,
+ RequestTimeout: m.RequestTimeout,
+ ThinkingLevel: m.ThinkingLevel,
+ }
+ expanded = append(expanded, additionalEntry)
+ fallbackNames = append(fallbackNames, expandedName)
+ }
+
+ // Create the primary entry with first key and fallbacks
+ primaryEntry := ModelConfig{
+ ModelName: originalName,
+ Model: m.Model,
+ APIBase: m.APIBase,
+ APIKey: keys[0],
+ Proxy: m.Proxy,
+ AuthMethod: m.AuthMethod,
+ ConnectMode: m.ConnectMode,
+ Workspace: m.Workspace,
+ RPM: m.RPM,
+ MaxTokensField: m.MaxTokensField,
+ RequestTimeout: m.RequestTimeout,
+ ThinkingLevel: m.ThinkingLevel,
+ }
+
+ // Prepend new fallbacks to existing ones
+ if len(fallbackNames) > 0 {
+ primaryEntry.Fallbacks = append(fallbackNames, m.Fallbacks...)
+ } else if len(m.Fallbacks) > 0 {
+ primaryEntry.Fallbacks = m.Fallbacks
+ }
+
+ expanded = append(expanded, primaryEntry)
+ }
+
+ return expanded
+}
+
func (t *ToolsConfig) IsToolEnabled(name string) bool {
switch name {
case "web":
@@ -1076,6 +1338,8 @@ func (t *ToolsConfig) IsToolEnabled(name string) bool {
return t.ReadFile.Enabled
case "spawn":
return t.Spawn.Enabled
+ case "spawn_status":
+ return t.SpawnStatus.Enabled
case "spi":
return t.SPI.Enabled
case "subagent":
diff --git a/pkg/config/config_test.go b/pkg/config/config_test.go
index caab8a152..88ab1ed51 100644
--- a/pkg/config/config_test.go
+++ b/pkg/config/config_test.go
@@ -7,8 +7,22 @@ import (
"runtime"
"strings"
"testing"
+
+ "github.com/sipeed/picoclaw/pkg/credential"
)
+// mustSetupSSHKey generates a temporary Ed25519 SSH key in t.TempDir() and sets
+// PICOCLAW_SSH_KEY_PATH to its path for the duration of the test. This is required
+// whenever a test exercises encryption/decryption via credential.Encrypt or SaveConfig.
+func mustSetupSSHKey(t *testing.T) {
+ t.Helper()
+ keyPath := filepath.Join(t.TempDir(), "picoclaw_ed25519.key")
+ if err := credential.GenerateSSHKey(keyPath); err != nil {
+ t.Fatalf("mustSetupSSHKey: %v", err)
+ }
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", keyPath)
+}
+
func TestAgentModelConfig_UnmarshalString(t *testing.T) {
var m AgentModelConfig
if err := json.Unmarshal([]byte(`"gpt-4"`), &m); err != nil {
@@ -63,6 +77,22 @@ func TestAgentModelConfig_MarshalObject(t *testing.T) {
}
}
+func TestProvidersConfig_IsEmpty(t *testing.T) {
+ var empty ProvidersConfig
+ if !empty.IsEmpty() {
+ t.Fatal("empty ProvidersConfig should report empty")
+ }
+
+ novita := ProvidersConfig{
+ Novita: ProviderConfig{
+ APIKey: "test-key",
+ },
+ }
+ if novita.IsEmpty() {
+ t.Fatal("ProvidersConfig with novita settings should not report empty")
+ }
+}
+
func TestAgentConfig_FullParse(t *testing.T) {
jsonData := `{
"agents": {
@@ -253,6 +283,9 @@ func TestDefaultConfig_Gateway(t *testing.T) {
if cfg.Gateway.Port == 0 {
t.Error("Gateway port should have default value")
}
+ if cfg.Gateway.HotReload {
+ t.Error("Gateway hot reload should be disabled by default")
+ }
}
// TestDefaultConfig_Providers verifies provider structure
@@ -384,6 +417,45 @@ func TestDefaultConfig_OpenAIWebSearchEnabled(t *testing.T) {
}
}
+func TestDefaultConfig_WebPreferNativeEnabled(t *testing.T) {
+ cfg := DefaultConfig()
+ if !cfg.Tools.Web.PreferNative {
+ t.Fatal("DefaultConfig().Tools.Web.PreferNative should be true")
+ }
+}
+
+func TestLoadConfig_WebPreferNativeDefaultsTrueWhenUnset(t *testing.T) {
+ dir := t.TempDir()
+ configPath := filepath.Join(dir, "config.json")
+ if err := os.WriteFile(configPath, []byte(`{"tools":{"web":{"enabled":true}}}`), 0o600); err != nil {
+ t.Fatalf("WriteFile() error: %v", err)
+ }
+
+ cfg, err := LoadConfig(configPath)
+ if err != nil {
+ t.Fatalf("LoadConfig() error: %v", err)
+ }
+ if !cfg.Tools.Web.PreferNative {
+ t.Fatal("PreferNative should remain true when unset in config file")
+ }
+}
+
+func TestLoadConfig_WebPreferNativeCanBeDisabled(t *testing.T) {
+ dir := t.TempDir()
+ configPath := filepath.Join(dir, "config.json")
+ if err := os.WriteFile(configPath, []byte(`{"tools":{"web":{"prefer_native":false}}}`), 0o600); err != nil {
+ t.Fatalf("WriteFile() error: %v", err)
+ }
+
+ cfg, err := LoadConfig(configPath)
+ if err != nil {
+ t.Fatalf("LoadConfig() error: %v", err)
+ }
+ if cfg.Tools.Web.PreferNative {
+ t.Fatal("PreferNative should be false when disabled in config file")
+ }
+}
+
func TestDefaultConfig_ExecAllowRemoteEnabled(t *testing.T) {
cfg := DefaultConfig()
if !cfg.Tools.Exec.AllowRemote {
@@ -391,6 +463,13 @@ func TestDefaultConfig_ExecAllowRemoteEnabled(t *testing.T) {
}
}
+func TestDefaultConfig_CronAllowCommandEnabled(t *testing.T) {
+ cfg := DefaultConfig()
+ if !cfg.Tools.Cron.AllowCommand {
+ t.Fatal("DefaultConfig().Tools.Cron.AllowCommand should be true")
+ }
+}
+
func TestDefaultConfig_HooksDefaults(t *testing.T) {
cfg := DefaultConfig()
if !cfg.Hooks.Enabled {
@@ -407,6 +486,13 @@ func TestDefaultConfig_HooksDefaults(t *testing.T) {
}
}
+func TestDefaultConfig_LogLevel(t *testing.T) {
+ cfg := DefaultConfig()
+ if cfg.Agents.Defaults.LogLevel != "fatal" {
+ t.Errorf("LogLevel = %q, want \"fatal\"", cfg.Agents.Defaults.LogLevel)
+ }
+}
+
func TestLoadConfig_OpenAIWebSearchDefaultsTrueWhenUnset(t *testing.T) {
dir := t.TempDir()
configPath := filepath.Join(dir, "config.json")
@@ -439,6 +525,22 @@ func TestLoadConfig_ExecAllowRemoteDefaultsTrueWhenUnset(t *testing.T) {
}
}
+func TestLoadConfig_CronAllowCommandDefaultsTrueWhenUnset(t *testing.T) {
+ dir := t.TempDir()
+ configPath := filepath.Join(dir, "config.json")
+ if err := os.WriteFile(configPath, []byte(`{"tools":{"cron":{"exec_timeout_minutes":5}}}`), 0o600); err != nil {
+ t.Fatalf("WriteFile() error: %v", err)
+ }
+
+ cfg, err := LoadConfig(configPath)
+ if err != nil {
+ t.Fatalf("LoadConfig() error: %v", err)
+ }
+ if !cfg.Tools.Cron.AllowCommand {
+ t.Fatal("tools.cron.allow_command should remain true when unset in config file")
+ }
+}
+
func TestLoadConfig_OpenAIWebSearchCanBeDisabled(t *testing.T) {
dir := t.TempDir()
configPath := filepath.Join(dir, "config.json")
@@ -580,13 +682,19 @@ func TestDefaultConfig_DMScope(t *testing.T) {
}
func TestDefaultConfig_WorkspacePath_Default(t *testing.T) {
- // Unset to ensure we test the default
t.Setenv("PICOCLAW_HOME", "")
- // Set a known home for consistent test results
- t.Setenv("HOME", "/tmp/home")
+
+ var fakeHome string
+ if runtime.GOOS == "windows" {
+ fakeHome = `C:\tmp\home`
+ t.Setenv("USERPROFILE", fakeHome)
+ } else {
+ fakeHome = "/tmp/home"
+ t.Setenv("HOME", fakeHome)
+ }
cfg := DefaultConfig()
- want := filepath.Join("/tmp/home", ".picoclaw", "workspace")
+ want := filepath.Join(fakeHome, ".picoclaw", "workspace")
if cfg.Agents.Defaults.Workspace != want {
t.Errorf("Default workspace path = %q, want %q", cfg.Agents.Defaults.Workspace, want)
@@ -597,7 +705,7 @@ func TestDefaultConfig_WorkspacePath_WithPicoclawHome(t *testing.T) {
t.Setenv("PICOCLAW_HOME", "/custom/picoclaw/home")
cfg := DefaultConfig()
- want := "/custom/picoclaw/home/workspace"
+ want := filepath.Join("/custom/picoclaw/home", "workspace")
if cfg.Agents.Defaults.Workspace != want {
t.Errorf("Workspace path with PICOCLAW_HOME = %q, want %q", cfg.Agents.Defaults.Workspace, want)
@@ -719,3 +827,373 @@ func TestFlexibleStringSlice_UnmarshalText_EmptySliceConsistency(t *testing.T) {
}
})
}
+
+// TestLoadConfig_WarnsForPlaintextAPIKey verifies that LoadConfig resolves a plaintext
+// api_key into memory but does NOT rewrite the config file. File writes are the sole
+// responsibility of SaveConfig.
+func TestLoadConfig_WarnsForPlaintextAPIKey(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+ const original = `{"model_list":[{"model_name":"test","model":"openai/gpt-4","api_key":"sk-plaintext"}]}`
+ if err := os.WriteFile(cfgPath, []byte(original), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "test-passphrase")
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", "")
+
+ cfg, err := LoadConfig(cfgPath)
+ if err != nil {
+ t.Fatalf("LoadConfig: %v", err)
+ }
+ // In-memory value must be the resolved plaintext.
+ if cfg.ModelList[0].APIKey != "sk-plaintext" {
+ t.Errorf("in-memory api_key = %q, want %q", cfg.ModelList[0].APIKey, "sk-plaintext")
+ }
+ // The file on disk must remain unchanged — LoadConfig must not write anything.
+ raw, _ := os.ReadFile(cfgPath)
+ if string(raw) != original {
+ t.Errorf("LoadConfig must not modify the config file; got:\n%s", string(raw))
+ }
+}
+
+// TestSaveConfig_EncryptsPlaintextAPIKey verifies that SaveConfig writes enc:// ciphertext
+// to disk and that a subsequent LoadConfig decrypts it back to the original plaintext.
+func TestSaveConfig_EncryptsPlaintextAPIKey(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "test-passphrase")
+ mustSetupSSHKey(t)
+
+ cfg := DefaultConfig()
+ cfg.ModelList = []ModelConfig{
+ {ModelName: "test", Model: "openai/gpt-4", APIKey: "sk-plaintext"},
+ }
+ if err := SaveConfig(cfgPath, cfg); err != nil {
+ t.Fatalf("SaveConfig: %v", err)
+ }
+
+ // Disk must contain enc://, not the raw key.
+ raw, _ := os.ReadFile(cfgPath)
+ if !strings.Contains(string(raw), "enc://") {
+ t.Errorf("saved file should contain enc://, got:\n%s", string(raw))
+ }
+ if strings.Contains(string(raw), "sk-plaintext") {
+ t.Errorf("saved file must not contain the plaintext key")
+ }
+
+ // A fresh load must decrypt back to the original plaintext.
+ cfg2, err := LoadConfig(cfgPath)
+ if err != nil {
+ t.Fatalf("LoadConfig after SaveConfig: %v", err)
+ }
+ if cfg2.ModelList[0].APIKey != "sk-plaintext" {
+ t.Errorf("loaded api_key = %q, want %q", cfg2.ModelList[0].APIKey, "sk-plaintext")
+ }
+}
+
+// TestLoadConfig_NoSealWithoutPassphrase verifies that api_key values are left
+// unchanged when PICOCLAW_KEY_PASSPHRASE is not set.
+func TestLoadConfig_NoSealWithoutPassphrase(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+ data := `{"model_list":[{"model_name":"test","model":"openai/gpt-4","api_key":"sk-plaintext"}]}`
+ if err := os.WriteFile(cfgPath, []byte(data), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "")
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", "")
+
+ if _, err := LoadConfig(cfgPath); err != nil {
+ t.Fatalf("LoadConfig: %v", err)
+ }
+
+ raw, _ := os.ReadFile(cfgPath)
+ if strings.Contains(string(raw), "enc://") {
+ t.Error("config file must not be modified when no passphrase is set")
+ }
+}
+
+// TestLoadConfig_FileRefNotSealed verifies that file:// api_key references are not
+// converted to enc:// values (they are resolved at runtime by the Resolver).
+func TestLoadConfig_FileRefNotSealed(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+ keyFile := filepath.Join(dir, "openai.key")
+ if err := os.WriteFile(keyFile, []byte("sk-from-file"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+ data := `{"model_list":[{"model_name":"test","model":"openai/gpt-4","api_key":"file://openai.key"}]}`
+ if err := os.WriteFile(cfgPath, []byte(data), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "test-passphrase")
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", "")
+
+ if _, err := LoadConfig(cfgPath); err != nil {
+ t.Fatalf("LoadConfig: %v", err)
+ }
+
+ raw, _ := os.ReadFile(cfgPath)
+ if !strings.Contains(string(raw), "file://openai.key") {
+ t.Error("file:// reference should be preserved unchanged in the config file")
+ }
+ if strings.Contains(string(raw), "enc://") {
+ t.Error("file:// reference must not be converted to enc://")
+ }
+}
+
+// TestSaveConfig_MixedKeys verifies that SaveConfig encrypts only plaintext api_keys
+// and leaves already-encrypted (enc://) and file:// entries unchanged.
+func TestSaveConfig_MixedKeys(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "test-passphrase")
+ mustSetupSSHKey(t)
+
+ // Pre-encrypt one key so we have a genuine enc:// value to put in the config.
+ if err := SaveConfig(cfgPath, &Config{
+ ModelList: []ModelConfig{
+ {ModelName: "pre", Model: "openai/gpt-4", APIKey: "sk-already-plain"},
+ },
+ }); err != nil {
+ t.Fatalf("setup SaveConfig: %v", err)
+ }
+ raw, _ := os.ReadFile(cfgPath)
+ // Extract the enc:// value from the saved file.
+ var tmp struct {
+ ModelList []struct {
+ APIKey string `json:"api_key"`
+ } `json:"model_list"`
+ }
+ if err := json.Unmarshal(raw, &tmp); err != nil || len(tmp.ModelList) == 0 {
+ t.Fatalf("setup: could not parse saved config: %v", err)
+ }
+ alreadyEncrypted := tmp.ModelList[0].APIKey
+ if !strings.HasPrefix(alreadyEncrypted, "enc://") {
+ t.Fatalf("setup: expected enc:// key, got %q", alreadyEncrypted)
+ }
+
+ // Build a config with three models:
+ // 1. plaintext → must be encrypted by SaveConfig
+ // 2. enc:// → must be left unchanged (already encrypted)
+ // 3. file:// → must be left unchanged (file reference)
+ keyFile := filepath.Join(dir, "api.key")
+ if err := os.WriteFile(keyFile, []byte("sk-from-file"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+ cfg := &Config{
+ ModelList: []ModelConfig{
+ {ModelName: "plain", Model: "openai/gpt-4", APIKey: "sk-new-plaintext"},
+ {ModelName: "enc", Model: "openai/gpt-4", APIKey: alreadyEncrypted},
+ {ModelName: "file", Model: "openai/gpt-4", APIKey: "file://api.key"},
+ },
+ }
+ if err := SaveConfig(cfgPath, cfg); err != nil {
+ t.Fatalf("SaveConfig: %v", err)
+ }
+
+ raw, _ = os.ReadFile(cfgPath)
+ s := string(raw)
+
+ // 1. Plaintext must be encrypted.
+ if strings.Contains(s, "sk-new-plaintext") {
+ t.Error("plaintext key must not appear in saved file")
+ }
+ // 2. The pre-existing enc:// value must still be present (byte-for-byte unchanged).
+ if !strings.Contains(s, alreadyEncrypted) {
+ t.Error("pre-existing enc:// entry must be preserved unchanged")
+ }
+ // 3. file:// must be preserved.
+ if !strings.Contains(s, "file://api.key") {
+ t.Error("file:// reference must be preserved unchanged")
+ }
+
+ // Now load and verify all three decrypt/resolve correctly.
+ cfg2, err := LoadConfig(cfgPath)
+ if err != nil {
+ t.Fatalf("LoadConfig after SaveConfig: %v", err)
+ }
+ byName := make(map[string]string)
+ for _, m := range cfg2.ModelList {
+ byName[m.ModelName] = m.APIKey
+ }
+ if byName["plain"] != "sk-new-plaintext" {
+ t.Errorf("plain model api_key = %q, want %q", byName["plain"], "sk-new-plaintext")
+ }
+ if byName["enc"] != "sk-already-plain" {
+ t.Errorf("enc model api_key = %q, want %q", byName["enc"], "sk-already-plain")
+ }
+ if byName["file"] != "sk-from-file" {
+ t.Errorf("file model api_key = %q, want %q", byName["file"], "sk-from-file")
+ }
+}
+
+// TestLoadConfig_MixedKeys_NoPassphrase verifies that when PICOCLAW_KEY_PASSPHRASE
+// is not set, enc:// entries cause LoadConfig to return an error, while plaintext
+// and file:// entries in the same config are not affected.
+func TestLoadConfig_MixedKeys_NoPassphrase(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+
+ // First encrypt a key so we have a real enc:// value.
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "test-passphrase")
+ mustSetupSSHKey(t)
+ if err := SaveConfig(cfgPath, &Config{
+ ModelList: []ModelConfig{
+ {ModelName: "m", Model: "openai/gpt-4", APIKey: "sk-secret"},
+ },
+ }); err != nil {
+ t.Fatalf("setup SaveConfig: %v", err)
+ }
+ raw, _ := os.ReadFile(cfgPath)
+ var tmp struct {
+ ModelList []struct {
+ APIKey string `json:"api_key"`
+ } `json:"model_list"`
+ }
+ if err := json.Unmarshal(raw, &tmp); err != nil {
+ t.Fatalf("setup parse: %v", err)
+ }
+ encValue := tmp.ModelList[0].APIKey
+
+ // Write a mixed config: enc:// + plaintext + file://
+ keyFile := filepath.Join(dir, "api.key")
+ if err := os.WriteFile(keyFile, []byte("sk-from-file"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+ mixed, _ := json.Marshal(map[string]any{
+ "model_list": []map[string]any{
+ {"model_name": "enc", "model": "openai/gpt-4", "api_key": encValue},
+ {"model_name": "plain", "model": "openai/gpt-4", "api_key": "sk-plain"},
+ {"model_name": "file", "model": "openai/gpt-4", "api_key": "file://api.key"},
+ },
+ })
+ if err := os.WriteFile(cfgPath, mixed, 0o600); err != nil {
+ t.Fatalf("setup write: %v", err)
+ }
+
+ // Now clear the passphrase — LoadConfig must fail because enc:// cannot be decrypted.
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "")
+
+ _, err := LoadConfig(cfgPath)
+ if err == nil {
+ t.Fatal("LoadConfig should fail when enc:// key is present and no passphrase is set")
+ }
+ if !strings.Contains(err.Error(), "passphrase required") {
+ t.Errorf("error should mention passphrase required, got: %v", err)
+ }
+}
+
+// TestSaveConfig_UsesPassphraseProvider verifies that SaveConfig encrypts plaintext
+// api_keys using credential.PassphraseProvider() rather than os.Getenv directly.
+// This matters for the launcher, which clears the environment variable and redirects
+// PassphraseProvider to an in-memory SecureStore.
+func TestSaveConfig_UsesPassphraseProvider(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+
+ // Ensure the env var is empty — passphrase must come from PassphraseProvider only.
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "")
+ mustSetupSSHKey(t)
+
+ // Replace PassphraseProvider with an in-memory function (simulating SecureStore).
+ const testPassphrase = "provider-passphrase"
+ orig := credential.PassphraseProvider
+ credential.PassphraseProvider = func() string { return testPassphrase }
+ t.Cleanup(func() { credential.PassphraseProvider = orig })
+
+ cfg := DefaultConfig()
+ cfg.ModelList = []ModelConfig{
+ {ModelName: "test", Model: "openai/gpt-4", APIKey: "sk-plaintext"},
+ }
+ if err := SaveConfig(cfgPath, cfg); err != nil {
+ t.Fatalf("SaveConfig: %v", err)
+ }
+
+ raw, _ := os.ReadFile(cfgPath)
+ if !strings.Contains(string(raw), "enc://") {
+ t.Errorf("SaveConfig should have encrypted plaintext key via PassphraseProvider; got:\n%s", raw)
+ }
+}
+
+// TestLoadConfig_UsesPassphraseProvider verifies that LoadConfig decrypts enc:// keys
+// using credential.PassphraseProvider() rather than os.Getenv directly.
+func TestLoadConfig_UsesPassphraseProvider(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+
+ // Ensure the env var is empty throughout.
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "")
+ mustSetupSSHKey(t)
+
+ const testPassphrase = "provider-passphrase"
+ const plainKey = "sk-secret"
+
+ // First, encrypt the key using the same passphrase.
+ encrypted, err := credential.Encrypt(testPassphrase, "", plainKey)
+ if err != nil {
+ t.Fatalf("Encrypt: %v", err)
+ }
+
+ raw, _ := json.Marshal(map[string]any{
+ "model_list": []map[string]any{
+ {"model_name": "test", "model": "openai/gpt-4", "api_key": encrypted},
+ },
+ })
+ if err = os.WriteFile(cfgPath, raw, 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ // Redirect PassphraseProvider — env var is empty, so without this the load would fail.
+ orig := credential.PassphraseProvider
+ credential.PassphraseProvider = func() string { return testPassphrase }
+ t.Cleanup(func() { credential.PassphraseProvider = orig })
+
+ cfg, err := LoadConfig(cfgPath)
+ if err != nil {
+ t.Fatalf("LoadConfig: %v", err)
+ }
+ if cfg.ModelList[0].APIKey != plainKey {
+ t.Errorf("api_key = %q, want %q", cfg.ModelList[0].APIKey, plainKey)
+ }
+}
+
+func TestConfigParsesLogLevel(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+ data := `{"agents":{"defaults":{"log_level":"debug"}}}`
+ if err := os.WriteFile(cfgPath, []byte(data), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ cfg, err := LoadConfig(cfgPath)
+ if err != nil {
+ t.Fatalf("LoadConfig: %v", err)
+ }
+ if cfg.Agents.Defaults.LogLevel != "debug" {
+ t.Errorf("LogLevel = %q, want \"debug\"", cfg.Agents.Defaults.LogLevel)
+ }
+}
+
+func TestConfigLogLevelEmpty(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+ data := `{}`
+ if err := os.WriteFile(cfgPath, []byte(data), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ cfg, err := LoadConfig(cfgPath)
+ if err != nil {
+ t.Fatalf("LoadConfig: %v", err)
+ }
+ // When config omits log_level, the DefaultConfig value ("fatal") is preserved.
+ if cfg.Agents.Defaults.LogLevel != "fatal" {
+ t.Errorf("LogLevel = %q, want \"fatal\"", cfg.Agents.Defaults.LogLevel)
+ }
+}
diff --git a/pkg/config/defaults.go b/pkg/config/defaults.go
index bfb54fb97..28c1efb80 100644
--- a/pkg/config/defaults.go
+++ b/pkg/config/defaults.go
@@ -15,7 +15,7 @@ func DefaultConfig() *Config {
// Determine the base path for the workspace.
// Priority: $PICOCLAW_HOME > ~/.picoclaw
var homePath string
- if picoclawHome := os.Getenv("PICOCLAW_HOME"); picoclawHome != "" {
+ if picoclawHome := os.Getenv(EnvHome); picoclawHome != "" {
homePath = picoclawHome
} else {
userHome, _ := os.UserHomeDir()
@@ -26,6 +26,7 @@ func DefaultConfig() *Config {
return &Config{
Agents: AgentsConfig{
Defaults: AgentDefaults{
+ LogLevel: "fatal",
Workspace: workspacePath,
RestrictToWorkspace: true,
Provider: "",
@@ -36,6 +37,10 @@ func DefaultConfig() *Config {
SummarizeMessageThreshold: 20,
SummarizeTokenPercent: 75,
SteeringMode: "one-at-a-time",
+ ToolFeedback: ToolFeedbackConfig{
+ Enabled: true,
+ MaxArgsLength: 300,
+ },
},
},
Bindings: []AgentBinding{},
@@ -59,6 +64,8 @@ func DefaultConfig() *Config {
Enabled: true,
Text: "Thinking... 💭",
},
+ Streaming: StreamingConfig{Enabled: true, ThrottleSeconds: 3, MinGrowthChars: 200},
+ UseMarkdownV2: false,
},
Feishu: FeishuConfig{
Enabled: false,
@@ -81,11 +88,12 @@ func DefaultConfig() *Config {
AllowFrom: FlexibleStringSlice{},
},
QQ: QQConfig{
- Enabled: false,
- AppID: "",
- AppSecret: "",
- AllowFrom: FlexibleStringSlice{},
- MaxMessageLength: 2000,
+ Enabled: false,
+ AppID: "",
+ AppSecret: "",
+ AllowFrom: FlexibleStringSlice{},
+ MaxMessageLength: 2000,
+ MaxBase64FileSizeMiB: 0,
},
DingTalk: DingTalkConfig{
Enabled: false,
@@ -158,14 +166,15 @@ func DefaultConfig() *Config {
ReplyTimeout: 5,
},
WeComAIBot: WeComAIBotConfig{
- Enabled: false,
- Token: "",
- EncodingAESKey: "",
- WebhookPath: "/webhook/wecom-aibot",
- AllowFrom: FlexibleStringSlice{},
- ReplyTimeout: 5,
- MaxSteps: 10,
- WelcomeMessage: "Hello! I'm your AI assistant. How can I help you today?",
+ Enabled: false,
+ Token: "",
+ EncodingAESKey: "",
+ WebhookPath: "/webhook/wecom-aibot",
+ AllowFrom: FlexibleStringSlice{},
+ ReplyTimeout: 5,
+ MaxSteps: 10,
+ WelcomeMessage: "Hello! I'm your AI assistant. How can I help you today?",
+ ProcessingMessage: DefaultWeComAIBotProcessingMessage,
},
Pico: PicoConfig{
Enabled: false,
@@ -393,10 +402,20 @@ func DefaultConfig() *Config {
APIBase: "http://localhost:8000/v1",
APIKey: "",
},
+
+ // Azure OpenAI - https://portal.azure.com
+ // model_name is a user-friendly alias; the model field's path after "azure/" is your deployment name
+ {
+ ModelName: "azure-gpt5",
+ Model: "azure/my-gpt5-deployment",
+ APIBase: "https://your-resource.openai.azure.com",
+ APIKey: "",
+ },
},
Gateway: GatewayConfig{
- Host: "127.0.0.1",
- Port: 18790,
+ Host: "127.0.0.1",
+ Port: 18790,
+ HotReload: false,
},
Tools: ToolsConfig{
MediaCleanup: MediaCleanupConfig{
@@ -410,8 +429,10 @@ func DefaultConfig() *Config {
ToolConfig: ToolConfig{
Enabled: true,
},
+ PreferNative: true,
Proxy: "",
FetchLimitBytes: 10 * 1024 * 1024, // 10MB by default
+ Format: "plaintext",
Brave: BraveConfig{
Enabled: false,
APIKey: "",
@@ -452,6 +473,7 @@ func DefaultConfig() *Config {
Enabled: true,
},
ExecTimeoutMinutes: 5,
+ AllowCommand: true,
},
Exec: ExecConfig{
ToolConfig: ToolConfig{
@@ -521,6 +543,9 @@ func DefaultConfig() *Config {
Spawn: ToolConfig{
Enabled: true,
},
+ SpawnStatus: ToolConfig{
+ Enabled: false,
+ },
SPI: ToolConfig{
Enabled: false, // Hardware tool - Linux only
},
diff --git a/pkg/config/envkeys.go b/pkg/config/envkeys.go
new file mode 100644
index 000000000..b04ff19f5
--- /dev/null
+++ b/pkg/config/envkeys.go
@@ -0,0 +1,37 @@
+// PicoClaw - Ultra-lightweight personal AI agent
+// License: MIT
+//
+// Copyright (c) 2026 PicoClaw contributors
+
+package config
+
+// Runtime environment variable keys for the picoclaw process.
+// These control the location of files and binaries at runtime and are read
+// directly via os.Getenv / os.LookupEnv. All picoclaw-specific keys use the
+// PICOCLAW_ prefix. Reference these constants instead of inline string
+// literals to keep all supported knobs visible in one place and to prevent
+// typos.
+const (
+ // EnvHome overrides the base directory for all picoclaw data
+ // (config, workspace, skills, auth store, …).
+ // Default: ~/.picoclaw
+ EnvHome = "PICOCLAW_HOME"
+
+ // EnvConfig overrides the full path to the JSON config file.
+ // Default: $PICOCLAW_HOME/config.json
+ EnvConfig = "PICOCLAW_CONFIG"
+
+ // EnvBuiltinSkills overrides the directory from which built-in
+ // skills are loaded.
+ // Default: /skills
+ EnvBuiltinSkills = "PICOCLAW_BUILTIN_SKILLS"
+
+ // EnvBinary overrides the path to the picoclaw executable.
+ // Used by the web launcher when spawning the gateway subprocess.
+ // Default: resolved from the same directory as the current executable.
+ EnvBinary = "PICOCLAW_BINARY"
+
+ // EnvGatewayHost overrides the host address for the gateway server.
+ // Default: "127.0.0.1"
+ EnvGatewayHost = "PICOCLAW_GATEWAY_HOST"
+)
diff --git a/pkg/config/migration.go b/pkg/config/migration.go
index c7fc214d5..832d8bf17 100644
--- a/pkg/config/migration.go
+++ b/pkg/config/migration.go
@@ -468,3 +468,84 @@ func ConvertProvidersToModelList(cfg *Config) []ModelConfig {
return result
}
+
+// protocolProviderMapping maps a model protocol prefix (the part before "/" in
+// the Model field) to a function that extracts the corresponding ProviderConfig
+// from the legacy ProvidersConfig. Used by InheritProviderCredentials.
+var protocolProviderMapping = map[string]func(p ProvidersConfig) ProviderConfig{
+ "openai": func(p ProvidersConfig) ProviderConfig { return p.OpenAI.ProviderConfig },
+ "anthropic": func(p ProvidersConfig) ProviderConfig { return p.Anthropic },
+ "litellm": func(p ProvidersConfig) ProviderConfig { return p.LiteLLM },
+ "openrouter": func(p ProvidersConfig) ProviderConfig { return p.OpenRouter },
+ "groq": func(p ProvidersConfig) ProviderConfig { return p.Groq },
+ "zhipu": func(p ProvidersConfig) ProviderConfig { return p.Zhipu },
+ "vllm": func(p ProvidersConfig) ProviderConfig { return p.VLLM },
+ "gemini": func(p ProvidersConfig) ProviderConfig { return p.Gemini },
+ "nvidia": func(p ProvidersConfig) ProviderConfig { return p.Nvidia },
+ "ollama": func(p ProvidersConfig) ProviderConfig { return p.Ollama },
+ "moonshot": func(p ProvidersConfig) ProviderConfig { return p.Moonshot },
+ "shengsuanyun": func(p ProvidersConfig) ProviderConfig { return p.ShengSuanYun },
+ "deepseek": func(p ProvidersConfig) ProviderConfig { return p.DeepSeek },
+ "cerebras": func(p ProvidersConfig) ProviderConfig { return p.Cerebras },
+ "vivgrid": func(p ProvidersConfig) ProviderConfig { return p.Vivgrid },
+ "volcengine": func(p ProvidersConfig) ProviderConfig { return p.VolcEngine },
+ "github-copilot": func(p ProvidersConfig) ProviderConfig { return p.GitHubCopilot },
+ "antigravity": func(p ProvidersConfig) ProviderConfig { return p.Antigravity },
+ "qwen": func(p ProvidersConfig) ProviderConfig { return p.Qwen },
+ "mistral": func(p ProvidersConfig) ProviderConfig { return p.Mistral },
+ "avian": func(p ProvidersConfig) ProviderConfig { return p.Avian },
+ "minimax": func(p ProvidersConfig) ProviderConfig { return p.Minimax },
+ "longcat": func(p ProvidersConfig) ProviderConfig { return p.LongCat },
+ "modelscope": func(p ProvidersConfig) ProviderConfig { return p.ModelScope },
+ "novita": func(p ProvidersConfig) ProviderConfig { return p.Novita },
+}
+
+// InheritProviderCredentials fills in missing api_key, api_base, proxy, and
+// request_timeout on model_list entries from the matching legacy providers
+// configuration. The match is determined by the protocol prefix in the Model
+// field (e.g. "deepseek/deepseek-chat" matches providers.deepseek).
+//
+// Only empty fields are filled — any value explicitly set on a model_list entry
+// takes precedence. This function modifies the slice in place.
+//
+// This bridges the gap described in issue #1635: users who configure
+// credentials once in the providers section expect model_list entries using
+// the same protocol to "just work" without duplicating credentials.
+func InheritProviderCredentials(models []ModelConfig, providers ProvidersConfig) {
+ if providers.IsEmpty() {
+ return
+ }
+
+ for i := range models {
+ m := &models[i]
+
+ // Extract protocol prefix from Model field
+ protocol := ""
+ if idx := strings.Index(m.Model, "/"); idx > 0 {
+ protocol = strings.ToLower(m.Model[:idx])
+ }
+ if protocol == "" {
+ continue
+ }
+
+ getProvider, ok := protocolProviderMapping[protocol]
+ if !ok {
+ continue
+ }
+ pc := getProvider(providers)
+
+ // Only fill empty fields — explicit model_list values win
+ if m.APIKey == "" && pc.APIKey != "" {
+ m.APIKey = pc.APIKey
+ }
+ if m.APIBase == "" && pc.APIBase != "" {
+ m.APIBase = pc.APIBase
+ }
+ if m.Proxy == "" && pc.Proxy != "" {
+ m.Proxy = pc.Proxy
+ }
+ if m.RequestTimeout == 0 && pc.RequestTimeout != 0 {
+ m.RequestTimeout = pc.RequestTimeout
+ }
+ }
+}
diff --git a/pkg/config/migration_test.go b/pkg/config/migration_test.go
index 1b6e5b032..bea5b9034 100644
--- a/pkg/config/migration_test.go
+++ b/pkg/config/migration_test.go
@@ -613,3 +613,143 @@ func TestConvertProvidersToModelList_LegacyModelWithProtocolPrefix(t *testing.T)
t.Errorf("Model = %q, want %q (should not duplicate prefix)", result[0].Model, "openrouter/auto")
}
}
+
+// ---------- InheritProviderCredentials tests ----------
+
+func TestInheritProviderCredentials_FillsMissingAPIKey(t *testing.T) {
+ models := []ModelConfig{
+ {ModelName: "my-deepseek", Model: "deepseek/deepseek-chat"},
+ }
+ providers := ProvidersConfig{
+ DeepSeek: ProviderConfig{
+ APIKey: "sk-deepseek-from-providers",
+ APIBase: "https://api.deepseek.com/v1",
+ },
+ }
+
+ InheritProviderCredentials(models, providers)
+
+ if models[0].APIKey != "sk-deepseek-from-providers" {
+ t.Errorf("APIKey = %q, want %q", models[0].APIKey, "sk-deepseek-from-providers")
+ }
+ if models[0].APIBase != "https://api.deepseek.com/v1" {
+ t.Errorf("APIBase = %q, want %q", models[0].APIBase, "https://api.deepseek.com/v1")
+ }
+}
+
+func TestInheritProviderCredentials_ExplicitValuesTakePrecedence(t *testing.T) {
+ models := []ModelConfig{
+ {
+ ModelName: "my-openai",
+ Model: "openai/gpt-5.4",
+ APIKey: "sk-explicit-model-key",
+ APIBase: "https://my-custom-endpoint.com/v1",
+ },
+ }
+ providers := ProvidersConfig{
+ OpenAI: OpenAIProviderConfig{
+ ProviderConfig: ProviderConfig{
+ APIKey: "sk-provider-key",
+ APIBase: "https://api.openai.com/v1",
+ },
+ },
+ }
+
+ InheritProviderCredentials(models, providers)
+
+ if models[0].APIKey != "sk-explicit-model-key" {
+ t.Errorf("APIKey = %q, want %q (explicit should win)", models[0].APIKey, "sk-explicit-model-key")
+ }
+ if models[0].APIBase != "https://my-custom-endpoint.com/v1" {
+ t.Errorf("APIBase = %q, want %q (explicit should win)", models[0].APIBase, "https://my-custom-endpoint.com/v1")
+ }
+}
+
+func TestInheritProviderCredentials_MultipleModels(t *testing.T) {
+ models := []ModelConfig{
+ {ModelName: "groq-llama", Model: "groq/llama-3.1-70b"},
+ {ModelName: "zhipu-glm", Model: "zhipu/glm-4"},
+ {ModelName: "custom-openai", Model: "openai/gpt-5.4", APIKey: "sk-already-set"},
+ }
+ providers := ProvidersConfig{
+ Groq: ProviderConfig{APIKey: "gsk-groq-key", Proxy: "http://proxy:8080"},
+ Zhipu: ProviderConfig{APIKey: "zhipu-key-123", APIBase: "https://zhipu.example.com"},
+ OpenAI: OpenAIProviderConfig{
+ ProviderConfig: ProviderConfig{APIKey: "sk-should-not-override"},
+ },
+ }
+
+ InheritProviderCredentials(models, providers)
+
+ // groq model should inherit
+ if models[0].APIKey != "gsk-groq-key" {
+ t.Errorf("groq APIKey = %q, want %q", models[0].APIKey, "gsk-groq-key")
+ }
+ if models[0].Proxy != "http://proxy:8080" {
+ t.Errorf("groq Proxy = %q, want %q", models[0].Proxy, "http://proxy:8080")
+ }
+
+ // zhipu model should inherit
+ if models[1].APIKey != "zhipu-key-123" {
+ t.Errorf("zhipu APIKey = %q, want %q", models[1].APIKey, "zhipu-key-123")
+ }
+ if models[1].APIBase != "https://zhipu.example.com" {
+ t.Errorf("zhipu APIBase = %q, want %q", models[1].APIBase, "https://zhipu.example.com")
+ }
+
+ // openai model already has key — should NOT be overridden
+ if models[2].APIKey != "sk-already-set" {
+ t.Errorf("openai APIKey = %q, want %q (should not be overridden)", models[2].APIKey, "sk-already-set")
+ }
+}
+
+func TestInheritProviderCredentials_NoMatchingProvider(t *testing.T) {
+ models := []ModelConfig{
+ {ModelName: "my-model", Model: "novelai/some-model"},
+ }
+ providers := ProvidersConfig{
+ DeepSeek: ProviderConfig{APIKey: "sk-deepseek"},
+ }
+
+ InheritProviderCredentials(models, providers)
+
+ // No matching provider for "novelai" protocol — should stay empty
+ if models[0].APIKey != "" {
+ t.Errorf("APIKey = %q, want empty (no matching provider)", models[0].APIKey)
+ }
+}
+
+func TestInheritProviderCredentials_EmptyProviders(t *testing.T) {
+ models := []ModelConfig{
+ {ModelName: "my-model", Model: "openai/gpt-5.4"},
+ }
+ providers := ProvidersConfig{} // all empty
+
+ InheritProviderCredentials(models, providers)
+
+ // Empty providers — nothing to inherit
+ if models[0].APIKey != "" {
+ t.Errorf("APIKey = %q, want empty", models[0].APIKey)
+ }
+}
+
+func TestInheritProviderCredentials_InheritsRequestTimeout(t *testing.T) {
+ models := []ModelConfig{
+ {ModelName: "my-ollama", Model: "ollama/llama3.2:3b"},
+ }
+ providers := ProvidersConfig{
+ Ollama: ProviderConfig{
+ APIBase: "http://localhost:11434",
+ RequestTimeout: 120,
+ },
+ }
+
+ InheritProviderCredentials(models, providers)
+
+ if models[0].APIBase != "http://localhost:11434" {
+ t.Errorf("APIBase = %q, want %q", models[0].APIBase, "http://localhost:11434")
+ }
+ if models[0].RequestTimeout != 120 {
+ t.Errorf("RequestTimeout = %d, want 120", models[0].RequestTimeout)
+ }
+}
diff --git a/pkg/config/model_config_test.go b/pkg/config/model_config_test.go
index da6e506f8..9bc600ed9 100644
--- a/pkg/config/model_config_test.go
+++ b/pkg/config/model_config_test.go
@@ -80,6 +80,36 @@ func TestGetModelConfig_RoundRobin(t *testing.T) {
}
}
+func TestGetModelConfig_RoundRobinStartsFromFirstMatch(t *testing.T) {
+ rrCounter.Store(0)
+
+ cfg := &Config{
+ ModelList: []ModelConfig{
+ {ModelName: "lb-model", Model: "openai/gpt-4o-1", APIKey: "key1"},
+ {ModelName: "lb-model", Model: "openai/gpt-4o-2", APIKey: "key2"},
+ {ModelName: "lb-model", Model: "openai/gpt-4o-3", APIKey: "key3"},
+ },
+ }
+
+ wantOrder := []string{
+ "openai/gpt-4o-1",
+ "openai/gpt-4o-2",
+ "openai/gpt-4o-3",
+ "openai/gpt-4o-1",
+ "openai/gpt-4o-2",
+ }
+
+ for i, want := range wantOrder {
+ result, err := cfg.GetModelConfig("lb-model")
+ if err != nil {
+ t.Fatalf("GetModelConfig() call %d error = %v", i, err)
+ }
+ if result.Model != want {
+ t.Fatalf("GetModelConfig() call %d model = %q, want %q", i, result.Model, want)
+ }
+ }
+}
+
func TestGetModelConfig_Concurrent(t *testing.T) {
cfg := &Config{
ModelList: []ModelConfig{
diff --git a/pkg/config/multikey_test.go b/pkg/config/multikey_test.go
new file mode 100644
index 000000000..b899b991c
--- /dev/null
+++ b/pkg/config/multikey_test.go
@@ -0,0 +1,291 @@
+package config
+
+import (
+ "testing"
+)
+
+func TestExpandMultiKeyModels_SingleKey(t *testing.T) {
+ models := []ModelConfig{
+ {
+ ModelName: "gpt-4",
+ Model: "openai/gpt-4o",
+ APIKey: "single-key",
+ },
+ }
+
+ result := ExpandMultiKeyModels(models)
+
+ if len(result) != 1 {
+ t.Fatalf("expected 1 model, got %d", len(result))
+ }
+
+ if result[0].ModelName != "gpt-4" {
+ t.Errorf("expected model_name 'gpt-4', got %q", result[0].ModelName)
+ }
+
+ if result[0].APIKey != "single-key" {
+ t.Errorf("expected api_key 'single-key', got %q", result[0].APIKey)
+ }
+
+ if len(result[0].Fallbacks) != 0 {
+ t.Errorf("expected no fallbacks, got %v", result[0].Fallbacks)
+ }
+}
+
+func TestExpandMultiKeyModels_APIKeysOnly(t *testing.T) {
+ models := []ModelConfig{
+ {
+ ModelName: "glm-4.7",
+ Model: "zhipu/glm-4.7",
+ APIBase: "https://api.example.com",
+ APIKeys: []string{"key1", "key2", "key3"},
+ },
+ }
+
+ result := ExpandMultiKeyModels(models)
+
+ // Should expand to 3 models
+ if len(result) != 3 {
+ t.Fatalf("expected 3 models, got %d", len(result))
+ }
+
+ // First entry should be the primary with key1 and fallbacks
+ primary := result[2] // Primary is added last
+ if primary.ModelName != "glm-4.7" {
+ t.Errorf("expected primary model_name 'glm-4.7', got %q", primary.ModelName)
+ }
+ if primary.APIKey != "key1" {
+ t.Errorf("expected primary api_key 'key1', got %q", primary.APIKey)
+ }
+ if len(primary.Fallbacks) != 2 {
+ t.Errorf("expected 2 fallbacks, got %d", len(primary.Fallbacks))
+ }
+ if primary.Fallbacks[0] != "glm-4.7__key_1" {
+ t.Errorf("expected first fallback 'glm-4.7__key_1', got %q", primary.Fallbacks[0])
+ }
+ if primary.Fallbacks[1] != "glm-4.7__key_2" {
+ t.Errorf("expected second fallback 'glm-4.7__key_2', got %q", primary.Fallbacks[1])
+ }
+
+ // Second entry should be key2
+ second := result[0]
+ if second.ModelName != "glm-4.7__key_1" {
+ t.Errorf("expected second model_name 'glm-4.7__key_1', got %q", second.ModelName)
+ }
+ if second.APIKey != "key2" {
+ t.Errorf("expected second api_key 'key2', got %q", second.APIKey)
+ }
+
+ // Third entry should be key3
+ third := result[1]
+ if third.ModelName != "glm-4.7__key_2" {
+ t.Errorf("expected third model_name 'glm-4.7__key_2', got %q", third.ModelName)
+ }
+ if third.APIKey != "key3" {
+ t.Errorf("expected third api_key 'key3', got %q", third.APIKey)
+ }
+}
+
+func TestExpandMultiKeyModels_APIKeyAndAPIKeys(t *testing.T) {
+ models := []ModelConfig{
+ {
+ ModelName: "gpt-4",
+ Model: "openai/gpt-4o",
+ APIKey: "key0",
+ APIKeys: []string{"key1", "key2"},
+ },
+ }
+
+ result := ExpandMultiKeyModels(models)
+
+ // Should expand to 3 models (key0 from APIKey + key1, key2 from APIKeys)
+ if len(result) != 3 {
+ t.Fatalf("expected 3 models, got %d", len(result))
+ }
+
+ // Primary should use key0
+ primary := result[2]
+ if primary.APIKey != "key0" {
+ t.Errorf("expected primary api_key 'key0', got %q", primary.APIKey)
+ }
+ if len(primary.Fallbacks) != 2 {
+ t.Errorf("expected 2 fallbacks, got %d", len(primary.Fallbacks))
+ }
+}
+
+func TestExpandMultiKeyModels_WithExistingFallbacks(t *testing.T) {
+ models := []ModelConfig{
+ {
+ ModelName: "gpt-4",
+ Model: "openai/gpt-4o",
+ APIKeys: []string{"key1", "key2"},
+ Fallbacks: []string{"claude-3"},
+ },
+ }
+
+ result := ExpandMultiKeyModels(models)
+
+ primary := result[1]
+ // With 2 keys, we get 1 key fallback + 1 existing fallback = 2 total
+ if len(primary.Fallbacks) != 2 {
+ t.Fatalf("expected 2 fallbacks, got %d: %v", len(primary.Fallbacks), primary.Fallbacks)
+ }
+
+ // Key fallbacks should come first, then existing fallbacks
+ if primary.Fallbacks[0] != "gpt-4__key_1" {
+ t.Errorf("expected first fallback 'gpt-4__key_1', got %q", primary.Fallbacks[0])
+ }
+ if primary.Fallbacks[1] != "claude-3" {
+ t.Errorf("expected second fallback 'claude-3', got %q", primary.Fallbacks[1])
+ }
+}
+
+func TestExpandMultiKeyModels_EmptyAPIKeys(t *testing.T) {
+ models := []ModelConfig{
+ {
+ ModelName: "gpt-4",
+ Model: "openai/gpt-4o",
+ APIKey: "",
+ APIKeys: []string{},
+ },
+ }
+
+ result := ExpandMultiKeyModels(models)
+
+ // Should keep as-is with no changes
+ if len(result) != 1 {
+ t.Fatalf("expected 1 model, got %d", len(result))
+ }
+
+ if result[0].ModelName != "gpt-4" {
+ t.Errorf("expected model_name 'gpt-4', got %q", result[0].ModelName)
+ }
+}
+
+func TestExpandMultiKeyModels_Deduplication(t *testing.T) {
+ models := []ModelConfig{
+ {
+ ModelName: "gpt-4",
+ Model: "openai/gpt-4o",
+ APIKey: "key1",
+ APIKeys: []string{"key1", "key2", "key1"}, // Duplicate key1
+ },
+ }
+
+ result := ExpandMultiKeyModels(models)
+
+ // Should only create 2 models (deduplicated keys)
+ if len(result) != 2 {
+ t.Fatalf("expected 2 models (deduplicated), got %d", len(result))
+ }
+
+ primary := result[1]
+ if primary.APIKey != "key1" {
+ t.Errorf("expected primary api_key 'key1', got %q", primary.APIKey)
+ }
+ if len(primary.Fallbacks) != 1 {
+ t.Errorf("expected 1 fallback, got %d", len(primary.Fallbacks))
+ }
+}
+
+func TestExpandMultiKeyModels_PreservesOtherFields(t *testing.T) {
+ models := []ModelConfig{
+ {
+ ModelName: "gpt-4",
+ Model: "openai/gpt-4o",
+ APIBase: "https://api.example.com",
+ APIKeys: []string{"key1", "key2"},
+ Proxy: "http://proxy:8080",
+ RPM: 60,
+ MaxTokensField: "max_completion_tokens",
+ RequestTimeout: 30,
+ ThinkingLevel: "high",
+ },
+ }
+
+ result := ExpandMultiKeyModels(models)
+
+ // Check primary entry preserves all fields
+ primary := result[1]
+ if primary.APIBase != "https://api.example.com" {
+ t.Errorf("expected api_base preserved, got %q", primary.APIBase)
+ }
+ if primary.Proxy != "http://proxy:8080" {
+ t.Errorf("expected proxy preserved, got %q", primary.Proxy)
+ }
+ if primary.RPM != 60 {
+ t.Errorf("expected rpm preserved, got %d", primary.RPM)
+ }
+ if primary.MaxTokensField != "max_completion_tokens" {
+ t.Errorf("expected max_tokens_field preserved, got %q", primary.MaxTokensField)
+ }
+ if primary.RequestTimeout != 30 {
+ t.Errorf("expected request_timeout preserved, got %d", primary.RequestTimeout)
+ }
+ if primary.ThinkingLevel != "high" {
+ t.Errorf("expected thinking_level preserved, got %q", primary.ThinkingLevel)
+ }
+
+ // Check additional entry also preserves fields
+ additional := result[0]
+ if additional.APIBase != "https://api.example.com" {
+ t.Errorf("expected additional api_base preserved, got %q", additional.APIBase)
+ }
+ if additional.RPM != 60 {
+ t.Errorf("expected additional rpm preserved, got %d", additional.RPM)
+ }
+}
+
+func TestMergeAPIKeys(t *testing.T) {
+ tests := []struct {
+ name string
+ apiKey string
+ apiKeys []string
+ expected []string
+ }{
+ {
+ name: "both empty",
+ apiKey: "",
+ apiKeys: nil,
+ expected: nil,
+ },
+ {
+ name: "only apiKey",
+ apiKey: "key1",
+ apiKeys: nil,
+ expected: []string{"key1"},
+ },
+ {
+ name: "only apiKeys",
+ apiKey: "",
+ apiKeys: []string{"key1", "key2"},
+ expected: []string{"key1", "key2"},
+ },
+ {
+ name: "both with overlap",
+ apiKey: "key1",
+ apiKeys: []string{"key1", "key2", "key3"},
+ expected: []string{"key1", "key2", "key3"},
+ },
+ {
+ name: "with whitespace",
+ apiKey: " key1 ",
+ apiKeys: []string{" key2 ", " key1 "},
+ expected: []string{"key1", "key2"},
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ result := MergeAPIKeys(tt.apiKey, tt.apiKeys)
+ if len(result) != len(tt.expected) {
+ t.Fatalf("expected %d keys, got %d", len(tt.expected), len(result))
+ }
+ for i, k := range result {
+ if k != tt.expected[i] {
+ t.Errorf("expected key[%d] = %q, got %q", i, tt.expected[i], k)
+ }
+ }
+ })
+ }
+}
diff --git a/pkg/credential/credential.go b/pkg/credential/credential.go
new file mode 100644
index 000000000..b65c19446
--- /dev/null
+++ b/pkg/credential/credential.go
@@ -0,0 +1,342 @@
+// Package credential resolves API credential values for model_list entries.
+//
+// An API key is a form of authorization credential. This package centralizes
+// how raw credential strings—plaintext or file references—are resolved into
+// their actual values, keeping that logic out of the config loader.
+//
+// Supported formats for the api_key field:
+//
+// - Plaintext: "sk-abc123" → returned as-is
+// - File ref: "file://filename.key" → content read from configDir/filename.key
+// - Encrypted: "enc://" → AES-256-GCM decrypt via PICOCLAW_KEY_PASSPHRASE
+// - Empty: "" → returned as-is (auth_method=oauth etc.)
+//
+// Encryption uses AES-256-GCM with HKDF-SHA256 key derivation (< 1ms, safe for embedded Linux).
+// An SSH private key is required for both encryption and decryption.
+// Key derivation:
+//
+// HKDF-SHA256(ikm=HMAC-SHA256(SHA256(sshKeyBytes), passphrase), salt, info)
+//
+// SSH key path resolution priority:
+//
+// 1. sshKeyPath argument to Encrypt (explicit)
+// 2. PICOCLAW_SSH_KEY_PATH env var
+// 3. ~/.ssh/picoclaw_ed25519.key (os.UserHomeDir is cross-platform)
+package credential
+
+import (
+ "crypto/aes"
+ "crypto/cipher"
+ "crypto/hkdf"
+ "crypto/hmac"
+ "crypto/rand"
+ "crypto/sha256"
+ "encoding/base64"
+ "errors"
+ "fmt"
+ "io"
+ "os"
+ "path/filepath"
+ "strings"
+)
+
+// PassphraseEnvVar is the environment variable that holds the encryption passphrase.
+// Other packages (e.g. config) reference this constant to avoid duplicating the string.
+const PassphraseEnvVar = "PICOCLAW_KEY_PASSPHRASE"
+
+// PassphraseProvider is the function used to retrieve the passphrase for enc://
+// credential decryption. It defaults to reading PICOCLAW_KEY_PASSPHRASE from the
+// process environment. Replace it at startup to use a different source, such as
+// an in-memory SecureStore, so that all LoadConfig() calls everywhere share the
+// same passphrase source without needing os.Environ.
+//
+// Example (launcher main.go):
+//
+// credential.PassphraseProvider = apiHandler.passphraseStore.Get
+var PassphraseProvider func() string = func() string {
+ return os.Getenv(PassphraseEnvVar)
+}
+
+// ErrPassphraseRequired is returned when an enc:// credential is encountered but
+// no passphrase is available from PassphraseProvider. Callers can detect this
+// with errors.Is to distinguish a missing-passphrase condition from other errors.
+var ErrPassphraseRequired = errors.New("credential: enc:// passphrase required")
+
+// ErrDecryptionFailed is returned when an enc:// credential cannot be decrypted,
+// indicating a wrong passphrase or SSH key. Callers can detect this with errors.Is.
+var ErrDecryptionFailed = errors.New("credential: enc:// decryption failed (wrong passphrase or SSH key?)")
+
+// SSHKeyPathEnvVar is the environment variable that specifies the path to the
+// SSH private key used for enc:// credential encryption and decryption.
+const SSHKeyPathEnvVar = "PICOCLAW_SSH_KEY_PATH"
+
+// picoclawHome is a package-local copy of config.EnvHome. It is kept here to
+// avoid a circular import between pkg/credential and pkg/config.
+const picoclawHome = "PICOCLAW_HOME"
+
+const (
+ fileScheme = "file://"
+ encScheme = "enc://"
+ hkdfInfo = "picoclaw-credential-v1"
+ saltLen = 16
+ nonceLen = 12
+ keyLen = 32
+)
+
+// Resolver resolves raw credential strings for model_list api_key fields.
+// File references are resolved relative to the directory of the config file.
+type Resolver struct {
+ configDir string
+ resolvedConfigDir string // symlink-resolved form of configDir
+}
+
+// NewResolver returns a Resolver that resolves file:// references relative to
+// configDir (typically filepath.Dir of the config file path).
+func NewResolver(configDir string) *Resolver {
+ resolved := configDir
+ if configDir != "" {
+ if linkedPath, err := filepath.EvalSymlinks(configDir); err == nil {
+ resolved = linkedPath
+ }
+ }
+ return &Resolver{configDir: configDir, resolvedConfigDir: resolved}
+}
+
+// Resolve returns the actual credential value for raw:
+//
+// - "" → "" (no error; auth_method=oauth needs no key)
+// - "file://name.key" → trimmed content of configDir/name.key
+// - anything else → raw unchanged (plaintext credential)
+func (r *Resolver) Resolve(raw string) (string, error) {
+ if raw == "" {
+ return "", nil
+ }
+
+ if strings.HasPrefix(raw, fileScheme) {
+ fileName := strings.TrimSpace(strings.TrimPrefix(raw, fileScheme))
+ if fileName == "" {
+ return "", fmt.Errorf("credential: file:// reference has no filename")
+ }
+
+ baseDir := r.resolvedConfigDir
+ if baseDir == "" {
+ baseDir = r.configDir
+ }
+ keyPath := filepath.Join(baseDir, fileName)
+ // Resolve symlinks before enforcing containment to prevent escaping via symlinks.
+ realKeyPath, err := filepath.EvalSymlinks(keyPath)
+ if err != nil {
+ return "", fmt.Errorf("credential: failed to resolve credential file path %q: %w", keyPath, err)
+ }
+ if !isWithinDir(realKeyPath, baseDir) {
+ return "", fmt.Errorf("credential: file:// path escapes config directory")
+ }
+ data, err := os.ReadFile(realKeyPath)
+ if err != nil {
+ return "", fmt.Errorf("credential: failed to read credential file %q: %w", realKeyPath, err)
+ }
+
+ value := strings.TrimSpace(string(data))
+ if value == "" {
+ return "", fmt.Errorf("credential: credential file %q is empty", realKeyPath)
+ }
+
+ return value, nil
+ }
+
+ if strings.HasPrefix(raw, encScheme) {
+ return resolveEncrypted(raw)
+ }
+
+ // Plaintext credential — return unchanged.
+ return raw, nil
+}
+
+// resolveEncrypted decrypts an enc:// credential using PassphraseProvider.
+func resolveEncrypted(raw string) (string, error) {
+ passphrase := PassphraseProvider()
+ if passphrase == "" {
+ return "", ErrPassphraseRequired
+ }
+
+ sshKeyPath := pickSSHKeyPath("") // override="": consult env then auto-detect
+
+ b64 := strings.TrimPrefix(raw, encScheme)
+ blob, err := base64.StdEncoding.DecodeString(b64)
+ if err != nil {
+ return "", fmt.Errorf("credential: enc:// invalid base64: %w", err)
+ }
+ if len(blob) < saltLen+nonceLen+1 {
+ return "", fmt.Errorf("credential: enc:// payload too short")
+ }
+
+ salt := blob[:saltLen]
+ nonce := blob[saltLen : saltLen+nonceLen]
+ ciphertext := blob[saltLen+nonceLen:]
+
+ key, err := deriveKey(passphrase, sshKeyPath, salt)
+ if err != nil {
+ return "", err
+ }
+ block, err := aes.NewCipher(key)
+ if err != nil {
+ return "", fmt.Errorf("credential: enc:// cipher init: %w", err)
+ }
+ gcm, err := cipher.NewGCM(block)
+ if err != nil {
+ return "", fmt.Errorf("credential: enc:// gcm init: %w", err)
+ }
+
+ plaintext, err := gcm.Open(nil, nonce, ciphertext, nil)
+ if err != nil {
+ return "", fmt.Errorf("%w: %w", ErrDecryptionFailed, err)
+ }
+ return string(plaintext), nil
+}
+
+// Encrypt encrypts plaintext and returns an enc:// credential string.
+//
+// passphrase is required (PICOCLAW_KEY_PASSPHRASE value).
+// sshKeyPath is the SSH private key file to use; pass "" to auto-detect via
+// PICOCLAW_SSH_KEY_PATH env var or ~/.ssh/picoclaw_ed25519.key.
+// An SSH private key must be resolvable or Encrypt returns an error.
+func Encrypt(passphrase, sshKeyPath, plaintext string) (string, error) {
+ if passphrase == "" {
+ return "", fmt.Errorf("credential: passphrase must not be empty")
+ }
+ sshKeyPath = pickSSHKeyPath(sshKeyPath)
+
+ salt := make([]byte, saltLen)
+ if _, err := io.ReadFull(rand.Reader, salt); err != nil {
+ return "", fmt.Errorf("credential: failed to generate salt: %w", err)
+ }
+
+ key, err := deriveKey(passphrase, sshKeyPath, salt)
+ if err != nil {
+ return "", err
+ }
+ block, err := aes.NewCipher(key)
+ if err != nil {
+ return "", fmt.Errorf("credential: cipher init: %w", err)
+ }
+ gcm, err := cipher.NewGCM(block)
+ if err != nil {
+ return "", fmt.Errorf("credential: gcm init: %w", err)
+ }
+
+ nonce := make([]byte, nonceLen)
+ if _, err := io.ReadFull(rand.Reader, nonce); err != nil {
+ return "", fmt.Errorf("credential: failed to generate nonce: %w", err)
+ }
+
+ ciphertext := gcm.Seal(nil, nonce, []byte(plaintext), nil)
+ blob := make([]byte, 0, saltLen+nonceLen+len(ciphertext))
+ blob = append(blob, salt...)
+ blob = append(blob, nonce...)
+ blob = append(blob, ciphertext...)
+ return encScheme + base64.StdEncoding.EncodeToString(blob), nil
+}
+
+// isWithinDir reports whether path is contained within (or equal to) dir.
+// Uses filepath.IsLocal on the relative path for robust cross-platform traversal detection.
+func isWithinDir(path, dir string) bool {
+ rel, err := filepath.Rel(filepath.Clean(dir), filepath.Clean(path))
+ return err == nil && filepath.IsLocal(rel)
+}
+
+// allowedSSHKeyPath reports whether path is in a permitted location for SSH key files:
+// - exact match with PICOCLAW_SSH_KEY_PATH env var
+// - within the PICOCLAW_HOME env var directory
+// - within ~/.ssh/
+func allowedSSHKeyPath(path string) bool {
+ if path == "" {
+ return true // passphrase-only mode; no file will be read
+ }
+ clean := filepath.Clean(path)
+
+ // Exact match with PICOCLAW_SSH_KEY_PATH.
+ if envPath, ok := os.LookupEnv(SSHKeyPathEnvVar); ok && envPath != "" {
+ if clean == filepath.Clean(envPath) {
+ return true
+ }
+ }
+
+ // Within PICOCLAW_HOME.
+ if picoHome := os.Getenv(picoclawHome); picoHome != "" {
+ if isWithinDir(clean, picoHome) {
+ return true
+ }
+ }
+
+ // Within ~/.ssh/.
+ if userHome, err := os.UserHomeDir(); err == nil {
+ if isWithinDir(clean, filepath.Join(userHome, ".ssh")) {
+ return true
+ }
+ }
+
+ return false
+}
+
+// deriveKey derives a 32-byte AES-256 key from passphrase and SSH private key.
+//
+// ikm = HMAC-SHA256(key=SHA256(sshKeyBytes), msg=passphrase)
+// Final key: HKDF-SHA256(ikm, salt, info="picoclaw-credential-v1", 32 bytes)
+// sshKeyPath must be non-empty; returns an error otherwise.
+func deriveKey(passphrase, sshKeyPath string, salt []byte) ([]byte, error) {
+ if sshKeyPath == "" {
+ return nil, fmt.Errorf(
+ "credential: SSH private key is required but not found" +
+ " (set PICOCLAW_SSH_KEY_PATH or place key at ~/.ssh/picoclaw_ed25519.key)")
+ }
+ if !allowedSSHKeyPath(sshKeyPath) {
+ return nil, fmt.Errorf(
+ "credential: SSH key path %q is not in an allowed location (PICOCLAW_SSH_KEY_PATH, PICOCLAW_HOME, or ~/.ssh/)",
+ sshKeyPath,
+ )
+ }
+ sshBytes, err := os.ReadFile(sshKeyPath)
+ if err != nil {
+ return nil, fmt.Errorf("credential: cannot read SSH key %q: %w", sshKeyPath, err)
+ }
+ sshHash := sha256.Sum256(sshBytes)
+ mac := hmac.New(sha256.New, sshHash[:])
+ mac.Write([]byte(passphrase))
+ ikm := mac.Sum(nil)
+
+ key, err := hkdf.Key(sha256.New, ikm, salt, hkdfInfo, keyLen)
+ if err != nil {
+ return nil, fmt.Errorf("credential: HKDF expand failed: %w", err)
+ }
+ return key, nil
+}
+
+// pickSSHKeyPath returns the SSH private key path to use for encryption/decryption.
+//
+// Priority:
+// 1. override (non-empty explicit argument)
+// 2. PICOCLAW_SSH_KEY_PATH env var
+// 3. ~/.ssh/picoclaw_ed25519.key (auto-detection)
+//
+// Returns "" when no key is found; deriveKey will return an error in that case.
+func pickSSHKeyPath(override string) string {
+ if override != "" {
+ return override
+ }
+ if p, ok := os.LookupEnv(SSHKeyPathEnvVar); ok {
+ return p // respect explicit setting, even if ""
+ }
+ return findDefaultSSHKey()
+}
+
+// findDefaultSSHKey returns the picoclaw-specific SSH key path if it exists.
+func findDefaultSSHKey() string {
+ p, err := DefaultSSHKeyPath()
+ if err != nil {
+ return ""
+ }
+ if _, err := os.Stat(p); err == nil {
+ return p
+ }
+ return ""
+}
diff --git a/pkg/credential/credential_test.go b/pkg/credential/credential_test.go
new file mode 100644
index 000000000..138af3134
--- /dev/null
+++ b/pkg/credential/credential_test.go
@@ -0,0 +1,283 @@
+package credential_test
+
+import (
+ "os"
+ "path/filepath"
+ "testing"
+
+ "github.com/sipeed/picoclaw/pkg/credential"
+)
+
+func TestResolve_PlainKey(t *testing.T) {
+ r := credential.NewResolver(t.TempDir())
+ got, err := r.Resolve("sk-plaintext-key")
+ if err != nil {
+ t.Fatalf("unexpected error: %v", err)
+ }
+ if got != "sk-plaintext-key" {
+ t.Fatalf("got %q, want %q", got, "sk-plaintext-key")
+ }
+}
+
+func TestResolve_FileKey_Success(t *testing.T) {
+ dir := t.TempDir()
+ keyFile := "openai_plain.key"
+ if err := os.WriteFile(filepath.Join(dir, keyFile), []byte("sk-from-file\n"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ r := credential.NewResolver(dir)
+ got, err := r.Resolve("file://" + keyFile)
+ if err != nil {
+ t.Fatalf("unexpected error: %v", err)
+ }
+ if got != "sk-from-file" {
+ t.Fatalf("got %q, want %q", got, "sk-from-file")
+ }
+}
+
+func TestResolve_FileKey_NotFound(t *testing.T) {
+ r := credential.NewResolver(t.TempDir())
+ _, err := r.Resolve("file://missing.key")
+ if err == nil {
+ t.Fatal("expected error for missing file, got nil")
+ }
+}
+
+func TestResolve_FileKey_Empty(t *testing.T) {
+ dir := t.TempDir()
+ keyFile := "empty.key"
+ if err := os.WriteFile(filepath.Join(dir, keyFile), []byte(" \n"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ r := credential.NewResolver(dir)
+ _, err := r.Resolve("file://" + keyFile)
+ if err == nil {
+ t.Fatal("expected error for empty credential file, got nil")
+ }
+}
+
+// TestResolve_EncKey_RoundTrip tests basic encryption/decryption round-trip with an SSH key.
+func TestResolve_EncKey_RoundTrip(t *testing.T) {
+ dir := t.TempDir()
+ sshKeyPath := filepath.Join(dir, "picoclaw_ed25519.key")
+ if err := os.WriteFile(sshKeyPath, []byte("fake-ssh-key-material\n"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ const passphrase = "test-passphrase-32bytes-long-ok!"
+ const plaintext = "sk-encrypted-secret"
+
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", sshKeyPath)
+
+ enc, err := credential.Encrypt(passphrase, "", plaintext)
+ if err != nil {
+ t.Fatalf("Encrypt: %v", err)
+ }
+
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", passphrase)
+
+ r := credential.NewResolver(t.TempDir())
+ got, err := r.Resolve(enc)
+ if err != nil {
+ t.Fatalf("Resolve: %v", err)
+ }
+ if got != plaintext {
+ t.Fatalf("got %q, want %q", got, plaintext)
+ }
+}
+
+// TestResolve_EncKey_WithSSHKey tests that the SSH key file is incorporated into key derivation.
+func TestResolve_EncKey_WithSSHKey(t *testing.T) {
+ dir := t.TempDir()
+ sshKeyPath := filepath.Join(dir, "picoclaw_ed25519.key")
+ if err := os.WriteFile(sshKeyPath, []byte("fake-ssh-private-key-material\n"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ const passphrase = "test-passphrase"
+ const plaintext = "sk-ssh-protected-secret"
+
+ // Set PICOCLAW_SSH_KEY_PATH before Encrypt so the path passes allowedSSHKeyPath validation.
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", passphrase)
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", sshKeyPath)
+
+ enc, err := credential.Encrypt(passphrase, sshKeyPath, plaintext)
+ if err != nil {
+ t.Fatalf("Encrypt: %v", err)
+ }
+
+ r := credential.NewResolver(t.TempDir())
+ got, err := r.Resolve(enc)
+ if err != nil {
+ t.Fatalf("Resolve: %v", err)
+ }
+ if got != plaintext {
+ t.Fatalf("got %q, want %q", got, plaintext)
+ }
+}
+
+func TestResolve_EncKey_NoPassphrase(t *testing.T) {
+ dir := t.TempDir()
+ sshKeyPath := filepath.Join(dir, "picoclaw_ed25519.key")
+ if err := os.WriteFile(sshKeyPath, []byte("fake-ssh-key\n"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", sshKeyPath)
+
+ enc, err := credential.Encrypt("some-passphrase", "", "sk-secret")
+ if err != nil {
+ t.Fatalf("Encrypt: %v", err)
+ }
+
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "")
+
+ r := credential.NewResolver(t.TempDir())
+ _, err = r.Resolve(enc)
+ if err == nil {
+ t.Fatal("expected error when PICOCLAW_KEY_PASSPHRASE is unset, got nil")
+ }
+}
+
+func TestResolve_EncKey_BadCiphertext(t *testing.T) {
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "some-passphrase")
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", "")
+
+ r := credential.NewResolver(t.TempDir())
+ _, err := r.Resolve("enc://!!not-valid-base64!!")
+ if err == nil {
+ t.Fatal("expected error for invalid enc:// payload, got nil")
+ }
+}
+
+func TestResolve_EncKey_PayloadTooShort(t *testing.T) {
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "some-passphrase")
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", "")
+
+ // Valid base64 but fewer bytes than salt(16)+nonce(12)+1 minimum.
+ import64 := "dG9vc2hvcnQ=" // "tooshort" = 8 bytes
+ r := credential.NewResolver(t.TempDir())
+ _, err := r.Resolve("enc://" + import64)
+ if err == nil {
+ t.Fatal("expected error for too-short enc:// payload, got nil")
+ }
+}
+
+func TestResolve_EncKey_WrongPassphrase(t *testing.T) {
+ dir := t.TempDir()
+ sshKeyPath := filepath.Join(dir, "picoclaw_ed25519.key")
+ if err := os.WriteFile(sshKeyPath, []byte("fake-ssh-key\n"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", sshKeyPath)
+
+ enc, err := credential.Encrypt("correct-passphrase", "", "sk-secret")
+ if err != nil {
+ t.Fatalf("Encrypt: %v", err)
+ }
+
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "wrong-passphrase")
+
+ r := credential.NewResolver(t.TempDir())
+ _, err = r.Resolve(enc)
+ if err == nil {
+ t.Fatal("expected decryption error for wrong passphrase, got nil")
+ }
+}
+
+func TestEncrypt_EmptyPassphrase(t *testing.T) {
+ _, err := credential.Encrypt("", "", "sk-secret")
+ if err == nil {
+ t.Fatal("expected error for empty passphrase, got nil")
+ }
+}
+
+func TestDeriveKey_SSHKeyNotFound(t *testing.T) {
+ // Encrypt with a real SSH key path, then try to decrypt with a missing path.
+ dir := t.TempDir()
+ sshKeyPath := filepath.Join(dir, "picoclaw_ed25519.key")
+ if err := os.WriteFile(sshKeyPath, []byte("fake-key\n"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ // Register the real key path so allowedSSHKeyPath validation passes for Encrypt.
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", sshKeyPath)
+
+ enc, err := credential.Encrypt("passphrase", sshKeyPath, "sk-secret")
+ if err != nil {
+ t.Fatalf("Encrypt: %v", err)
+ }
+
+ // Point to a non-existent SSH key so deriveKey's ReadFile fails.
+ // The path is still under the same dir, so allowedSSHKeyPath passes (exact env match).
+ t.Setenv("PICOCLAW_KEY_PASSPHRASE", "passphrase")
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", filepath.Join(dir, "nonexistent_key"))
+
+ r := credential.NewResolver(t.TempDir())
+ _, err = r.Resolve(enc)
+ if err == nil {
+ t.Fatal("expected error when SSH key file is missing, got nil")
+ }
+}
+
+// TestResolve_FileRef_PathTraversal verifies that file:// references cannot escape configDir
+// via relative traversal ("../../etc/passwd") or absolute paths ("/abs/path").
+func TestResolve_FileRef_PathTraversal(t *testing.T) {
+ dir := t.TempDir()
+ cfgPath := filepath.Join(dir, "config.json")
+ // Create a file outside configDir that the traversal would point to.
+ outsideFile := filepath.Join(t.TempDir(), "secret.key")
+ if err := os.WriteFile(outsideFile, []byte("stolen"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ r := credential.NewResolver(filepath.Dir(cfgPath))
+
+ cases := []string{
+ "file://../../secret.key",
+ "file://../secret.key",
+ "file://" + outsideFile, // absolute path
+ }
+ for _, raw := range cases {
+ _, err := r.Resolve(raw)
+ if err == nil {
+ t.Errorf("Resolve(%q): expected path traversal error, got nil", raw)
+ }
+ }
+}
+
+// TestResolve_FileRef_withinConfigDir verifies that a legitimate relative file:// ref works.
+func TestResolve_FileRef_withinConfigDir(t *testing.T) {
+ dir := t.TempDir()
+ if err := os.WriteFile(filepath.Join(dir, "my.key"), []byte("sk-valid\n"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+ r := credential.NewResolver(dir)
+ got, err := r.Resolve("file://my.key")
+ if err != nil {
+ t.Fatalf("unexpected error: %v", err)
+ }
+ if got != "sk-valid" {
+ t.Fatalf("got %q, want %q", got, "sk-valid")
+ }
+}
+
+// TestEncrypt_SSHKeyOutsideAllowedDirs verifies that Encrypt rejects SSH key paths
+// that are not under PICOCLAW_SSH_KEY_PATH, PICOCLAW_HOME, or ~/.ssh/.
+func TestEncrypt_SSHKeyOutsideAllowedDirs(t *testing.T) {
+ dir := t.TempDir()
+ sshKeyPath := filepath.Join(dir, "picoclaw_ed25519.key")
+ if err := os.WriteFile(sshKeyPath, []byte("fake-key\n"), 0o600); err != nil {
+ t.Fatalf("setup: %v", err)
+ }
+
+ // Make sure none of the allowed env vars point here.
+ t.Setenv("PICOCLAW_SSH_KEY_PATH", "")
+ t.Setenv("PICOCLAW_HOME", "")
+
+ _, err := credential.Encrypt("passphrase", sshKeyPath, "sk-secret")
+ if err == nil {
+ t.Fatal("expected error for SSH key outside allowed directories, got nil")
+ }
+}
diff --git a/pkg/credential/keygen.go b/pkg/credential/keygen.go
new file mode 100644
index 000000000..c57564a76
--- /dev/null
+++ b/pkg/credential/keygen.go
@@ -0,0 +1,62 @@
+package credential
+
+import (
+ "crypto/ed25519"
+ "crypto/rand"
+ "encoding/pem"
+ "fmt"
+ "os"
+ "path/filepath"
+
+ "golang.org/x/crypto/ssh"
+)
+
+// DefaultSSHKeyPath returns the canonical path for the picoclaw-specific SSH key.
+// The path is always ~/.ssh/picoclaw_ed25519.key (os.UserHomeDir is cross-platform).
+func DefaultSSHKeyPath() (string, error) {
+ home, err := os.UserHomeDir()
+ if err != nil {
+ return "", fmt.Errorf("credential: cannot determine home directory: %w", err)
+ }
+ return filepath.Join(home, ".ssh", "picoclaw_ed25519.key"), nil
+}
+
+// GenerateSSHKey generates an Ed25519 SSH key pair and writes the private key
+// to path (permissions 0600) and the public key to path+".pub" (permissions 0644).
+// The ~/.ssh/ directory is created with 0700 if it does not exist.
+// If the files already exist they are overwritten.
+func GenerateSSHKey(path string) error {
+ if err := os.MkdirAll(filepath.Dir(path), 0o700); err != nil {
+ return fmt.Errorf("credential: keygen: cannot create directory %q: %w", filepath.Dir(path), err)
+ }
+
+ pubRaw, privRaw, err := ed25519.GenerateKey(rand.Reader)
+ if err != nil {
+ return fmt.Errorf("credential: keygen: ed25519 key generation failed: %w", err)
+ }
+
+ // Marshal private key as OpenSSH PEM.
+ block, err := ssh.MarshalPrivateKey(privRaw, "")
+ if err != nil {
+ return fmt.Errorf("credential: keygen: marshal private key: %w", err)
+ }
+ privPEM := pem.EncodeToMemory(block)
+
+ if err = os.WriteFile(path, privPEM, 0o600); err != nil {
+ return fmt.Errorf("credential: keygen: write private key %q: %w", path, err)
+ }
+
+ // Marshal public key as authorized_keys line.
+ sshPub, err := ssh.NewPublicKey(pubRaw)
+ if err != nil {
+ return fmt.Errorf("credential: keygen: marshal public key: %w", err)
+ }
+ pubLine := ssh.MarshalAuthorizedKey(sshPub)
+
+ pubPath := path + ".pub"
+ if err := os.WriteFile(pubPath, pubLine, 0o644); err != nil {
+ return fmt.Errorf("credential: keygen: write public key %q: %w", pubPath, err)
+ }
+
+ return nil
+}
diff --git a/pkg/credential/keygen_test.go b/pkg/credential/keygen_test.go
new file mode 100644
index 000000000..1e21ea0b9
--- /dev/null
+++ b/pkg/credential/keygen_test.go
@@ -0,0 +1,115 @@
+package credential
+
+import (
+ "crypto/ed25519"
+ "os"
+ "path/filepath"
+ "runtime"
+ "testing"
+
+ "golang.org/x/crypto/ssh"
+)
+
+func TestGenerateSSHKey_CreatesFiles(t *testing.T) {
+ dir := t.TempDir()
+ keyPath := filepath.Join(dir, "test_ed25519.key")
+
+ if err := GenerateSSHKey(keyPath); err != nil {
+ t.Fatalf("GenerateSSHKey() error = %v", err)
+ }
+
+ // Private key must exist.
+ privInfo, err := os.Stat(keyPath)
+ if err != nil {
+ t.Fatalf("private key file missing: %v", err)
+ }
+
+ // Check permissions on non-Windows (Windows does not support Unix permission bits).
+ if runtime.GOOS != "windows" {
+ if got := privInfo.Mode().Perm(); got != 0o600 {
+ t.Errorf("private key permissions = %04o, want 0600", got)
+ }
+ }
+
+ // Public key must exist.
+ pubPath := keyPath + ".pub"
+ pubInfo, err := os.Stat(pubPath)
+ if err != nil {
+ t.Fatalf("public key file missing: %v", err)
+ }
+ if runtime.GOOS != "windows" {
+ if got := pubInfo.Mode().Perm(); got != 0o644 {
+ t.Errorf("public key permissions = %04o, want 0644", got)
+ }
+ }
+
+ // Private key must be parseable as an OpenSSH ed25519 key.
+ privPEM, err := os.ReadFile(keyPath)
+ if err != nil {
+ t.Fatalf("read private key: %v", err)
+ }
+ privKey, err := ssh.ParseRawPrivateKey(privPEM)
+ if err != nil {
+ t.Fatalf("parse private key: %v", err)
+ }
+ if _, ok := privKey.(*ed25519.PrivateKey); !ok {
+ t.Errorf("private key type = %T, want *ed25519.PrivateKey", privKey)
+ }
+
+ // Public key must be parseable as authorized_keys line.
+ pubBytes, err := os.ReadFile(pubPath)
+ if err != nil {
+ t.Fatalf("read public key: %v", err)
+ }
+ pubKey, _, _, rest, err := ssh.ParseAuthorizedKey(pubBytes)
+ if err != nil {
+ t.Fatalf("parse public key: %v", err)
+ }
+ if pubKey == nil {
+ t.Fatal("expected non-nil public key")
+ }
+ if len(rest) > 0 {
+ t.Errorf("unexpected trailing bytes after public key: %d bytes", len(rest))
+ }
+}
+
+func TestGenerateSSHKey_OverwritesExisting(t *testing.T) {
+ dir := t.TempDir()
+ keyPath := filepath.Join(dir, "test_ed25519.key")
+
+ // Generate twice; second call must not error and must produce a different key.
+ if err := GenerateSSHKey(keyPath); err != nil {
+ t.Fatalf("first GenerateSSHKey() error = %v", err)
+ }
+ first, err := os.ReadFile(keyPath)
+ if err != nil {
+ t.Fatalf("read first key: %v", err)
+ }
+
+ if err = GenerateSSHKey(keyPath); err != nil {
+ t.Fatalf("second GenerateSSHKey() error = %v", err)
+ }
+ second, err := os.ReadFile(keyPath)
+ if err != nil {
+ t.Fatalf("read second key: %v", err)
+ }
+
+ // Two independently generated Ed25519 keys must differ.
+ if string(first) == string(second) {
+ t.Error("expected overwritten key to differ from original")
+ }
+}
+
+func TestGenerateSSHKey_CreatesDirectory(t *testing.T) {
+ dir := t.TempDir()
+ // Nested directory that does not yet exist.
+ keyPath := filepath.Join(dir, "subdir", ".ssh", "picoclaw_ed25519.key")
+
+ if err := GenerateSSHKey(keyPath); err != nil {
+ t.Fatalf("GenerateSSHKey() error = %v", err)
+ }
+
+ if _, err := os.Stat(keyPath); err != nil {
+ t.Fatalf("private key not created: %v", err)
+ }
+}
diff --git a/pkg/credential/store.go b/pkg/credential/store.go
new file mode 100644
index 000000000..9c72974b0
--- /dev/null
+++ b/pkg/credential/store.go
@@ -0,0 +1,44 @@
+package credential
+
+import "sync/atomic"
+
+// SecureStore holds a passphrase in memory.
+//
+// Uses atomic.Pointer so reads and writes are lock-free.
+// The passphrase is never written to disk; callers decide how to
+// transport it outside this store (e.g., via cmd.Env or os.Environ).
+type SecureStore struct {
+ val atomic.Pointer[string]
+}
+
+// NewSecureStore creates an empty SecureStore.
+func NewSecureStore() *SecureStore {
+ return &SecureStore{}
+}
+
+// SetString stores the passphrase. An empty string clears the store.
+func (s *SecureStore) SetString(passphrase string) {
+ if passphrase == "" {
+ s.val.Store(nil)
+ return
+ }
+ s.val.Store(&passphrase)
+}
+
+// Get returns the stored passphrase, or "" if not set.
+func (s *SecureStore) Get() string {
+ if p := s.val.Load(); p != nil {
+ return *p
+ }
+ return ""
+}
+
+// IsSet reports whether a passphrase is currently stored.
+func (s *SecureStore) IsSet() bool {
+ return s.val.Load() != nil
+}
+
+// Clear removes the stored passphrase.
+func (s *SecureStore) Clear() {
+ s.val.Store(nil)
+}
diff --git a/pkg/credential/store_test.go b/pkg/credential/store_test.go
new file mode 100644
index 000000000..63299743a
--- /dev/null
+++ b/pkg/credential/store_test.go
@@ -0,0 +1,81 @@
+package credential
+
+import (
+ "sync"
+ "testing"
+)
+
+func TestSecureStore_SetGet(t *testing.T) {
+ s := NewSecureStore()
+ if s.IsSet() {
+ t.Error("expected empty store")
+ }
+
+ s.SetString("hunter2")
+ if !s.IsSet() {
+ t.Error("expected store to be set")
+ }
+ if got := s.Get(); got != "hunter2" {
+ t.Errorf("Get() = %q, want %q", got, "hunter2")
+ }
+}
+
+func TestSecureStore_Clear(t *testing.T) {
+ s := NewSecureStore()
+ s.SetString("secret")
+ s.Clear()
+
+ if s.IsSet() {
+ t.Error("expected store to be empty after Clear()")
+ }
+ if got := s.Get(); got != "" {
+ t.Errorf("Get() after Clear() = %q, want empty", got)
+ }
+}
+
+func TestSecureStore_SetOverwrites(t *testing.T) {
+ s := NewSecureStore()
+ s.SetString("first")
+ s.SetString("second")
+
+ if got := s.Get(); got != "second" {
+ t.Errorf("Get() = %q, want %q", got, "second")
+ }
+}
+
+func TestSecureStore_EmptyPassphrase(t *testing.T) {
+ s := NewSecureStore()
+ s.SetString("") // empty → should not mark as set
+
+ if s.IsSet() {
+ t.Error("empty passphrase should not mark store as set")
+ }
+}
+
+func TestSecureStore_ConcurrentSetGet(t *testing.T) {
+ s := NewSecureStore()
+ const goroutines = 10
+ const iterations = 1000
+
+ var wg sync.WaitGroup
+ wg.Add(goroutines)
+ for i := 0; i < goroutines; i++ {
+ go func(id int) {
+ defer wg.Done()
+ for j := 0; j < iterations; j++ {
+ if id%2 == 0 {
+ s.SetString("even")
+ } else {
+ s.SetString("odd")
+ }
+ _ = s.Get()
+ }
+ }(i)
+ }
+ wg.Wait()
+
+ final := s.Get()
+ if final != "" && final != "even" && final != "odd" {
+ t.Errorf("Get() returned unexpected value %q after concurrent Set/Get", final)
+ }
+}
diff --git a/pkg/cron/service.go b/pkg/cron/service.go
index 04775ac42..77a413133 100644
--- a/pkg/cron/service.go
+++ b/pkg/cron/service.go
@@ -65,6 +65,7 @@ type CronService struct {
mu sync.RWMutex
running bool
stopChan chan struct{}
+ wakeChan chan struct{}
gronx *gronx.Gronx
}
@@ -73,6 +74,7 @@ func NewCronService(storePath string, onJob JobHandler) *CronService {
storePath: storePath,
onJob: onJob,
gronx: gronx.New(),
+ wakeChan: make(chan struct{}),
}
// Initialize and load store on creation
cs.loadStore()
@@ -97,6 +99,9 @@ func (cs *CronService) Start() error {
}
cs.stopChan = make(chan struct{})
+ if cs.wakeChan == nil {
+ cs.wakeChan = make(chan struct{})
+ }
cs.running = true
go cs.runLoop(cs.stopChan)
@@ -119,14 +124,47 @@ func (cs *CronService) Stop() {
}
func (cs *CronService) runLoop(stopChan chan struct{}) {
- ticker := time.NewTicker(1 * time.Second)
- defer ticker.Stop()
+ timer := time.NewTimer(time.Hour)
+ if !timer.Stop() {
+ <-timer.C
+ }
+ defer timer.Stop()
for {
+ // every loop, recalculate the next wake time
+ cs.mu.RLock()
+ nextWake := cs.getNextWakeMS()
+ cs.mu.RUnlock()
+
+ var delay time.Duration
+ now := time.Now().UnixMilli()
+
+ if nextWake == nil {
+ // no jobs, sleep for a long time (or until a new job is added)
+ delay = time.Hour
+ } else {
+ diff := *nextWake - now
+ if diff <= 0 {
+ delay = 0
+ } else {
+ delay = time.Duration(diff) * time.Millisecond
+ }
+ }
+
+ timer.Reset(delay)
+
select {
case <-stopChan:
return
- case <-ticker.C:
+ case <-cs.wakeChan: // wake on new job or update
+ if !timer.Stop() {
+ select {
+ case <-timer.C:
+ default:
+ }
+ }
+ continue
+ case <-timer.C:
cs.checkJobs()
}
}
@@ -264,22 +302,19 @@ func (cs *CronService) executeJobByID(jobID string) {
}
func (cs *CronService) computeNextRun(schedule *CronSchedule, nowMS int64) *int64 {
- if schedule.Kind == "at" {
+ switch schedule.Kind {
+ case "at":
if schedule.AtMS != nil && *schedule.AtMS > nowMS {
return schedule.AtMS
}
return nil
- }
-
- if schedule.Kind == "every" {
+ case "every":
if schedule.EveryMS == nil || *schedule.EveryMS <= 0 {
return nil
}
next := nowMS + *schedule.EveryMS
return &next
- }
-
- if schedule.Kind == "cron" {
+ case "cron":
if schedule.Expr == "" {
return nil
}
@@ -294,9 +329,19 @@ func (cs *CronService) computeNextRun(schedule *CronSchedule, nowMS int64) *int6
nextMS := nextTime.UnixMilli()
return &nextMS
+ default:
+ log.Printf("[cron] unknown schedule kind '%s'", schedule.Kind)
+ return nil
}
+}
- return nil
+// wake up the loop to re-evaluate next wake time immediately (e.g. after add/update/remove jobs)
+func (cs *CronService) notify() {
+ select {
+ case cs.wakeChan <- struct{}{}:
+ default:
+ // if the channel is full, it means the loop will wake up soon anyway, so we can skip sending
+ }
}
func (cs *CronService) recomputeNextRuns() {
@@ -400,6 +445,8 @@ func (cs *CronService) AddJob(
return nil, err
}
+ cs.notify()
+
return &job, nil
}
@@ -411,6 +458,9 @@ func (cs *CronService) UpdateJob(job *CronJob) error {
if cs.store.Jobs[i].ID == job.ID {
cs.store.Jobs[i] = *job
cs.store.Jobs[i].UpdatedAtMS = time.Now().UnixMilli()
+
+ cs.notify()
+
return cs.saveStoreUnsafe()
}
}
@@ -441,6 +491,8 @@ func (cs *CronService) removeJobUnsafe(jobID string) bool {
}
}
+ cs.notify()
+
return removed
}
@@ -463,6 +515,9 @@ func (cs *CronService) EnableJob(jobID string, enabled bool) *CronJob {
if err := cs.saveStoreUnsafe(); err != nil {
log.Printf("[cron] failed to save store after enable: %v", err)
}
+
+ cs.notify()
+
return job
}
}
diff --git a/pkg/cron/service_test.go b/pkg/cron/service_test.go
index 1a0dd1829..c55e62174 100644
--- a/pkg/cron/service_test.go
+++ b/pkg/cron/service_test.go
@@ -1,10 +1,13 @@
package cron
import (
+ "fmt"
"os"
"path/filepath"
"runtime"
+ "sync"
"testing"
+ "time"
)
func TestSaveStore_FilePermissions(t *testing.T) {
@@ -36,3 +39,199 @@ func TestSaveStore_FilePermissions(t *testing.T) {
func int64Ptr(v int64) *int64 {
return &v
}
+
+func setupService(handler JobHandler) (*CronService, string) {
+ tmpFile := fmt.Sprintf("test_cron_%d.json", time.Now().UnixNano())
+ cs := NewCronService(tmpFile, handler)
+ return cs, tmpFile
+}
+
+func TestCronService_CRUD(t *testing.T) {
+ cs, path := setupService(nil)
+ defer os.Remove(path)
+
+ // Test AddJob
+ at := time.Now().Add(time.Hour).UnixMilli()
+ job, err := cs.AddJob("Task1", CronSchedule{Kind: "at", AtMS: &at}, "msg", true, "ch", "to")
+ if err != nil || job.ID == "" {
+ t.Fatalf("AddJob failed: %v", err)
+ }
+
+ // Test ListJobs
+ if len(cs.ListJobs(true)) != 1 {
+ t.Error("ListJobs should return 1 job")
+ }
+
+ // Test UpdateJob
+ job.Name = "UpdatedName"
+ err = cs.UpdateJob(job)
+ if err != nil || cs.store.Jobs[0].Name != "UpdatedName" {
+ t.Error("UpdateJob failed")
+ }
+
+ // Test EnableJob
+ cs.EnableJob(job.ID, false)
+ if cs.store.Jobs[0].Enabled != false || cs.store.Jobs[0].State.NextRunAtMS != nil {
+ t.Error("EnableJob(false) failed to clear state")
+ }
+
+ // Test RemoveJob
+ removed := cs.RemoveJob(job.ID)
+ if !removed || len(cs.store.Jobs) != 0 {
+ t.Error("RemoveJob failed")
+ }
+}
+
+// 2. Test Cron Expression Calculation Logic
+func TestCronService_ComputeNextRun(t *testing.T) {
+ cs, path := setupService(nil)
+ defer os.Remove(path)
+
+ now := time.Date(2024, 1, 1, 12, 0, 0, 0, time.UTC).UnixMilli()
+
+ tests := []struct {
+ name string
+ schedule CronSchedule
+ wantNil bool
+ }{
+ {"Valid Cron", CronSchedule{Kind: "cron", Expr: "0 * * * *"}, false},
+ {"Invalid Cron", CronSchedule{Kind: "cron", Expr: "invalid"}, true},
+ {"Every MS", CronSchedule{Kind: "every", EveryMS: int64Ptr(5000)}, false},
+ {"At Future", CronSchedule{Kind: "at", AtMS: int64Ptr(now + 1000)}, false},
+ {"At Past", CronSchedule{Kind: "at", AtMS: int64Ptr(now - 1000)}, true},
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ got := cs.computeNextRun(&tt.schedule, now)
+ if (got == nil) != tt.wantNil {
+ t.Errorf("%s: got %v, wantNil %v", tt.name, got, tt.wantNil)
+ }
+ })
+ }
+}
+
+// 3. Test Execution Flow
+func TestCronService_ExecutionFlow(t *testing.T) {
+ var mu sync.Mutex
+ executedJobs := make(map[string]bool)
+
+ handler := func(job *CronJob) (string, error) {
+ mu.Lock()
+ executedJobs[job.ID] = true
+ mu.Unlock()
+ return "ok", nil
+ }
+
+ cs, path := setupService(handler)
+ defer os.Remove(path)
+
+ // Start the service
+ if err := cs.Start(); err != nil {
+ t.Fatalf("Start failed: %v", err)
+ }
+ defer cs.Stop()
+
+ // Add a job then runs 100ms from now
+ target := time.Now().Add(100 * time.Millisecond).UnixMilli()
+ job, _ := cs.AddJob("FastJob", CronSchedule{Kind: "at", AtMS: &target}, "", false, "", "")
+
+ // Check for job execution with a timeout
+ success := false
+ for range 20 {
+ mu.Lock()
+ if executedJobs[job.ID] {
+ success = true
+ mu.Unlock()
+ break
+ }
+ mu.Unlock()
+ time.Sleep(100 * time.Millisecond)
+ }
+
+ if !success {
+ t.Error("Job was not executed in time")
+ }
+
+ // check that the job is removed after execution (DeleteAfterRun = true)
+ status := cs.Status()
+ if status["jobs"].(int) != 0 {
+ t.Errorf("Job should be deleted after run, got count: %v", status["jobs"])
+ }
+}
+
+func TestCronService_PersistenceIntegrity(t *testing.T) {
+ tmpFile := "persist_test.json"
+ defer os.Remove(tmpFile)
+
+ // write a job and persist
+ cs1 := NewCronService(tmpFile, nil)
+ at := int64(2000000000000)
+ cs1.AddJob("PersistMe", CronSchedule{Kind: "at", AtMS: &at}, "payload", true, "ch1", "")
+
+ // check file exists
+ if _, err := os.Stat(tmpFile); os.IsNotExist(err) {
+ t.Fatal("Store file was not created")
+ }
+
+ // reload and check data integrity
+ cs2 := NewCronService(tmpFile, nil)
+ if err := cs2.Load(); err != nil {
+ t.Fatalf("Failed to load store: %v", err)
+ }
+
+ jobs := cs2.ListJobs(true)
+ if len(jobs) != 1 || jobs[0].Name != "PersistMe" {
+ t.Errorf("Data corruption after reload. Got: %+v", jobs)
+ }
+
+ // test loading invalid JSON
+ os.WriteFile(tmpFile, []byte("{invalid json}"), 0o644)
+ cs3 := NewCronService(tmpFile, nil)
+ err := cs3.loadStore()
+ if err == nil {
+ t.Error("Should return error when loading invalid JSON")
+ }
+}
+
+func TestCronService_ConcurrentAccess(t *testing.T) {
+ cs, path := setupService(nil)
+ defer os.Remove(path)
+
+ cs.Start()
+ defer cs.Stop()
+
+ var wg sync.WaitGroup
+ workers := 10
+ iterations := 50
+
+ wg.Add(workers * 2)
+
+ // add jobs concurrently
+ for i := range workers {
+ go func(id int) {
+ defer wg.Done()
+ for j := range iterations {
+ at := time.Now().Add(time.Hour).UnixMilli()
+ cs.AddJob(fmt.Sprintf("Job-%d-%d", id, j), CronSchedule{Kind: "at", AtMS: &at}, "", false, "", "")
+ time.Sleep(100 * time.Microsecond)
+ }
+ }(i)
+ }
+
+ // read and update jobs concurrently
+ for range workers {
+ go func() {
+ defer wg.Done()
+ for j := range iterations {
+ jobs := cs.ListJobs(true)
+ if len(jobs) > 0 {
+ cs.EnableJob(jobs[0].ID, j%2 == 0)
+ }
+ time.Sleep(100 * time.Microsecond)
+ }
+ }()
+ }
+
+ wg.Wait()
+}
diff --git a/cmd/picoclaw/internal/gateway/helpers.go b/pkg/gateway/gateway.go
similarity index 56%
rename from cmd/picoclaw/internal/gateway/helpers.go
rename to pkg/gateway/gateway.go
index 3562f03ef..4ad4e950e 100644
--- a/cmd/picoclaw/internal/gateway/helpers.go
+++ b/pkg/gateway/gateway.go
@@ -7,9 +7,10 @@ import (
"os/signal"
"path/filepath"
"sync"
+ "sync/atomic"
+ "syscall"
"time"
- "github.com/sipeed/picoclaw/cmd/picoclaw/internal"
"github.com/sipeed/picoclaw/pkg/agent"
"github.com/sipeed/picoclaw/pkg/bus"
"github.com/sipeed/picoclaw/pkg/channels"
@@ -41,42 +42,60 @@ import (
"github.com/sipeed/picoclaw/pkg/voice"
)
-// Timeout constants for service operations
const (
- serviceRestartTimeout = 30 * time.Second
serviceShutdownTimeout = 30 * time.Second
providerReloadTimeout = 30 * time.Second
gracefulShutdownTimeout = 15 * time.Second
)
-// gatewayServices holds references to all running services
-type gatewayServices struct {
+type services struct {
CronService *cron.CronService
HeartbeatService *heartbeat.HeartbeatService
MediaStore media.MediaStore
ChannelManager *channels.Manager
DeviceService *devices.Service
HealthServer *health.Server
+ manualReloadChan chan struct{}
+ reloading atomic.Bool
}
-func gatewayCmd(debug bool) error {
+type startupBlockedProvider struct {
+ reason string
+}
+
+func (p *startupBlockedProvider) Chat(
+ _ context.Context,
+ _ []providers.Message,
+ _ []providers.ToolDefinition,
+ _ string,
+ _ map[string]any,
+) (*providers.LLMResponse, error) {
+ return nil, fmt.Errorf("%s", p.reason)
+}
+
+func (p *startupBlockedProvider) GetDefaultModel() string {
+ return ""
+}
+
+// Run starts the gateway runtime using the configuration loaded from configPath.
+func Run(debug bool, configPath string, allowEmptyStartup bool) error {
+ cfg, err := config.LoadConfig(configPath)
+ if err != nil {
+ return fmt.Errorf("error loading config: %w", err)
+ }
+
+ logger.SetLevelFromString(cfg.Agents.Defaults.LogLevel)
+
if debug {
logger.SetLevel(logger.DEBUG)
fmt.Println("🔍 Debug mode enabled")
}
- configPath := internal.GetConfigPath()
- cfg, err := internal.LoadConfig()
- if err != nil {
- return fmt.Errorf("error loading config: %w", err)
- }
-
- provider, modelID, err := providers.CreateProvider(cfg)
+ provider, modelID, err := createStartupProvider(cfg, allowEmptyStartup)
if err != nil {
return fmt.Errorf("error creating provider: %w", err)
}
- // Use the resolved model ID from provider creation
if modelID != "" {
cfg.Agents.Defaults.ModelName = modelID
}
@@ -84,17 +103,13 @@ func gatewayCmd(debug bool) error {
msgBus := bus.NewMessageBus()
agentLoop := agent.NewAgentLoop(cfg, msgBus, provider)
- // Print agent startup info
fmt.Println("\n📦 Agent Status:")
startupInfo := agentLoop.GetStartupInfo()
toolsInfo := startupInfo["tools"].(map[string]any)
skillsInfo := startupInfo["skills"].(map[string]any)
fmt.Printf(" • Tools: %d loaded\n", toolsInfo["count"])
- fmt.Printf(" • Skills: %d/%d available\n",
- skillsInfo["available"],
- skillsInfo["total"])
+ fmt.Printf(" • Skills: %d/%d available\n", skillsInfo["available"], skillsInfo["total"])
- // Log to file as well
logger.InfoCF("agent", "Agent initialized",
map[string]any{
"tools_count": toolsInfo["count"],
@@ -102,12 +117,30 @@ func gatewayCmd(debug bool) error {
"skills_available": skillsInfo["available"],
})
- // Setup and start all services
- services, err := setupAndStartServices(cfg, agentLoop, msgBus)
+ runningServices, err := setupAndStartServices(cfg, agentLoop, msgBus)
if err != nil {
return err
}
+ // Setup manual reload channel for /reload endpoint
+ manualReloadChan := make(chan struct{}, 1)
+ runningServices.manualReloadChan = manualReloadChan
+ reloadTrigger := func() error {
+ if !runningServices.reloading.CompareAndSwap(false, true) {
+ return fmt.Errorf("reload already in progress")
+ }
+ select {
+ case manualReloadChan <- struct{}{}:
+ return nil
+ default:
+ // Should not happen, but reset flag if channel is full
+ runningServices.reloading.Store(false)
+ return fmt.Errorf("reload already queued")
+ }
+ }
+ runningServices.HealthServer.SetReloadFunc(reloadTrigger)
+ agentLoop.SetReloadFunc(reloadTrigger)
+
fmt.Printf("✓ Gateway started on %s:%d\n", cfg.Gateway.Host, cfg.Gateway.Port)
fmt.Println("Press Ctrl+C to stop")
@@ -116,41 +149,95 @@ func gatewayCmd(debug bool) error {
go agentLoop.Run(ctx)
- // Setup config file watcher for hot reload
- configReloadChan, stopWatch := setupConfigWatcherPolling(configPath, debug)
+ var configReloadChan <-chan *config.Config
+ stopWatch := func() {}
+ if cfg.Gateway.HotReload {
+ configReloadChan, stopWatch = setupConfigWatcherPolling(configPath, debug)
+ logger.Info("Config hot reload enabled")
+ }
defer stopWatch()
sigChan := make(chan os.Signal, 1)
- signal.Notify(sigChan, os.Interrupt)
+ signal.Notify(sigChan, os.Interrupt, syscall.SIGTERM)
- // Main event loop - wait for signals or config changes
for {
select {
case <-sigChan:
logger.Info("Shutting down...")
- shutdownGateway(services, agentLoop, provider, true)
+ shutdownGateway(runningServices, agentLoop, provider, true)
return nil
-
case newCfg := <-configReloadChan:
- err := handleConfigReload(ctx, agentLoop, newCfg, &provider, services, msgBus)
+ if !runningServices.reloading.CompareAndSwap(false, true) {
+ logger.Warn("Config reload skipped: another reload is in progress")
+ continue
+ }
+ err := executeReload(ctx, agentLoop, newCfg, &provider, runningServices, msgBus, allowEmptyStartup)
if err != nil {
logger.Errorf("Config reload failed: %v", err)
}
+ case <-manualReloadChan:
+ logger.Info("Manual reload triggered via /reload endpoint")
+ newCfg, err := config.LoadConfig(configPath)
+ if err != nil {
+ logger.Errorf("Error loading config for manual reload: %v", err)
+ runningServices.reloading.Store(false)
+ continue
+ }
+ if err = newCfg.ValidateModelList(); err != nil {
+ logger.Errorf("Config validation failed: %v", err)
+ runningServices.reloading.Store(false)
+ continue
+ }
+ err = executeReload(ctx, agentLoop, newCfg, &provider, runningServices, msgBus, allowEmptyStartup)
+ if err != nil {
+ logger.Errorf("Manual reload failed: %v", err)
+ } else {
+ logger.Info("Manual reload completed successfully")
+ }
}
}
}
-// setupAndStartServices initializes and starts all services
+func executeReload(
+ ctx context.Context,
+ agentLoop *agent.AgentLoop,
+ newCfg *config.Config,
+ provider *providers.LLMProvider,
+ runningServices *services,
+ msgBus *bus.MessageBus,
+ allowEmptyStartup bool,
+) error {
+ defer runningServices.reloading.Store(false)
+ return handleConfigReload(ctx, agentLoop, newCfg, provider, runningServices, msgBus, allowEmptyStartup)
+}
+
+func createStartupProvider(
+ cfg *config.Config,
+ allowEmptyStartup bool,
+) (providers.LLMProvider, string, error) {
+ modelName := cfg.Agents.Defaults.GetModelName()
+ if modelName == "" && allowEmptyStartup {
+ reason := "no default model configured; gateway started in limited mode"
+ fmt.Printf("⚠ Warning: %s\n", reason)
+ logger.WarnCF("gateway", "Gateway started without default model", map[string]any{
+ "limited_mode": true,
+ })
+ return &startupBlockedProvider{reason: reason}, "", nil
+ }
+
+ return providers.CreateProvider(cfg)
+}
+
func setupAndStartServices(
cfg *config.Config,
agentLoop *agent.AgentLoop,
msgBus *bus.MessageBus,
-) (*gatewayServices, error) {
- services := &gatewayServices{}
+) (*services, error) {
+ runningServices := &services{}
- // Setup cron tool and service
execTimeout := time.Duration(cfg.Tools.Cron.ExecTimeoutMinutes) * time.Minute
- services.CronService = setupCronTool(
+ var err error
+ runningServices.CronService, err = setupCronTool(
agentLoop,
msgBus,
cfg.WorkspacePath(),
@@ -158,139 +245,113 @@ func setupAndStartServices(
execTimeout,
cfg,
)
- if err := services.CronService.Start(); err != nil {
+ if err != nil {
+ return nil, fmt.Errorf("error setting up cron service: %w", err)
+ }
+ if err = runningServices.CronService.Start(); err != nil {
return nil, fmt.Errorf("error starting cron service: %w", err)
}
fmt.Println("✓ Cron service started")
- // Setup heartbeat service
- services.HeartbeatService = heartbeat.NewHeartbeatService(
+ runningServices.HeartbeatService = heartbeat.NewHeartbeatService(
cfg.WorkspacePath(),
cfg.Heartbeat.Interval,
cfg.Heartbeat.Enabled,
)
- services.HeartbeatService.SetBus(msgBus)
- services.HeartbeatService.SetHandler(func(prompt, channel, chatID string) *tools.ToolResult {
- // Use cli:direct as fallback if no valid channel
- if channel == "" || chatID == "" {
- channel, chatID = "cli", "direct"
- }
- // Use ProcessHeartbeat - no session history, each heartbeat is independent
- var response string
- var err error
- response, err = agentLoop.ProcessHeartbeat(context.Background(), prompt, channel, chatID)
- if err != nil {
- return tools.ErrorResult(fmt.Sprintf("Heartbeat error: %v", err))
- }
- if response == "HEARTBEAT_OK" {
- return tools.SilentResult("Heartbeat OK")
- }
- // For heartbeat, always return silent - the subagent result will be
- // sent to user via processSystemMessage when the async task completes
- return tools.SilentResult(response)
- })
- if err := services.HeartbeatService.Start(); err != nil {
+ runningServices.HeartbeatService.SetBus(msgBus)
+ runningServices.HeartbeatService.SetHandler(createHeartbeatHandler(agentLoop))
+ if err = runningServices.HeartbeatService.Start(); err != nil {
return nil, fmt.Errorf("error starting heartbeat service: %w", err)
}
fmt.Println("✓ Heartbeat service started")
- // Create media store for file lifecycle management with TTL cleanup
- services.MediaStore = media.NewFileMediaStoreWithCleanup(media.MediaCleanerConfig{
+ runningServices.MediaStore = media.NewFileMediaStoreWithCleanup(media.MediaCleanerConfig{
Enabled: cfg.Tools.MediaCleanup.Enabled,
MaxAge: time.Duration(cfg.Tools.MediaCleanup.MaxAge) * time.Minute,
Interval: time.Duration(cfg.Tools.MediaCleanup.Interval) * time.Minute,
})
- // Start the media store if it's a FileMediaStore with cleanup
- if fms, ok := services.MediaStore.(*media.FileMediaStore); ok {
+ if fms, ok := runningServices.MediaStore.(*media.FileMediaStore); ok {
fms.Start()
}
- // Create channel manager
- var err error
- services.ChannelManager, err = channels.NewManager(cfg, msgBus, services.MediaStore)
+ runningServices.ChannelManager, err = channels.NewManager(cfg, msgBus, runningServices.MediaStore)
if err != nil {
- // Stop the media store if it's a FileMediaStore with cleanup
- if fms, ok := services.MediaStore.(*media.FileMediaStore); ok {
+ if fms, ok := runningServices.MediaStore.(*media.FileMediaStore); ok {
fms.Stop()
}
return nil, fmt.Errorf("error creating channel manager: %w", err)
}
- // Inject channel manager and media store into agent loop
- agentLoop.SetChannelManager(services.ChannelManager)
- agentLoop.SetMediaStore(services.MediaStore)
+ agentLoop.SetChannelManager(runningServices.ChannelManager)
+ agentLoop.SetMediaStore(runningServices.MediaStore)
- // Wire up voice transcription if a supported provider is configured.
if transcriber := voice.DetectTranscriber(cfg); transcriber != nil {
agentLoop.SetTranscriber(transcriber)
logger.InfoCF("voice", "Transcription enabled (agent-level)", map[string]any{"provider": transcriber.Name()})
}
- enabledChannels := services.ChannelManager.GetEnabledChannels()
+ enabledChannels := runningServices.ChannelManager.GetEnabledChannels()
if len(enabledChannels) > 0 {
fmt.Printf("✓ Channels enabled: %s\n", enabledChannels)
} else {
fmt.Println("⚠ Warning: No channels enabled")
}
- // Setup shared HTTP server with health endpoints and webhook handlers
addr := fmt.Sprintf("%s:%d", cfg.Gateway.Host, cfg.Gateway.Port)
- services.HealthServer = health.NewServer(cfg.Gateway.Host, cfg.Gateway.Port)
- services.ChannelManager.SetupHTTPServer(addr, services.HealthServer)
+ runningServices.HealthServer = health.NewServer(cfg.Gateway.Host, cfg.Gateway.Port)
+ runningServices.ChannelManager.SetupHTTPServer(addr, runningServices.HealthServer)
- if err := services.ChannelManager.StartAll(context.Background()); err != nil {
+ if err = runningServices.ChannelManager.StartAll(context.Background()); err != nil {
return nil, fmt.Errorf("error starting channels: %w", err)
}
- fmt.Printf("✓ Health endpoints available at http://%s:%d/health and /ready\n", cfg.Gateway.Host, cfg.Gateway.Port)
+ fmt.Printf(
+ "✓ Health endpoints available at http://%s:%d/health, /ready and /reload (POST)\n",
+ cfg.Gateway.Host,
+ cfg.Gateway.Port,
+ )
- // Setup state manager and device service
stateManager := state.NewManager(cfg.WorkspacePath())
- services.DeviceService = devices.NewService(devices.Config{
+ runningServices.DeviceService = devices.NewService(devices.Config{
Enabled: cfg.Devices.Enabled,
MonitorUSB: cfg.Devices.MonitorUSB,
}, stateManager)
- services.DeviceService.SetBus(msgBus)
- if err := services.DeviceService.Start(context.Background()); err != nil {
+ runningServices.DeviceService.SetBus(msgBus)
+ if err = runningServices.DeviceService.Start(context.Background()); err != nil {
logger.ErrorCF("device", "Error starting device service", map[string]any{"error": err.Error()})
} else if cfg.Devices.Enabled {
fmt.Println("✓ Device event service started")
}
- return services, nil
+ return runningServices, nil
}
-// stopAndCleanupServices stops all services and cleans up resources
-func stopAndCleanupServices(
- services *gatewayServices,
- shutdownTimeout time.Duration,
-) {
+func stopAndCleanupServices(runningServices *services, shutdownTimeout time.Duration, isReload bool) {
shutdownCtx, shutdownCancel := context.WithTimeout(context.Background(), shutdownTimeout)
defer shutdownCancel()
- if services.ChannelManager != nil {
- services.ChannelManager.StopAll(shutdownCtx)
+ // reload should not stop channel manager
+ if !isReload && runningServices.ChannelManager != nil {
+ runningServices.ChannelManager.StopAll(shutdownCtx)
}
- if services.DeviceService != nil {
- services.DeviceService.Stop()
+ if runningServices.DeviceService != nil {
+ runningServices.DeviceService.Stop()
}
- if services.HeartbeatService != nil {
- services.HeartbeatService.Stop()
+ if runningServices.HeartbeatService != nil {
+ runningServices.HeartbeatService.Stop()
}
- if services.CronService != nil {
- services.CronService.Stop()
+ if runningServices.CronService != nil {
+ runningServices.CronService.Stop()
}
- if services.MediaStore != nil {
- // Stop the media store if it's a FileMediaStore with cleanup
- if fms, ok := services.MediaStore.(*media.FileMediaStore); ok {
+ if runningServices.MediaStore != nil {
+ if fms, ok := runningServices.MediaStore.(*media.FileMediaStore); ok {
fms.Stop()
}
}
}
-// shutdownGateway performs a complete gateway shutdown
func shutdownGateway(
- services *gatewayServices,
+ runningServices *services,
agentLoop *agent.AgentLoop,
provider providers.LLMProvider,
fullShutdown bool,
@@ -299,7 +360,7 @@ func shutdownGateway(
cp.Close()
}
- stopAndCleanupServices(services, gracefulShutdownTimeout)
+ stopAndCleanupServices(runningServices, gracefulShutdownTimeout, false)
agentLoop.Stop()
agentLoop.Close()
@@ -307,15 +368,14 @@ func shutdownGateway(
logger.Info("✓ Gateway stopped")
}
-// handleConfigReload handles config file reload by stopping all services,
-// reloading the provider and config, and restarting services with the new config.
func handleConfigReload(
ctx context.Context,
al *agent.AgentLoop,
newCfg *config.Config,
providerRef *providers.LLMProvider,
- services *gatewayServices,
+ runningServices *services,
msgBus *bus.MessageBus,
+ allowEmptyStartup bool,
) error {
logger.Info("🔄 Config file changed, reloading...")
@@ -326,18 +386,14 @@ func handleConfigReload(
logger.Infof(" New model is '%s', recreating provider...", newModel)
- // Stop all services before reloading
logger.Info(" Stopping all services...")
- stopAndCleanupServices(services, serviceShutdownTimeout)
+ stopAndCleanupServices(runningServices, serviceShutdownTimeout, true)
- // Create new provider from updated config first to ensure validity
- // This will use the correct API key and settings from newCfg.ModelList
- newProvider, newModelID, err := providers.CreateProvider(newCfg)
+ newProvider, newModelID, err := createStartupProvider(newCfg, allowEmptyStartup)
if err != nil {
logger.Errorf(" ⚠ Error creating new provider: %v", err)
logger.Warn(" Attempting to restart services with old provider and config...")
- // Try to restart services with old configuration
- if restartErr := restartServices(al, services, msgBus); restartErr != nil {
+ if restartErr := restartServices(al, runningServices, msgBus); restartErr != nil {
logger.Errorf(" ⚠ Failed to restart services: %v", restartErr)
}
return fmt.Errorf("error creating new provider: %w", err)
@@ -347,31 +403,25 @@ func handleConfigReload(
newCfg.Agents.Defaults.ModelName = newModelID
}
- // Use the atomic reload method on AgentLoop to safely swap provider and config.
- // This handles locking internally to prevent races with in-flight LLM calls
- // and concurrent reads of registry/config while the swap occurs.
reloadCtx, reloadCancel := context.WithTimeout(context.Background(), providerReloadTimeout)
defer reloadCancel()
if err := al.ReloadProviderAndConfig(reloadCtx, newProvider, newCfg); err != nil {
logger.Errorf(" ⚠ Error reloading agent loop: %v", err)
- // Close the newly created provider since it wasn't adopted
if cp, ok := newProvider.(providers.StatefulProvider); ok {
cp.Close()
}
logger.Warn(" Attempting to restart services with old provider and config...")
- if restartErr := restartServices(al, services, msgBus); restartErr != nil {
+ if restartErr := restartServices(al, runningServices, msgBus); restartErr != nil {
logger.Errorf(" ⚠ Failed to restart services: %v", restartErr)
}
return fmt.Errorf("error reloading agent loop: %w", err)
}
- // Update local provider reference only after successful atomic reload
*providerRef = newProvider
- // Restart all services with new config
logger.Info(" Restarting all services with new configuration...")
- if err := restartServices(al, services, msgBus); err != nil {
+ if err := restartServices(al, runningServices, msgBus); err != nil {
logger.Errorf(" ⚠ Error restarting services: %v", err)
return fmt.Errorf("error restarting services: %w", err)
}
@@ -380,23 +430,16 @@ func handleConfigReload(
return nil
}
-// restartServices restarts all services after a config reload
func restartServices(
al *agent.AgentLoop,
- services *gatewayServices,
+ runningServices *services,
msgBus *bus.MessageBus,
) error {
- // Create an independent context with timeout for service restart
- // This prevents cancellation from the main loop context during reload
- ctx, cancel := context.WithTimeout(context.Background(), serviceRestartTimeout)
- defer cancel()
-
- // Get current config from agent loop (which has been updated if this is a reload)
cfg := al.GetConfig()
- // Re-create and start cron service with new config
execTimeout := time.Duration(cfg.Tools.Cron.ExecTimeoutMinutes) * time.Minute
- services.CronService = setupCronTool(
+ var err error
+ runningServices.CronService, err = setupCronTool(
al,
msgBus,
cfg.WorkspacePath(),
@@ -404,104 +447,75 @@ func restartServices(
execTimeout,
cfg,
)
- if err := services.CronService.Start(); err != nil {
+ if err != nil {
+ return fmt.Errorf("error restarting cron service: %w", err)
+ }
+ if err = runningServices.CronService.Start(); err != nil {
return fmt.Errorf("error restarting cron service: %w", err)
}
fmt.Println(" ✓ Cron service restarted")
- // Re-create and start heartbeat service with new config
- services.HeartbeatService = heartbeat.NewHeartbeatService(
+ runningServices.HeartbeatService = heartbeat.NewHeartbeatService(
cfg.WorkspacePath(),
cfg.Heartbeat.Interval,
cfg.Heartbeat.Enabled,
)
- services.HeartbeatService.SetBus(msgBus)
- services.HeartbeatService.SetHandler(func(prompt, channel, chatID string) *tools.ToolResult {
- if channel == "" || chatID == "" {
- channel, chatID = "cli", "direct"
- }
- var response string
- var err error
- response, err = al.ProcessHeartbeat(context.Background(), prompt, channel, chatID)
- if err != nil {
- return tools.ErrorResult(fmt.Sprintf("Heartbeat error: %v", err))
- }
- if response == "HEARTBEAT_OK" {
- return tools.SilentResult("Heartbeat OK")
- }
- return tools.SilentResult(response)
- })
- if err := services.HeartbeatService.Start(); err != nil {
+ runningServices.HeartbeatService.SetBus(msgBus)
+ runningServices.HeartbeatService.SetHandler(createHeartbeatHandler(al))
+ if err = runningServices.HeartbeatService.Start(); err != nil {
return fmt.Errorf("error restarting heartbeat service: %w", err)
}
fmt.Println(" ✓ Heartbeat service restarted")
- // Stop the old media store before creating a new one
- if fms, ok := services.MediaStore.(*media.FileMediaStore); ok {
- fms.Stop()
- }
-
- // Re-create media store with new config
- services.MediaStore = media.NewFileMediaStoreWithCleanup(media.MediaCleanerConfig{
+ runningServices.MediaStore = media.NewFileMediaStoreWithCleanup(media.MediaCleanerConfig{
Enabled: cfg.Tools.MediaCleanup.Enabled,
MaxAge: time.Duration(cfg.Tools.MediaCleanup.MaxAge) * time.Minute,
Interval: time.Duration(cfg.Tools.MediaCleanup.Interval) * time.Minute,
})
- // Start the media store if it's a FileMediaStore with cleanup
- if fms, ok := services.MediaStore.(*media.FileMediaStore); ok {
+ if fms, ok := runningServices.MediaStore.(*media.FileMediaStore); ok {
fms.Start()
}
- al.SetMediaStore(services.MediaStore)
+ al.SetMediaStore(runningServices.MediaStore)
- // Re-create channel manager with new config
- var err error
- services.ChannelManager, err = channels.NewManager(cfg, msgBus, services.MediaStore)
+ runningServices.ChannelManager, err = channels.NewManager(cfg, msgBus, runningServices.MediaStore)
if err != nil {
- // Stop the media store if it's a FileMediaStore with cleanup
- if fms, ok := services.MediaStore.(*media.FileMediaStore); ok {
- fms.Stop()
- }
return fmt.Errorf("error recreating channel manager: %w", err)
}
- al.SetChannelManager(services.ChannelManager)
+ al.SetChannelManager(runningServices.ChannelManager)
- enabledChannels := services.ChannelManager.GetEnabledChannels()
+ enabledChannels := runningServices.ChannelManager.GetEnabledChannels()
if len(enabledChannels) > 0 {
fmt.Printf(" ✓ Channels enabled: %s\n", enabledChannels)
} else {
fmt.Println(" ⚠ Warning: No channels enabled")
}
- // Setup HTTP server with new config
addr := fmt.Sprintf("%s:%d", cfg.Gateway.Host, cfg.Gateway.Port)
- services.HealthServer = health.NewServer(cfg.Gateway.Host, cfg.Gateway.Port)
- services.ChannelManager.SetupHTTPServer(addr, services.HealthServer)
-
- if err := services.ChannelManager.StartAll(ctx); err != nil {
- return fmt.Errorf("error restarting channels: %w", err)
+ // Reuse existing HealthServer to preserve reloadFunc
+ if runningServices.HealthServer == nil {
+ runningServices.HealthServer = health.NewServer(cfg.Gateway.Host, cfg.Gateway.Port)
}
- fmt.Printf(
- " ✓ Channels restarted, health endpoints at http://%s:%d/health and ready\n",
- cfg.Gateway.Host,
- cfg.Gateway.Port,
- )
+ runningServices.ChannelManager.SetupHTTPServer(addr, runningServices.HealthServer)
+
+ if err = runningServices.ChannelManager.Reload(context.Background(), cfg); err != nil {
+ return fmt.Errorf("error reload channels: %w", err)
+ }
+ fmt.Println(" ✓ Channels restarted.")
- // Re-create device service with new config
stateManager := state.NewManager(cfg.WorkspacePath())
- services.DeviceService = devices.NewService(devices.Config{
+ runningServices.DeviceService = devices.NewService(devices.Config{
Enabled: cfg.Devices.Enabled,
MonitorUSB: cfg.Devices.MonitorUSB,
}, stateManager)
- services.DeviceService.SetBus(msgBus)
- if err := services.DeviceService.Start(ctx); err != nil {
+ runningServices.DeviceService.SetBus(msgBus)
+ if err := runningServices.DeviceService.Start(context.Background()); err != nil {
logger.WarnCF("device", "Failed to restart device service", map[string]any{"error": err.Error()})
} else if cfg.Devices.Enabled {
fmt.Println(" ✓ Device event service restarted")
}
- // Wire up voice transcription with new config
transcriber := voice.DetectTranscriber(cfg)
- al.SetTranscriber(transcriber) // This will set it to nil if disabled
+ al.SetTranscriber(transcriber)
if transcriber != nil {
logger.InfoCF("voice", "Transcription re-enabled (agent-level)", map[string]any{"provider": transcriber.Name()})
} else {
@@ -511,8 +525,6 @@ func restartServices(
return nil
}
-// setupConfigWatcherPolling sets up a simple polling-based config file watcher
-// Returns a channel for config updates and a stop function
func setupConfigWatcherPolling(configPath string, debug bool) (chan *config.Config, func()) {
configChan := make(chan *config.Config, 1)
stop := make(chan struct{})
@@ -522,11 +534,10 @@ func setupConfigWatcherPolling(configPath string, debug bool) (chan *config.Conf
go func() {
defer wg.Done()
- // Get initial file info
lastModTime := getFileModTime(configPath)
lastSize := getFileSize(configPath)
- ticker := time.NewTicker(2 * time.Second) // Check every 2 seconds
+ ticker := time.NewTicker(2 * time.Second)
defer ticker.Stop()
for {
@@ -535,16 +546,16 @@ func setupConfigWatcherPolling(configPath string, debug bool) (chan *config.Conf
currentModTime := getFileModTime(configPath)
currentSize := getFileSize(configPath)
- // Check if file changed (modification time or size changed)
if currentModTime.After(lastModTime) || currentSize != lastSize {
if debug {
logger.Debugf("🔍 Config file change detected")
}
- // Debounce - wait a bit to ensure file write is complete
time.Sleep(500 * time.Millisecond)
- // Validate and load new config
+ lastModTime = currentModTime
+ lastSize = currentSize
+
newCfg, err := config.LoadConfig(configPath)
if err != nil {
logger.Errorf("⚠ Error loading new config: %v", err)
@@ -552,7 +563,6 @@ func setupConfigWatcherPolling(configPath string, debug bool) (chan *config.Conf
continue
}
- // Validate the new config
if err := newCfg.ValidateModelList(); err != nil {
logger.Errorf(" ⚠ New config validation failed: %v", err)
logger.Warn(" Using previous valid config")
@@ -561,19 +571,12 @@ func setupConfigWatcherPolling(configPath string, debug bool) (chan *config.Conf
logger.Info("✓ Config file validated and loaded")
- // Update last known state
- lastModTime = currentModTime
- lastSize = currentSize
-
- // Send new config to main loop (non-blocking)
select {
case configChan <- newCfg:
default:
- // Channel full, skip this update
logger.Warn("⚠ Previous config reload still in progress, skipping")
}
}
-
case <-stop:
return
}
@@ -588,7 +591,6 @@ func setupConfigWatcherPolling(configPath string, debug bool) (chan *config.Conf
return configChan, stopFunc
}
-// getFileModTime returns the modification time of a file, or zero time if file doesn't exist
func getFileModTime(path string) time.Time {
info, err := os.Stat(path)
if err != nil {
@@ -597,7 +599,6 @@ func getFileModTime(path string) time.Time {
return info.ModTime()
}
-// getFileSize returns the size of a file, or 0 if file doesn't exist
func getFileSize(path string) int64 {
info, err := os.Stat(path)
if err != nil {
@@ -613,25 +614,22 @@ func setupCronTool(
restrict bool,
execTimeout time.Duration,
cfg *config.Config,
-) *cron.CronService {
+) (*cron.CronService, error) {
cronStorePath := filepath.Join(workspace, "cron", "jobs.json")
- // Create cron service
cronService := cron.NewCronService(cronStorePath, nil)
- // Create and register CronTool if enabled
var cronTool *tools.CronTool
if cfg.Tools.IsToolEnabled("cron") {
var err error
cronTool, err = tools.NewCronTool(cronService, agentLoop, msgBus, workspace, restrict, execTimeout, cfg)
if err != nil {
- logger.Fatalf("Critical error during CronTool initialization: %v", err)
+ return nil, fmt.Errorf("critical error during CronTool initialization: %w", err)
}
agentLoop.RegisterTool(cronTool)
}
- // Set onJob handler
if cronTool != nil {
cronService.SetOnJob(func(job *cron.CronJob) (string, error) {
result := cronTool.ExecuteJob(context.Background(), job)
@@ -639,5 +637,22 @@ func setupCronTool(
})
}
- return cronService
+ return cronService, nil
+}
+
+func createHeartbeatHandler(agentLoop *agent.AgentLoop) func(prompt, channel, chatID string) *tools.ToolResult {
+ return func(prompt, channel, chatID string) *tools.ToolResult {
+ if channel == "" || chatID == "" {
+ channel, chatID = "cli", "direct"
+ }
+
+ response, err := agentLoop.ProcessHeartbeat(context.Background(), prompt, channel, chatID)
+ if err != nil {
+ return tools.ErrorResult(fmt.Sprintf("Heartbeat error: %v", err))
+ }
+ if response == "HEARTBEAT_OK" {
+ return tools.SilentResult("Heartbeat OK")
+ }
+ return tools.SilentResult(response)
+ }
}
diff --git a/pkg/health/server.go b/pkg/health/server.go
index 5609ebdf6..fe20e4b94 100644
--- a/pkg/health/server.go
+++ b/pkg/health/server.go
@@ -6,16 +6,18 @@ import (
"fmt"
"maps"
"net/http"
+ "os"
"sync"
"time"
)
type Server struct {
- server *http.Server
- mu sync.RWMutex
- ready bool
- checks map[string]Check
- startTime time.Time
+ server *http.Server
+ mu sync.RWMutex
+ ready bool
+ checks map[string]Check
+ startTime time.Time
+ reloadFunc func() error
}
type Check struct {
@@ -29,6 +31,7 @@ type StatusResponse struct {
Status string `json:"status"`
Uptime string `json:"uptime"`
Checks map[string]Check `json:"checks,omitempty"`
+ Pid int `json:"pid"`
}
func NewServer(host string, port int) *Server {
@@ -41,6 +44,7 @@ func NewServer(host string, port int) *Server {
mux.HandleFunc("/health", s.healthHandler)
mux.HandleFunc("/ready", s.readyHandler)
+ mux.HandleFunc("/reload", s.reloadHandler)
addr := fmt.Sprintf("%s:%d", host, port)
s.server = &http.Server{
@@ -104,6 +108,44 @@ func (s *Server) RegisterCheck(name string, checkFn func() (bool, string)) {
}
}
+// SetReloadFunc sets the callback function for config reload.
+func (s *Server) SetReloadFunc(fn func() error) {
+ s.mu.Lock()
+ defer s.mu.Unlock()
+ s.reloadFunc = fn
+}
+
+func (s *Server) reloadHandler(w http.ResponseWriter, r *http.Request) {
+ if r.Method != http.MethodPost {
+ w.Header().Set("Content-Type", "application/json")
+ w.WriteHeader(http.StatusMethodNotAllowed)
+ json.NewEncoder(w).Encode(map[string]string{"error": "method not allowed, use POST"})
+ return
+ }
+
+ s.mu.Lock()
+ reloadFunc := s.reloadFunc
+ s.mu.Unlock()
+
+ if reloadFunc == nil {
+ w.Header().Set("Content-Type", "application/json")
+ w.WriteHeader(http.StatusServiceUnavailable)
+ json.NewEncoder(w).Encode(map[string]string{"error": "reload not configured"})
+ return
+ }
+
+ if err := reloadFunc(); err != nil {
+ w.Header().Set("Content-Type", "application/json")
+ w.WriteHeader(http.StatusInternalServerError)
+ json.NewEncoder(w).Encode(map[string]string{"error": err.Error()})
+ return
+ }
+
+ w.Header().Set("Content-Type", "application/json")
+ w.WriteHeader(http.StatusOK)
+ json.NewEncoder(w).Encode(map[string]string{"status": "reload triggered"})
+}
+
func (s *Server) healthHandler(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(http.StatusOK)
@@ -112,6 +154,7 @@ func (s *Server) healthHandler(w http.ResponseWriter, r *http.Request) {
resp := StatusResponse{
Status: "ok",
Uptime: uptime.String(),
+ Pid: os.Getpid(),
}
json.NewEncoder(w).Encode(resp)
@@ -155,11 +198,12 @@ func (s *Server) readyHandler(w http.ResponseWriter, r *http.Request) {
})
}
-// RegisterOnMux registers /health and /ready handlers onto the given mux.
+// RegisterOnMux registers /health, /ready and /reload handlers onto the given mux.
// This allows the health endpoints to be served by a shared HTTP server.
func (s *Server) RegisterOnMux(mux *http.ServeMux) {
mux.HandleFunc("/health", s.healthHandler)
mux.HandleFunc("/ready", s.readyHandler)
+ mux.HandleFunc("/reload", s.reloadHandler)
}
func statusString(ok bool) string {
diff --git a/pkg/heartbeat/service.go b/pkg/heartbeat/service.go
index 09c93fc6b..5dda78ea9 100644
--- a/pkg/heartbeat/service.go
+++ b/pkg/heartbeat/service.go
@@ -26,6 +26,7 @@ import (
const (
minIntervalMinutes = 5
defaultIntervalMinutes = 30
+ userTasksMarker = "Add your heartbeat tasks below this line:"
)
// HeartbeatHandler is the function type for handling heartbeat.
@@ -232,7 +233,7 @@ func (hs *HeartbeatService) buildPrompt() string {
}
content := string(data)
- if len(content) == 0 {
+ if !heartbeatHasUserTasks(content) {
return ""
}
@@ -284,6 +285,32 @@ Add your heartbeat tasks below this line:
}
}
+func heartbeatHasUserTasks(content string) bool {
+ trimmed := strings.TrimSpace(content)
+ if trimmed == "" {
+ return false
+ }
+
+ markerIdx := strings.Index(content, userTasksMarker)
+ if markerIdx < 0 {
+ return true
+ }
+
+ tasksSection := content[markerIdx+len(userTasksMarker):]
+ for _, line := range strings.Split(tasksSection, "\n") {
+ trimmedLine := strings.TrimSpace(line)
+ if trimmedLine == "" {
+ continue
+ }
+ if strings.HasPrefix(trimmedLine, "#") {
+ continue
+ }
+ return true
+ }
+
+ return false
+}
+
// sendResponse sends the heartbeat response to the last channel
func (hs *HeartbeatService) sendResponse(response string) {
hs.mu.RLock()
diff --git a/pkg/heartbeat/service_test.go b/pkg/heartbeat/service_test.go
index 3b7eeeefb..309b4378f 100644
--- a/pkg/heartbeat/service_test.go
+++ b/pkg/heartbeat/service_test.go
@@ -3,6 +3,7 @@ package heartbeat
import (
"os"
"path/filepath"
+ "strings"
"testing"
"time"
@@ -203,3 +204,47 @@ func TestHeartbeatFilePath(t *testing.T) {
t.Errorf("Expected HEARTBEAT.md at %s, but it doesn't exist", expectedPath)
}
}
+
+func TestBuildPrompt_DefaultTemplateStaysIdle(t *testing.T) {
+ tmpDir, err := os.MkdirTemp("", "heartbeat-test-*")
+ if err != nil {
+ t.Fatalf("Failed to create temp dir: %v", err)
+ }
+ defer os.RemoveAll(tmpDir)
+
+ hs := NewHeartbeatService(tmpDir, 30, true)
+ hs.createDefaultHeartbeatTemplate()
+
+ if prompt := hs.buildPrompt(); prompt != "" {
+ t.Fatalf("buildPrompt() = %q, want empty prompt for untouched default template", prompt)
+ }
+}
+
+func TestBuildPrompt_UserTasksAfterMarkerProducePrompt(t *testing.T) {
+ tmpDir, err := os.MkdirTemp("", "heartbeat-test-*")
+ if err != nil {
+ t.Fatalf("Failed to create temp dir: %v", err)
+ }
+ defer os.RemoveAll(tmpDir)
+
+ hs := NewHeartbeatService(tmpDir, 30, true)
+ hs.createDefaultHeartbeatTemplate()
+
+ path := filepath.Join(tmpDir, "HEARTBEAT.md")
+ data, err := os.ReadFile(path)
+ if err != nil {
+ t.Fatalf("Failed to read HEARTBEAT.md: %v", err)
+ }
+ updated := string(data) + "\n- Check unread Feishu messages\n"
+ if err := os.WriteFile(path, []byte(updated), 0o644); err != nil {
+ t.Fatalf("Failed to update HEARTBEAT.md: %v", err)
+ }
+
+ prompt := hs.buildPrompt()
+ if prompt == "" {
+ t.Fatal("buildPrompt() = empty, want non-empty prompt when user tasks are present")
+ }
+ if !strings.Contains(prompt, "Check unread Feishu messages") {
+ t.Fatalf("prompt = %q, want user task content", prompt)
+ }
+}
diff --git a/pkg/identity/identity.go b/pkg/identity/identity.go
index 372bbe38b..045725a8d 100644
--- a/pkg/identity/identity.go
+++ b/pkg/identity/identity.go
@@ -94,13 +94,18 @@ func MatchAllowed(sender bus.SenderInfo, allowed string) bool {
return false
}
-// isNumeric returns true if s consists entirely of digits.
+// isNumeric returns true if s consists entirely of digits, allowing for an optional leading minus sign
+// (required for Telegram group/channel IDs like -1001234567890).
func isNumeric(s string) bool {
if s == "" {
return false
}
- for _, r := range s {
- if r < '0' || r > '9' {
+ start := 0
+ if s[0] == '-' && len(s) > 1 {
+ start = 1
+ }
+ for i := start; i < len(s); i++ {
+ if s[i] < '0' || s[i] > '9' {
return false
}
}
diff --git a/pkg/identity/identity_test.go b/pkg/identity/identity_test.go
index a588f1484..c60402d19 100644
--- a/pkg/identity/identity_test.go
+++ b/pkg/identity/identity_test.go
@@ -97,6 +97,15 @@ func TestMatchAllowed(t *testing.T) {
allowed: "654321",
want: false,
},
+ {
+ name: "negative numeric ID matches PlatformID",
+ sender: bus.SenderInfo{
+ Platform: "telegram",
+ PlatformID: "-1001234567890",
+ },
+ allowed: "-1001234567890",
+ want: true,
+ },
// Username matching
{
name: "@username matches Username",
@@ -238,6 +247,9 @@ func TestIsNumeric(t *testing.T) {
{"abc", false},
{"12a34", false},
{"telegram", false},
+ {"-1001234567890", true},
+ {"-", false},
+ {"-12a34", false},
}
for _, tt := range tests {
diff --git a/pkg/logger/logger.go b/pkg/logger/logger.go
index 302613f33..179804607 100644
--- a/pkg/logger/logger.go
+++ b/pkg/logger/logger.go
@@ -5,6 +5,7 @@ import (
"os"
"path/filepath"
"runtime"
+ "strconv"
"strings"
"sync"
@@ -45,13 +46,47 @@ func init() {
consoleWriter := zerolog.ConsoleWriter{
Out: os.Stdout,
TimeFormat: "15:04:05", // TODO: make it configurable???
+
+ // Custom formatter to handle multiline strings and JSON objects
+ FormatFieldValue: formatFieldValue,
}
- logger = zerolog.New(consoleWriter).With().Timestamp().Logger()
+ logger = zerolog.New(consoleWriter).With().Timestamp().Caller().Logger()
fileLogger = zerolog.Logger{}
})
}
+func formatFieldValue(i any) string {
+ var s string
+
+ switch val := i.(type) {
+ case string:
+ s = val
+ case []byte:
+ s = string(val)
+ default:
+ return fmt.Sprintf("%v", i)
+ }
+
+ if unquoted, err := strconv.Unquote(s); err == nil {
+ s = unquoted
+ }
+
+ if strings.Contains(s, "\n") {
+ return fmt.Sprintf("\n%s", s)
+ }
+
+ if strings.Contains(s, " ") {
+ if (strings.HasPrefix(s, "{") && strings.HasSuffix(s, "}")) ||
+ (strings.HasPrefix(s, "[") && strings.HasSuffix(s, "]")) {
+ return s
+ }
+ return fmt.Sprintf("%q", s)
+ }
+
+ return s
+}
+
func SetLevel(level LogLevel) {
mu.Lock()
defer mu.Unlock()
@@ -59,12 +94,48 @@ func SetLevel(level LogLevel) {
zerolog.SetGlobalLevel(level)
}
+func SetConsoleLevel(level LogLevel) {
+ mu.Lock()
+ defer mu.Unlock()
+ logger = logger.Level(level)
+}
+
func GetLevel() LogLevel {
mu.RLock()
defer mu.RUnlock()
return currentLevel
}
+// ParseLevel converts a case-insensitive level name to a LogLevel.
+// Returns the level and true if valid, or (INFO, false) if unrecognized.
+func ParseLevel(s string) (LogLevel, bool) {
+ switch strings.ToLower(strings.TrimSpace(s)) {
+ case "debug":
+ return DEBUG, true
+ case "info":
+ return INFO, true
+ case "warn", "warning":
+ return WARN, true
+ case "error":
+ return ERROR, true
+ case "fatal":
+ return FATAL, true
+ default:
+ return INFO, false
+ }
+}
+
+// SetLevelFromString sets the log level from a string value.
+// If the string is empty or not a recognized level name, the current level is kept.
+func SetLevelFromString(s string) {
+ if s == "" {
+ return
+ }
+ if level, ok := ParseLevel(s); ok {
+ SetLevel(level)
+ }
+}
+
func EnableFileLogging(filePath string) error {
mu.Lock()
defer mu.Unlock()
@@ -99,9 +170,9 @@ func DisableFileLogging() {
fileLogger = zerolog.Logger{}
}
-func getCallerInfo() (string, int, string) {
+func getCallerSkip() int {
for i := 2; i < 15; i++ {
- pc, file, line, ok := runtime.Caller(i)
+ pc, file, _, ok := runtime.Caller(i)
if !ok {
continue
}
@@ -123,10 +194,10 @@ func getCallerInfo() (string, int, string) {
continue
}
- return filepath.Base(file), line, filepath.Base(funcName)
+ return i - 1
}
- return "???", 0, "???"
+ return 3
}
//nolint:zerologlint
@@ -152,22 +223,16 @@ func logMessage(level LogLevel, component string, message string, fields map[str
return
}
- callerFile, callerLine, callerFunc := getCallerInfo()
+ skip := getCallerSkip()
event := getEvent(logger, level)
- // Build combined field with component and caller
if component != "" {
- event.Str("caller", fmt.Sprintf("%-6s %s:%d (%s)", component, callerFile, callerLine, callerFunc))
- } else {
- event.Str("caller", fmt.Sprintf(" %s:%d (%s)", callerFile, callerLine, callerFunc))
+ event.Str("component", component)
}
- for k, v := range fields {
- event.Interface(k, v)
- }
-
- event.Msg(message)
+ appendFields(event, fields)
+ event.CallerSkipFrame(skip).Msg(message)
// Also log to file if enabled
if fileLogger.GetLevel() != zerolog.NoLevel {
@@ -176,10 +241,10 @@ func logMessage(level LogLevel, component string, message string, fields map[str
if component != "" {
fileEvent.Str("component", component)
}
- for k, v := range fields {
- fileEvent.Interface(k, v)
- }
- fileEvent.Msg(message)
+ // fileEvent.Str("caller", fmt.Sprintf("%s:%d (%s)", callerFile, callerLine, callerFunc))
+
+ appendFields(fileEvent, fields)
+ fileEvent.CallerSkipFrame(skip).Msg(message)
}
if level == FATAL {
@@ -187,6 +252,26 @@ func logMessage(level LogLevel, component string, message string, fields map[str
}
}
+func appendFields(event *zerolog.Event, fields map[string]any) {
+ for k, v := range fields {
+ // Type switch to avoid double JSON serialization of strings
+ switch val := v.(type) {
+ case string:
+ event.Str(k, val)
+ case int:
+ event.Int(k, val)
+ case int64:
+ event.Int64(k, val)
+ case float64:
+ event.Float64(k, val)
+ case bool:
+ event.Bool(k, val)
+ default:
+ event.Interface(k, v) // Fallback for struct, slice and maps
+ }
+ }
+}
+
func Debug(message string) {
logMessage(DEBUG, "", message, nil)
}
diff --git a/pkg/logger/logger_3rd_party.go b/pkg/logger/logger_3rd_party.go
index da50d686a..d0cb178c5 100644
--- a/pkg/logger/logger_3rd_party.go
+++ b/pkg/logger/logger_3rd_party.go
@@ -2,7 +2,20 @@
package logger
-import "fmt"
+import (
+ "fmt"
+ "regexp"
+)
+
+// botTokenRe matches the bot ID prefix and the secret part of a Telegram bot token.
+// Groups: 1 = "bot:", 2 = first 4 chars of secret, 3 = middle, 4 = last 4 chars.
+var botTokenRe = regexp.MustCompile(`(bot\d+:)([A-Za-z0-9_-]{4})[A-Za-z0-9_-]{12,}([A-Za-z0-9_-]{4})`)
+
+// maskSecrets replaces any embedded bot tokens in s with a redacted placeholder
+// that keeps the first and last 4 characters of the secret for identification.
+func maskSecrets(s string) string {
+ return botTokenRe.ReplaceAllString(s, "${1}${2}****${3}")
+}
// Logger implements common Logger interface
type Logger struct {
@@ -12,52 +25,52 @@ type Logger struct {
// Debug logs debug messages
func (b *Logger) Debug(v ...any) {
- logMessage(DEBUG, b.component, fmt.Sprint(v...), nil)
+ logMessage(DEBUG, b.component, maskSecrets(fmt.Sprint(v...)), nil)
}
// Info logs info messages
func (b *Logger) Info(v ...any) {
- logMessage(INFO, b.component, fmt.Sprint(v...), nil)
+ logMessage(INFO, b.component, maskSecrets(fmt.Sprint(v...)), nil)
}
// Warn logs warning messages
func (b *Logger) Warn(v ...any) {
- logMessage(WARN, b.component, fmt.Sprint(v...), nil)
+ logMessage(WARN, b.component, maskSecrets(fmt.Sprint(v...)), nil)
}
// Error logs error messages
func (b *Logger) Error(v ...any) {
- logMessage(ERROR, b.component, fmt.Sprint(v...), nil)
+ logMessage(ERROR, b.component, maskSecrets(fmt.Sprint(v...)), nil)
}
// Debugf logs formatted debug messages
func (b *Logger) Debugf(format string, v ...any) {
- logMessage(DEBUG, b.component, fmt.Sprintf(format, v...), nil)
+ logMessage(DEBUG, b.component, maskSecrets(fmt.Sprintf(format, v...)), nil)
}
// Infof logs formatted info messages
func (b *Logger) Infof(format string, v ...any) {
- logMessage(INFO, b.component, fmt.Sprintf(format, v...), nil)
+ logMessage(INFO, b.component, maskSecrets(fmt.Sprintf(format, v...)), nil)
}
// Warnf logs formatted warning messages
func (b *Logger) Warnf(format string, v ...any) {
- logMessage(WARN, b.component, fmt.Sprintf(format, v...), nil)
+ logMessage(WARN, b.component, maskSecrets(fmt.Sprintf(format, v...)), nil)
}
// Warningf logs formatted warning messages
func (b *Logger) Warningf(format string, v ...any) {
- logMessage(WARN, b.component, fmt.Sprintf(format, v...), nil)
+ logMessage(WARN, b.component, maskSecrets(fmt.Sprintf(format, v...)), nil)
}
// Errorf logs formatted error messages
func (b *Logger) Errorf(format string, v ...any) {
- logMessage(ERROR, b.component, fmt.Sprintf(format, v...), nil)
+ logMessage(ERROR, b.component, maskSecrets(fmt.Sprintf(format, v...)), nil)
}
// Fatalf logs formatted fatal messages and exits
func (b *Logger) Fatalf(format string, v ...any) {
- logMessage(FATAL, b.component, fmt.Sprintf(format, v...), nil)
+ logMessage(FATAL, b.component, maskSecrets(fmt.Sprintf(format, v...)), nil)
}
// Log logs a message at a given level with caller information
@@ -75,7 +88,7 @@ func (b *Logger) Log(msgL, caller int, format string, a ...any) {
level = lvl
}
}
- logMessage(level, b.component, fmt.Sprintf(format, a...), nil)
+ logMessage(level, b.component, maskSecrets(fmt.Sprintf(format, a...)), nil)
}
// Sync flushes log buffer (no-op for this implementation)
diff --git a/pkg/logger/logger_test.go b/pkg/logger/logger_test.go
index 8170a618b..e551db58e 100644
--- a/pkg/logger/logger_test.go
+++ b/pkg/logger/logger_test.go
@@ -141,3 +141,199 @@ func TestLoggerHelperFunctions(t *testing.T) {
Debugf("test from %v", "Debugf")
WarnF("Warning with fields", map[string]any{"key": "value"})
}
+
+func TestFormatFieldValue(t *testing.T) {
+ tests := []struct {
+ name string
+ input any
+ expected string
+ }{
+ // Basic types test (default case of the switch)
+ {
+ name: "Integer Type",
+ input: 42,
+ expected: "42",
+ },
+ {
+ name: "Boolean Type",
+ input: true,
+ expected: "true",
+ },
+ {
+ name: "Unsupported Struct Type",
+ input: struct{ A int }{A: 1},
+ expected: "{1}",
+ },
+
+ // Simple strings and byte slices test
+ {
+ name: "Simple string without spaces",
+ input: "simple_value",
+ expected: "simple_value",
+ },
+ {
+ name: "Simple byte slice",
+ input: []byte("byte_value"),
+ expected: "byte_value",
+ },
+
+ // Unquoting test (strconv.Unquote)
+ {
+ name: "Quoted string",
+ input: `"quoted_value"`,
+ expected: "quoted_value",
+ },
+
+ // Strings with newline (\n) test
+ {
+ name: "String with newline",
+ input: "line1\nline2",
+ expected: "\nline1\nline2",
+ },
+ {
+ name: "Quoted string with newline (Unquote -> newline)",
+ input: `"line1\nline2"`, // Escaped \n that Unquote will resolve
+ expected: "\nline1\nline2",
+ },
+
+ // Strings with spaces test (which should be quoted)
+ {
+ name: "String with spaces",
+ input: "hello world",
+ expected: `"hello world"`,
+ },
+ {
+ name: "Quoted string with spaces (Unquote -> has spaces -> Re-quote)",
+ input: `"hello world"`,
+ expected: `"hello world"`,
+ },
+
+ // JSON formats test (strings with spaces that start/end with brackets)
+ {
+ name: "Valid JSON object",
+ input: `{"key": "value"}`,
+ expected: `{"key": "value"}`,
+ },
+ {
+ name: "Valid JSON array",
+ input: `[1, 2, "three"]`,
+ expected: `[1, 2, "three"]`,
+ },
+ {
+ name: "Fake JSON (starts with { but doesn't end with })",
+ input: `{"key": "value"`, // Missing closing bracket, has spaces
+ expected: `"{\"key\": \"value\""`,
+ },
+ {
+ name: "Empty JSON (object)",
+ input: `{ }`,
+ expected: `{ }`,
+ },
+
+ // 7. Edge Cases
+ {
+ name: "Empty string",
+ input: "",
+ expected: "",
+ },
+ {
+ name: "Whitespace only string",
+ input: " ",
+ expected: `" "`,
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ actual := formatFieldValue(tt.input)
+ if actual != tt.expected {
+ t.Errorf("formatFieldValue() = %q, expected %q", actual, tt.expected)
+ }
+ })
+ }
+}
+
+func TestDefaultLevelIsInfo(t *testing.T) {
+ // The package-level default (before any SetLevel call) should be INFO.
+ // Because earlier tests may have changed it, we just verify the constant is wired correctly.
+ if logLevelNames[INFO] != "INFO" {
+ t.Errorf("INFO constant mapped to %q, want \"INFO\"", logLevelNames[INFO])
+ }
+}
+
+func TestParseLevelValid(t *testing.T) {
+ tests := []struct {
+ input string
+ want LogLevel
+ }{
+ {"debug", DEBUG},
+ {"DEBUG", DEBUG},
+ {"Debug", DEBUG},
+ {"info", INFO},
+ {"INFO", INFO},
+ {"warn", WARN},
+ {"WARN", WARN},
+ {"warning", WARN},
+ {"WARNING", WARN},
+ {"error", ERROR},
+ {"ERROR", ERROR},
+ {"fatal", FATAL},
+ {"FATAL", FATAL},
+ {" info ", INFO},
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.input, func(t *testing.T) {
+ got, ok := ParseLevel(tt.input)
+ if !ok {
+ t.Fatalf("ParseLevel(%q) returned ok=false, want true", tt.input)
+ }
+ if got != tt.want {
+ t.Errorf("ParseLevel(%q) = %v, want %v", tt.input, got, tt.want)
+ }
+ })
+ }
+}
+
+func TestParseLevelInvalid(t *testing.T) {
+ tests := []string{"", "garbage", "verbose", "trace", "critical"}
+
+ for _, input := range tests {
+ t.Run(input, func(t *testing.T) {
+ _, ok := ParseLevel(input)
+ if ok {
+ t.Errorf("ParseLevel(%q) returned ok=true, want false", input)
+ }
+ })
+ }
+}
+
+func TestSetLevelFromString(t *testing.T) {
+ initialLevel := GetLevel()
+ defer SetLevel(initialLevel)
+
+ // Valid string changes the level
+ SetLevel(INFO)
+ SetLevelFromString("error")
+ if got := GetLevel(); got != ERROR {
+ t.Errorf("after SetLevelFromString(\"error\"): GetLevel() = %v, want ERROR", got)
+ }
+
+ // Empty string is a no-op
+ SetLevelFromString("")
+ if got := GetLevel(); got != ERROR {
+ t.Errorf("after SetLevelFromString(\"\"): GetLevel() = %v, want ERROR (unchanged)", got)
+ }
+
+ // Invalid string is a no-op
+ SetLevelFromString("garbage")
+ if got := GetLevel(); got != ERROR {
+ t.Errorf("after SetLevelFromString(\"garbage\"): GetLevel() = %v, want ERROR (unchanged)", got)
+ }
+
+ // Case-insensitive
+ SetLevelFromString("FATAL")
+ if got := GetLevel(); got != FATAL {
+ t.Errorf("after SetLevelFromString(\"FATAL\"): GetLevel() = %v, want FATAL", got)
+ }
+}
diff --git a/pkg/media/tempdir.go b/pkg/media/tempdir.go
new file mode 100644
index 000000000..45942b34f
--- /dev/null
+++ b/pkg/media/tempdir.go
@@ -0,0 +1,13 @@
+package media
+
+import (
+ "os"
+ "path/filepath"
+)
+
+const TempDirName = "picoclaw_media"
+
+// TempDir returns the shared temporary directory used for downloaded media.
+func TempDir() string {
+ return filepath.Join(os.TempDir(), TempDirName)
+}
diff --git a/pkg/migrate/internal/common.go b/pkg/migrate/internal/common.go
index c77ab9f26..75aef5dc2 100644
--- a/pkg/migrate/internal/common.go
+++ b/pkg/migrate/internal/common.go
@@ -5,13 +5,15 @@ import (
"io"
"os"
"path/filepath"
+
+ "github.com/sipeed/picoclaw/pkg/config"
)
func ResolveTargetHome(override string) (string, error) {
if override != "" {
return ExpandHome(override), nil
}
- if envHome := os.Getenv("PICOCLAW_HOME"); envHome != "" {
+ if envHome := os.Getenv(config.EnvHome); envHome != "" {
return ExpandHome(envHome), nil
}
home, err := os.UserHomeDir()
diff --git a/pkg/migrate/sources/openclaw/openclaw_config.go b/pkg/migrate/sources/openclaw/openclaw_config.go
index e95c2f3ec..317bd3e84 100644
--- a/pkg/migrate/sources/openclaw/openclaw_config.go
+++ b/pkg/migrate/sources/openclaw/openclaw_config.go
@@ -132,11 +132,12 @@ type OpenClawChannels struct {
}
type OpenClawTelegramConfig struct {
- BotToken *string `json:"botToken"`
- AllowFrom []string `json:"allowFrom"`
- GroupPolicy *string `json:"groupPolicy"`
- DmPolicy *string `json:"dmPolicy"`
- Enabled *bool `json:"enabled"`
+ BotToken *string `json:"botToken"`
+ AllowFrom []string `json:"allowFrom"`
+ GroupPolicy *string `json:"groupPolicy"`
+ DmPolicy *string `json:"dmPolicy"`
+ Enabled *bool `json:"enabled"`
+ UseMarkdownV2 *bool `json:"useMarkdownV2"`
}
type OpenClawDiscordConfig struct {
@@ -645,10 +646,11 @@ type WhatsAppConfig struct {
}
type TelegramConfig struct {
- Enabled bool `json:"enabled"`
- Token string `json:"token"`
- Proxy string `json:"proxy"`
- AllowFrom []string `json:"allow_from"`
+ Enabled bool `json:"enabled"`
+ Token string `json:"token"`
+ Proxy string `json:"proxy"`
+ AllowFrom []string `json:"allow_from"`
+ UseMarkdownV2 bool `json:"use_markdown_v2"`
}
type FeishuConfig struct {
@@ -777,9 +779,11 @@ func (c *OpenClawConfig) convertChannels(warnings *[]string) ChannelsConfig {
if c.Channels.Telegram != nil {
enabled := c.Channels.Telegram.Enabled == nil || *c.Channels.Telegram.Enabled
+ useMarkdownV2 := c.Channels.Telegram.UseMarkdownV2 != nil && *c.Channels.Telegram.UseMarkdownV2
channels.Telegram = TelegramConfig{
- Enabled: enabled,
- AllowFrom: c.Channels.Telegram.AllowFrom,
+ Enabled: enabled,
+ AllowFrom: c.Channels.Telegram.AllowFrom,
+ UseMarkdownV2: useMarkdownV2,
}
if c.Channels.Telegram.BotToken != nil {
channels.Telegram.Token = *c.Channels.Telegram.BotToken
diff --git a/pkg/migrate/sources/openclaw/openclaw_handler.go b/pkg/migrate/sources/openclaw/openclaw_handler.go
index aaff119f1..5e5241268 100644
--- a/pkg/migrate/sources/openclaw/openclaw_handler.go
+++ b/pkg/migrate/sources/openclaw/openclaw_handler.go
@@ -10,6 +10,11 @@ import (
"github.com/sipeed/picoclaw/pkg/migrate/internal"
)
+// OpenclawHomeEnvVar is the environment variable that overrides the source
+// openclaw home directory when migrating from openclaw to picoclaw.
+// Default: ~/.openclaw
+const OpenclawHomeEnvVar = "OPENCLAW_HOME"
+
var providerMapping = map[string]string{
"anthropic": "anthropic",
"claude": "anthropic",
@@ -112,7 +117,7 @@ func resolveSourceHome(override string) (string, error) {
if override != "" {
return internal.ExpandHome(override), nil
}
- if envHome := os.Getenv("OPENCLAW_HOME"); envHome != "" {
+ if envHome := os.Getenv(OpenclawHomeEnvVar); envHome != "" {
return internal.ExpandHome(envHome), nil
}
home, err := os.UserHomeDir()
diff --git a/pkg/providers/anthropic/provider.go b/pkg/providers/anthropic/provider.go
index 242ded175..d4ceaab2c 100644
--- a/pkg/providers/anthropic/provider.go
+++ b/pkg/providers/anthropic/provider.go
@@ -180,6 +180,10 @@ func buildParams(
blocks = append(blocks, anthropic.NewTextBlock(msg.Content))
}
for _, tc := range msg.ToolCalls {
+ // Skip tool calls with empty names to avoid API errors
+ if tc.Name == "" {
+ continue
+ }
args := tc.Arguments
if args == nil && tc.Function != nil && tc.Function.Arguments != "" {
if err := json.Unmarshal([]byte(tc.Function.Arguments), &args); err != nil {
diff --git a/pkg/providers/anthropic_messages/provider.go b/pkg/providers/anthropic_messages/provider.go
index 8a83a7058..2b19e941a 100644
--- a/pkg/providers/anthropic_messages/provider.go
+++ b/pkg/providers/anthropic_messages/provider.go
@@ -221,11 +221,21 @@ func buildRequestBody(
// Add tool_use blocks
for _, tc := range msg.ToolCalls {
+ if strings.TrimSpace(tc.Name) == "" {
+ continue
+ }
+
+ // Handle nil Arguments (GLM-4 may return null input)
+ input := tc.Arguments
+ if input == nil {
+ input = map[string]any{}
+ }
+
toolUse := map[string]any{
"type": "tool_use",
"id": tc.ID,
"name": tc.Name,
- "input": tc.Arguments,
+ "input": input,
}
content = append(content, toolUse)
}
diff --git a/pkg/providers/anthropic_messages/provider_test.go b/pkg/providers/anthropic_messages/provider_test.go
index da4213e92..8eabc15fa 100644
--- a/pkg/providers/anthropic_messages/provider_test.go
+++ b/pkg/providers/anthropic_messages/provider_test.go
@@ -492,6 +492,20 @@ func TestBuildRequestBodyEdgeCases(t *testing.T) {
},
wantErr: false,
},
+ {
+ name: "skip tool calls with empty names",
+ messages: []Message{
+ {Role: "assistant", Content: "Calling tool", ToolCalls: []ToolCall{
+ {ID: "tool-empty", Name: "", Arguments: map[string]any{"ignored": true}},
+ {ID: "tool-valid", Name: "test_tool", Arguments: map[string]any{"arg": "value"}},
+ }},
+ },
+ model: "test-model",
+ options: map[string]any{
+ "max_tokens": 8192,
+ },
+ wantErr: false,
+ },
}
for _, tt := range tests {
@@ -513,6 +527,37 @@ func TestBuildRequestBodyEdgeCases(t *testing.T) {
if got["model"] != tt.model {
t.Errorf("model = %v, want %v", got["model"], tt.model)
}
+
+ if tt.name == "skip tool calls with empty names" {
+ messages, ok := got["messages"].([]any)
+ if !ok || len(messages) != 1 {
+ t.Fatalf("messages = %#v, want single assistant message", got["messages"])
+ }
+
+ assistantMsg, ok := messages[0].(map[string]any)
+ if !ok {
+ t.Fatalf("assistant message = %#v, want map", messages[0])
+ }
+
+ content, ok := assistantMsg["content"].([]any)
+ if !ok {
+ t.Fatalf("assistant content = %#v, want []any", assistantMsg["content"])
+ }
+ if len(content) != 2 {
+ t.Fatalf("assistant content length = %d, want 2", len(content))
+ }
+
+ toolUse, ok := content[1].(map[string]any)
+ if !ok {
+ t.Fatalf("tool_use block = %#v, want map", content[1])
+ }
+ if gotName := toolUse["name"]; gotName != "test_tool" {
+ t.Fatalf("tool_use name = %v, want %q", gotName, "test_tool")
+ }
+ if gotID := toolUse["id"]; gotID != "tool-valid" {
+ t.Fatalf("tool_use id = %v, want %q", gotID, "tool-valid")
+ }
+ }
})
}
}
diff --git a/pkg/providers/azure/provider.go b/pkg/providers/azure/provider.go
new file mode 100644
index 000000000..e0ddbbde4
--- /dev/null
+++ b/pkg/providers/azure/provider.go
@@ -0,0 +1,150 @@
+package azure
+
+import (
+ "bytes"
+ "context"
+ "encoding/json"
+ "fmt"
+ "net/http"
+ "net/url"
+ "strings"
+ "time"
+
+ "github.com/sipeed/picoclaw/pkg/providers/common"
+ "github.com/sipeed/picoclaw/pkg/providers/protocoltypes"
+)
+
+type (
+ LLMResponse = protocoltypes.LLMResponse
+ Message = protocoltypes.Message
+ ToolDefinition = protocoltypes.ToolDefinition
+)
+
+const (
+ // azureAPIVersion is the Azure OpenAI API version used for all requests.
+ azureAPIVersion = "2024-10-21"
+ defaultRequestTimeout = common.DefaultRequestTimeout
+)
+
+// Provider implements the LLM provider interface for Azure OpenAI endpoints.
+// It handles Azure-specific authentication (api-key header), URL construction
+// (deployment-based), and request body formatting (max_completion_tokens, no model field).
+type Provider struct {
+ apiKey string
+ apiBase string
+ httpClient *http.Client
+}
+
+// Option configures the Azure Provider.
+type Option func(*Provider)
+
+// WithRequestTimeout sets the HTTP request timeout.
+func WithRequestTimeout(timeout time.Duration) Option {
+ return func(p *Provider) {
+ if timeout > 0 {
+ p.httpClient.Timeout = timeout
+ }
+ }
+}
+
+// NewProvider creates a new Azure OpenAI provider.
+func NewProvider(apiKey, apiBase, proxy string, opts ...Option) *Provider {
+ p := &Provider{
+ apiKey: apiKey,
+ apiBase: strings.TrimRight(apiBase, "/"),
+ httpClient: common.NewHTTPClient(proxy),
+ }
+
+ for _, opt := range opts {
+ if opt != nil {
+ opt(p)
+ }
+ }
+
+ return p
+}
+
+// NewProviderWithTimeout creates a new Azure OpenAI provider with a custom request timeout in seconds.
+func NewProviderWithTimeout(apiKey, apiBase, proxy string, requestTimeoutSeconds int) *Provider {
+ return NewProvider(
+ apiKey, apiBase, proxy,
+ WithRequestTimeout(time.Duration(requestTimeoutSeconds)*time.Second),
+ )
+}
+
+// Chat sends a chat completion request to the Azure OpenAI endpoint.
+// The model parameter is used as the Azure deployment name in the URL.
+func (p *Provider) Chat(
+ ctx context.Context,
+ messages []Message,
+ tools []ToolDefinition,
+ model string,
+ options map[string]any,
+) (*LLMResponse, error) {
+ if p.apiBase == "" {
+ return nil, fmt.Errorf("Azure API base not configured")
+ }
+
+ // model is the deployment name for Azure OpenAI
+ deployment := model
+
+ // Build Azure-specific URL safely using url.JoinPath and query encoding
+ // to prevent path traversal or query injection via deployment names.
+ base, err := url.JoinPath(p.apiBase, "openai/deployments", deployment, "chat/completions")
+ if err != nil {
+ return nil, fmt.Errorf("failed to build Azure request URL: %w", err)
+ }
+ requestURL := base + "?api-version=" + azureAPIVersion
+
+ // Build request body — no "model" field (Azure infers from deployment URL)
+ requestBody := map[string]any{
+ "messages": common.SerializeMessages(messages),
+ }
+
+ if len(tools) > 0 {
+ requestBody["tools"] = tools
+ requestBody["tool_choice"] = "auto"
+ }
+
+ // Azure OpenAI always uses max_completion_tokens
+ if maxTokens, ok := common.AsInt(options["max_tokens"]); ok {
+ requestBody["max_completion_tokens"] = maxTokens
+ }
+
+ if temperature, ok := common.AsFloat(options["temperature"]); ok {
+ requestBody["temperature"] = temperature
+ }
+
+ jsonData, err := json.Marshal(requestBody)
+ if err != nil {
+ return nil, fmt.Errorf("failed to marshal request: %w", err)
+ }
+
+ req, err := http.NewRequestWithContext(ctx, "POST", requestURL, bytes.NewReader(jsonData))
+ if err != nil {
+ return nil, fmt.Errorf("failed to create request: %w", err)
+ }
+
+ // Azure uses api-key header instead of Authorization: Bearer
+ req.Header.Set("Content-Type", "application/json")
+ if p.apiKey != "" {
+ req.Header.Set("Api-Key", p.apiKey)
+ }
+
+ resp, err := p.httpClient.Do(req)
+ if err != nil {
+ return nil, fmt.Errorf("failed to send request: %w", err)
+ }
+ defer resp.Body.Close()
+
+ if resp.StatusCode != http.StatusOK {
+ return nil, common.HandleErrorResponse(resp, p.apiBase)
+ }
+
+ return common.ReadAndParseResponse(resp, p.apiBase)
+}
+
+// GetDefaultModel returns an empty string as Azure deployments are user-configured.
+func (p *Provider) GetDefaultModel() string {
+ return ""
+}
diff --git a/pkg/providers/azure/provider_test.go b/pkg/providers/azure/provider_test.go
new file mode 100644
index 000000000..531b81296
--- /dev/null
+++ b/pkg/providers/azure/provider_test.go
@@ -0,0 +1,232 @@
+package azure
+
+import (
+ "encoding/json"
+ "net/http"
+ "net/http/httptest"
+ "testing"
+ "time"
+)
+
+// writeValidResponse writes a minimal valid Azure OpenAI chat completion response.
+func writeValidResponse(w http.ResponseWriter) {
+ resp := map[string]any{
+ "choices": []map[string]any{
+ {
+ "message": map[string]any{"content": "ok"},
+ "finish_reason": "stop",
+ },
+ },
+ }
+ w.Header().Set("Content-Type", "application/json")
+ json.NewEncoder(w).Encode(resp)
+}
+
+func TestProviderChat_AzureURLConstruction(t *testing.T) {
+ var capturedPath string
+ var capturedAPIVersion string
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ capturedPath = r.URL.Path
+ capturedAPIVersion = r.URL.Query().Get("api-version")
+ writeValidResponse(w)
+ }))
+ defer server.Close()
+
+ p := NewProvider("test-key", server.URL, "")
+ _, err := p.Chat(t.Context(), []Message{{Role: "user", Content: "hi"}}, nil, "my-gpt5-deployment", nil)
+ if err != nil {
+ t.Fatalf("Chat() error = %v", err)
+ }
+
+ wantPath := "/openai/deployments/my-gpt5-deployment/chat/completions"
+ if capturedPath != wantPath {
+ t.Errorf("URL path = %q, want %q", capturedPath, wantPath)
+ }
+ if capturedAPIVersion != azureAPIVersion {
+ t.Errorf("api-version = %q, want %q", capturedAPIVersion, azureAPIVersion)
+ }
+}
+
+func TestProviderChat_AzureAuthHeader(t *testing.T) {
+ var capturedAPIKey string
+ var capturedAuth string
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ capturedAPIKey = r.Header.Get("Api-Key")
+ capturedAuth = r.Header.Get("Authorization")
+ writeValidResponse(w)
+ }))
+ defer server.Close()
+
+ p := NewProvider("test-azure-key", server.URL, "")
+ _, err := p.Chat(t.Context(), []Message{{Role: "user", Content: "hi"}}, nil, "deployment", nil)
+ if err != nil {
+ t.Fatalf("Chat() error = %v", err)
+ }
+
+ if capturedAPIKey != "test-azure-key" {
+ t.Errorf("api-key header = %q, want %q", capturedAPIKey, "test-azure-key")
+ }
+ if capturedAuth != "" {
+ t.Errorf("Authorization header should be empty, got %q", capturedAuth)
+ }
+}
+
+func TestProviderChat_AzureOmitsModelFromBody(t *testing.T) {
+ var requestBody map[string]any
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ json.NewDecoder(r.Body).Decode(&requestBody)
+ writeValidResponse(w)
+ }))
+ defer server.Close()
+
+ p := NewProvider("test-key", server.URL, "")
+ _, err := p.Chat(t.Context(), []Message{{Role: "user", Content: "hi"}}, nil, "deployment", nil)
+ if err != nil {
+ t.Fatalf("Chat() error = %v", err)
+ }
+
+ if _, exists := requestBody["model"]; exists {
+ t.Error("request body should not contain 'model' field for Azure OpenAI")
+ }
+}
+
+func TestProviderChat_AzureUsesMaxCompletionTokens(t *testing.T) {
+ var requestBody map[string]any
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ json.NewDecoder(r.Body).Decode(&requestBody)
+ writeValidResponse(w)
+ }))
+ defer server.Close()
+
+ p := NewProvider("test-key", server.URL, "")
+ _, err := p.Chat(
+ t.Context(),
+ []Message{{Role: "user", Content: "hi"}},
+ nil,
+ "deployment",
+ map[string]any{"max_tokens": 2048},
+ )
+ if err != nil {
+ t.Fatalf("Chat() error = %v", err)
+ }
+
+ if _, exists := requestBody["max_completion_tokens"]; !exists {
+ t.Error("request body should contain 'max_completion_tokens'")
+ }
+ if _, exists := requestBody["max_tokens"]; exists {
+ t.Error("request body should not contain 'max_tokens'")
+ }
+}
+
+func TestProviderChat_AzureHTTPError(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ http.Error(w, `{"error":"unauthorized"}`, http.StatusUnauthorized)
+ }))
+ defer server.Close()
+
+ p := NewProvider("bad-key", server.URL, "")
+ _, err := p.Chat(t.Context(), []Message{{Role: "user", Content: "hi"}}, nil, "deployment", nil)
+ if err == nil {
+ t.Fatal("expected error, got nil")
+ }
+}
+
+func TestProviderChat_AzureParseToolCalls(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ resp := map[string]any{
+ "choices": []map[string]any{
+ {
+ "message": map[string]any{
+ "content": "",
+ "tool_calls": []map[string]any{
+ {
+ "id": "call_1",
+ "type": "function",
+ "function": map[string]any{
+ "name": "get_weather",
+ "arguments": `{"city":"Seattle"}`,
+ },
+ },
+ },
+ },
+ "finish_reason": "tool_calls",
+ },
+ },
+ }
+ w.Header().Set("Content-Type", "application/json")
+ json.NewEncoder(w).Encode(resp)
+ }))
+ defer server.Close()
+
+ p := NewProvider("test-key", server.URL, "")
+ out, err := p.Chat(t.Context(), []Message{{Role: "user", Content: "weather?"}}, nil, "deployment", nil)
+ if err != nil {
+ t.Fatalf("Chat() error = %v", err)
+ }
+
+ if len(out.ToolCalls) != 1 {
+ t.Fatalf("len(ToolCalls) = %d, want 1", len(out.ToolCalls))
+ }
+ if out.ToolCalls[0].Name != "get_weather" {
+ t.Errorf("ToolCalls[0].Name = %q, want %q", out.ToolCalls[0].Name, "get_weather")
+ }
+}
+
+func TestProvider_AzureEmptyAPIBase(t *testing.T) {
+ p := NewProvider("test-key", "", "")
+ _, err := p.Chat(t.Context(), []Message{{Role: "user", Content: "hi"}}, nil, "deployment", nil)
+ if err == nil {
+ t.Fatal("expected error for empty API base")
+ }
+}
+
+func TestProvider_AzureRequestTimeoutDefault(t *testing.T) {
+ p := NewProvider("test-key", "https://example.com", "")
+ if p.httpClient.Timeout != defaultRequestTimeout {
+ t.Errorf("timeout = %v, want %v", p.httpClient.Timeout, defaultRequestTimeout)
+ }
+}
+
+func TestProvider_AzureRequestTimeoutOverride(t *testing.T) {
+ p := NewProvider("test-key", "https://example.com", "", WithRequestTimeout(300*time.Second))
+ if p.httpClient.Timeout != 300*time.Second {
+ t.Errorf("timeout = %v, want %v", p.httpClient.Timeout, 300*time.Second)
+ }
+}
+
+func TestProvider_AzureNewProviderWithTimeout(t *testing.T) {
+ p := NewProviderWithTimeout("test-key", "https://example.com", "", 180)
+ if p.httpClient.Timeout != 180*time.Second {
+ t.Errorf("timeout = %v, want %v", p.httpClient.Timeout, 180*time.Second)
+ }
+}
+
+func TestProviderChat_AzureDeploymentNameEscaped(t *testing.T) {
+ var capturedPath string
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ capturedPath = r.URL.RawPath // use RawPath to see percent-encoding
+ if capturedPath == "" {
+ capturedPath = r.URL.Path
+ }
+ writeValidResponse(w)
+ }))
+ defer server.Close()
+
+ p := NewProvider("test-key", server.URL, "")
+
+ // Deployment name with characters that could cause path injection
+ _, err := p.Chat(t.Context(), []Message{{Role: "user", Content: "hi"}}, nil, "my deploy/../../admin", nil)
+ if err != nil {
+ t.Fatalf("Chat() error = %v", err)
+ }
+
+ // The slash and special chars in the deployment name must be escaped, not treated as path separators
+ if capturedPath == "/openai/deployments/my deploy/../../admin/chat/completions" {
+ t.Fatal("deployment name was interpolated without escaping — path injection possible")
+ }
+}
diff --git a/pkg/providers/claude_cli_provider.go b/pkg/providers/claude_cli_provider.go
index 6c4f6a767..40b581490 100644
--- a/pkg/providers/claude_cli_provider.go
+++ b/pkg/providers/claude_cli_provider.go
@@ -50,10 +50,18 @@ func (p *ClaudeCliProvider) Chat(
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
- if stderrStr := stderr.String(); stderrStr != "" {
+ stderrStr := strings.TrimSpace(stderr.String())
+ stdoutStr := strings.TrimSpace(stdout.String())
+ switch {
+ case stderrStr != "" && stdoutStr != "":
+ return nil, fmt.Errorf("claude cli error: %w\nstderr: %s\nstdout: %s", err, stderrStr, stdoutStr)
+ case stderrStr != "":
return nil, fmt.Errorf("claude cli error: %s", stderrStr)
+ case stdoutStr != "":
+ return nil, fmt.Errorf("claude cli error: %w\noutput: %s", err, stdoutStr)
+ default:
+ return nil, fmt.Errorf("claude cli error: %w", err)
}
- return nil, fmt.Errorf("claude cli error: %w", err)
}
return p.parseClaudeCliResponse(stdout.String())
diff --git a/pkg/providers/codex_cli_credentials.go b/pkg/providers/codex_cli_credentials.go
index 40f3ee2a1..c5b25f040 100644
--- a/pkg/providers/codex_cli_credentials.go
+++ b/pkg/providers/codex_cli_credentials.go
@@ -8,6 +8,11 @@ import (
"time"
)
+// CodexHomeEnvVar is the environment variable that overrides the Codex CLI
+// home directory when resolving the codex auth.json credentials file.
+// Default: ~/.codex
+const CodexHomeEnvVar = "CODEX_HOME"
+
// CodexCliAuth represents the ~/.codex/auth.json file structure.
type CodexCliAuth struct {
Tokens struct {
@@ -69,7 +74,7 @@ func CreateCodexCliTokenSource() func() (string, string, error) {
}
func resolveCodexAuthPath() (string, error) {
- codexHome := os.Getenv("CODEX_HOME")
+ codexHome := os.Getenv(CodexHomeEnvVar)
if codexHome == "" {
home, err := os.UserHomeDir()
if err != nil {
diff --git a/pkg/providers/codex_provider.go b/pkg/providers/codex_provider.go
index cf5c2d876..4a6d61a4b 100644
--- a/pkg/providers/codex_provider.go
+++ b/pkg/providers/codex_provider.go
@@ -95,7 +95,10 @@ func (p *CodexProvider) Chat(
)
}
- params := buildCodexParams(messages, tools, resolvedModel, options, p.enableWebSearch)
+ // Respect tools.web.prefer_native: only inject native search when the agent
+ // loop requested it (options["native_search"]), so prefer_native: false
+ useNativeSearch := p.enableWebSearch && (options["native_search"] == true)
+ params := buildCodexParams(messages, tools, resolvedModel, options, useNativeSearch)
stream := p.client.Responses.NewStreaming(ctx, params, opts...)
defer stream.Close()
@@ -157,6 +160,10 @@ func (p *CodexProvider) GetDefaultModel() string {
return codexDefaultModel
}
+func (p *CodexProvider) SupportsNativeSearch() bool {
+ return p.enableWebSearch
+}
+
func resolveCodexModel(model string) (string, string) {
m := strings.ToLower(strings.TrimSpace(model))
if m == "" {
diff --git a/pkg/providers/codex_provider_test.go b/pkg/providers/codex_provider_test.go
index dd5ad2637..3a0da5e3b 100644
--- a/pkg/providers/codex_provider_test.go
+++ b/pkg/providers/codex_provider_test.go
@@ -355,7 +355,9 @@ func TestCodexProvider_ChatRoundTrip(t *testing.T) {
provider.client = createOpenAITestClient(server.URL, "test-token", "acc-123")
messages := []Message{{Role: "user", Content: "Hello"}}
- resp, err := provider.Chat(t.Context(), messages, nil, "gpt-4o", map[string]any{"max_tokens": 1024})
+ // Pass native_search so Codex injects built-in web search (mirrors agent loop when prefer_native is true).
+ opts := map[string]any{"max_tokens": 1024, "native_search": true}
+ resp, err := provider.Chat(t.Context(), messages, nil, "gpt-4o", opts)
if err != nil {
t.Fatalf("Chat() error: %v", err)
}
diff --git a/pkg/providers/common/common.go b/pkg/providers/common/common.go
new file mode 100644
index 000000000..9dfd7dc1d
--- /dev/null
+++ b/pkg/providers/common/common.go
@@ -0,0 +1,389 @@
+// PicoClaw - Ultra-lightweight personal AI agent
+// License: MIT
+//
+// Copyright (c) 2026 PicoClaw contributors
+
+// Package common provides shared utilities used by multiple LLM provider
+// implementations (openai_compat, azure, etc.).
+package common
+
+import (
+ "bufio"
+ "bytes"
+ "encoding/json"
+ "fmt"
+ "io"
+ "log"
+ "net/http"
+ "net/url"
+ "strings"
+ "time"
+
+ "github.com/sipeed/picoclaw/pkg/providers/protocoltypes"
+)
+
+// Re-export protocol types used across providers.
+type (
+ ToolCall = protocoltypes.ToolCall
+ FunctionCall = protocoltypes.FunctionCall
+ LLMResponse = protocoltypes.LLMResponse
+ UsageInfo = protocoltypes.UsageInfo
+ Message = protocoltypes.Message
+ ToolDefinition = protocoltypes.ToolDefinition
+ ToolFunctionDefinition = protocoltypes.ToolFunctionDefinition
+ ExtraContent = protocoltypes.ExtraContent
+ GoogleExtra = protocoltypes.GoogleExtra
+ ReasoningDetail = protocoltypes.ReasoningDetail
+)
+
+const DefaultRequestTimeout = 120 * time.Second
+
+// NewHTTPClient creates an *http.Client with an optional proxy and the default timeout.
+func NewHTTPClient(proxy string) *http.Client {
+ client := &http.Client{
+ Timeout: DefaultRequestTimeout,
+ }
+ if proxy != "" {
+ parsed, err := url.Parse(proxy)
+ if err == nil {
+ // Preserve http.DefaultTransport settings (TLS, HTTP/2, timeouts, etc.)
+ if base, ok := http.DefaultTransport.(*http.Transport); ok {
+ tr := base.Clone()
+ tr.Proxy = http.ProxyURL(parsed)
+ client.Transport = tr
+ } else {
+ // Fallback: minimal transport if DefaultTransport is not *http.Transport.
+ client.Transport = &http.Transport{
+ Proxy: http.ProxyURL(parsed),
+ }
+ }
+ } else {
+ log.Printf("common: invalid proxy URL %q: %v", proxy, err)
+ }
+ }
+ return client
+}
+
+// --- Message serialization ---
+
+// openaiMessage is the wire-format message for OpenAI-compatible APIs.
+// It mirrors protocoltypes.Message but omits SystemParts, which is an
+// internal field that would be unknown to third-party endpoints.
+type openaiMessage struct {
+ Role string `json:"role"`
+ Content string `json:"content"`
+ ReasoningContent string `json:"reasoning_content,omitempty"`
+ ToolCalls []ToolCall `json:"tool_calls,omitempty"`
+ ToolCallID string `json:"tool_call_id,omitempty"`
+}
+
+// SerializeMessages converts internal Message structs to the OpenAI wire format.
+// - Strips SystemParts (unknown to third-party endpoints)
+// - Converts messages with Media to multipart content format (text + image_url parts)
+// - Preserves ToolCallID, ToolCalls, and ReasoningContent for all messages
+func SerializeMessages(messages []Message) []any {
+ out := make([]any, 0, len(messages))
+ for _, m := range messages {
+ if len(m.Media) == 0 {
+ out = append(out, openaiMessage{
+ Role: m.Role,
+ Content: m.Content,
+ ReasoningContent: m.ReasoningContent,
+ ToolCalls: m.ToolCalls,
+ ToolCallID: m.ToolCallID,
+ })
+ continue
+ }
+
+ // Multipart content format for messages with media
+ parts := make([]map[string]any, 0, 1+len(m.Media))
+ if m.Content != "" {
+ parts = append(parts, map[string]any{
+ "type": "text",
+ "text": m.Content,
+ })
+ }
+ for _, mediaURL := range m.Media {
+ if strings.HasPrefix(mediaURL, "data:image/") {
+ parts = append(parts, map[string]any{
+ "type": "image_url",
+ "image_url": map[string]any{
+ "url": mediaURL,
+ },
+ })
+ }
+ }
+
+ msg := map[string]any{
+ "role": m.Role,
+ "content": parts,
+ }
+ if m.ToolCallID != "" {
+ msg["tool_call_id"] = m.ToolCallID
+ }
+ if len(m.ToolCalls) > 0 {
+ msg["tool_calls"] = m.ToolCalls
+ }
+ if m.ReasoningContent != "" {
+ msg["reasoning_content"] = m.ReasoningContent
+ }
+ out = append(out, msg)
+ }
+ return out
+}
+
+// --- Response parsing ---
+
+// ParseResponse parses a JSON chat completion response body into an LLMResponse.
+func ParseResponse(body io.Reader) (*LLMResponse, error) {
+ var apiResponse struct {
+ Choices []struct {
+ Message struct {
+ Content string `json:"content"`
+ ReasoningContent string `json:"reasoning_content"`
+ Reasoning string `json:"reasoning"`
+ ReasoningDetails []ReasoningDetail `json:"reasoning_details"`
+ ToolCalls []struct {
+ ID string `json:"id"`
+ Type string `json:"type"`
+ Function *struct {
+ Name string `json:"name"`
+ Arguments json.RawMessage `json:"arguments"`
+ } `json:"function"`
+ ExtraContent *struct {
+ Google *struct {
+ ThoughtSignature string `json:"thought_signature"`
+ } `json:"google"`
+ } `json:"extra_content"`
+ } `json:"tool_calls"`
+ } `json:"message"`
+ FinishReason string `json:"finish_reason"`
+ } `json:"choices"`
+ Usage *UsageInfo `json:"usage"`
+ }
+
+ if err := json.NewDecoder(body).Decode(&apiResponse); err != nil {
+ return nil, fmt.Errorf("failed to decode response: %w", err)
+ }
+
+ if len(apiResponse.Choices) == 0 {
+ return &LLMResponse{
+ Content: "",
+ FinishReason: "stop",
+ }, nil
+ }
+
+ choice := apiResponse.Choices[0]
+ toolCalls := make([]ToolCall, 0, len(choice.Message.ToolCalls))
+ for _, tc := range choice.Message.ToolCalls {
+ arguments := make(map[string]any)
+ name := ""
+
+ // Extract thought_signature from Gemini/Google-specific extra content
+ thoughtSignature := ""
+ if tc.ExtraContent != nil && tc.ExtraContent.Google != nil {
+ thoughtSignature = tc.ExtraContent.Google.ThoughtSignature
+ }
+
+ if tc.Function != nil {
+ name = tc.Function.Name
+ arguments = DecodeToolCallArguments(tc.Function.Arguments, name)
+ }
+
+ toolCall := ToolCall{
+ ID: tc.ID,
+ Name: name,
+ Arguments: arguments,
+ ThoughtSignature: thoughtSignature,
+ }
+
+ if thoughtSignature != "" {
+ toolCall.ExtraContent = &ExtraContent{
+ Google: &GoogleExtra{
+ ThoughtSignature: thoughtSignature,
+ },
+ }
+ }
+
+ toolCalls = append(toolCalls, toolCall)
+ }
+
+ return &LLMResponse{
+ Content: choice.Message.Content,
+ ReasoningContent: choice.Message.ReasoningContent,
+ Reasoning: choice.Message.Reasoning,
+ ReasoningDetails: choice.Message.ReasoningDetails,
+ ToolCalls: toolCalls,
+ FinishReason: normalizeFinishReason(choice.FinishReason),
+ Usage: apiResponse.Usage,
+ }, nil
+}
+
+// normalizeFinishReason normalizes finish_reason values across providers.
+// Converts "length" to "truncated" for consistent handling.
+func normalizeFinishReason(reason string) string {
+ if reason == "length" {
+ return "truncated"
+ }
+ return reason
+}
+
+// DecodeToolCallArguments decodes a tool call's arguments from raw JSON.
+func DecodeToolCallArguments(raw json.RawMessage, name string) map[string]any {
+ arguments := make(map[string]any)
+ raw = bytes.TrimSpace(raw)
+ if len(raw) == 0 || bytes.Equal(raw, []byte("null")) {
+ return arguments
+ }
+
+ var decoded any
+ if err := json.Unmarshal(raw, &decoded); err != nil {
+ log.Printf("common: failed to decode tool call arguments payload for %q: %v", name, err)
+ arguments["raw"] = string(raw)
+ return arguments
+ }
+
+ switch v := decoded.(type) {
+ case string:
+ if strings.TrimSpace(v) == "" {
+ return arguments
+ }
+ if err := json.Unmarshal([]byte(v), &arguments); err != nil {
+ log.Printf("common: failed to decode tool call arguments for %q: %v", name, err)
+ arguments["raw"] = v
+ }
+ return arguments
+ case map[string]any:
+ return v
+ default:
+ log.Printf("common: unsupported tool call arguments type for %q: %T", name, decoded)
+ arguments["raw"] = string(raw)
+ return arguments
+ }
+}
+
+// --- HTTP response helpers ---
+
+// HandleErrorResponse reads a non-200 response body and returns an appropriate error.
+func HandleErrorResponse(resp *http.Response, apiBase string) error {
+ contentType := resp.Header.Get("Content-Type")
+ body, readErr := io.ReadAll(io.LimitReader(resp.Body, 256))
+ if readErr != nil {
+ return fmt.Errorf("failed to read response: %w", readErr)
+ }
+ if LooksLikeHTML(body, contentType) {
+ return WrapHTMLResponseError(resp.StatusCode, body, contentType, apiBase)
+ }
+ return fmt.Errorf(
+ "API request failed:\n Status: %d\n Body: %s",
+ resp.StatusCode,
+ ResponsePreview(body, 128),
+ )
+}
+
+// ReadAndParseResponse peeks at the response body to detect HTML errors,
+// then parses the JSON response into an LLMResponse.
+func ReadAndParseResponse(resp *http.Response, apiBase string) (*LLMResponse, error) {
+ contentType := resp.Header.Get("Content-Type")
+ reader := bufio.NewReader(resp.Body)
+ prefix, err := reader.Peek(256)
+ if err != nil && err != io.EOF && err != bufio.ErrBufferFull {
+ return nil, fmt.Errorf("failed to inspect response: %w", err)
+ }
+ if LooksLikeHTML(prefix, contentType) {
+ return nil, WrapHTMLResponseError(resp.StatusCode, prefix, contentType, apiBase)
+ }
+ out, err := ParseResponse(reader)
+ if err != nil {
+ return nil, fmt.Errorf("failed to parse JSON response: %w", err)
+ }
+ return out, nil
+}
+
+// LooksLikeHTML checks if the response body appears to be HTML.
+func LooksLikeHTML(body []byte, contentType string) bool {
+ contentType = strings.ToLower(strings.TrimSpace(contentType))
+ if strings.Contains(contentType, "text/html") || strings.Contains(contentType, "application/xhtml+xml") {
+ return true
+ }
+ prefix := bytes.ToLower(leadingTrimmedPrefix(body, 128))
+ return bytes.HasPrefix(prefix, []byte(""
+ }
+ if len(trimmed) <= maxLen {
+ return string(trimmed)
+ }
+ return string(trimmed[:maxLen]) + "..."
+}
+
+func leadingTrimmedPrefix(body []byte, maxLen int) []byte {
+ i := 0
+ for i < len(body) {
+ switch body[i] {
+ case ' ', '\t', '\n', '\r', '\f', '\v':
+ i++
+ default:
+ end := i + maxLen
+ if end > len(body) {
+ end = len(body)
+ }
+ return body[i:end]
+ }
+ }
+ return nil
+}
+
+// --- Numeric helpers ---
+
+// AsInt converts various numeric types to int.
+func AsInt(v any) (int, bool) {
+ switch val := v.(type) {
+ case int:
+ return val, true
+ case int64:
+ return int(val), true
+ case float64:
+ return int(val), true
+ case float32:
+ return int(val), true
+ default:
+ return 0, false
+ }
+}
+
+// AsFloat converts various numeric types to float64.
+func AsFloat(v any) (float64, bool) {
+ switch val := v.(type) {
+ case float64:
+ return val, true
+ case float32:
+ return float64(val), true
+ case int:
+ return float64(val), true
+ case int64:
+ return float64(val), true
+ default:
+ return 0, false
+ }
+}
diff --git a/pkg/providers/common/common_test.go b/pkg/providers/common/common_test.go
new file mode 100644
index 000000000..bb7e7434d
--- /dev/null
+++ b/pkg/providers/common/common_test.go
@@ -0,0 +1,558 @@
+package common
+
+import (
+ "encoding/json"
+ "net/http"
+ "net/http/httptest"
+ "net/url"
+ "strings"
+ "testing"
+
+ "github.com/sipeed/picoclaw/pkg/providers/protocoltypes"
+)
+
+// --- NewHTTPClient tests ---
+
+func TestNewHTTPClient_DefaultTimeout(t *testing.T) {
+ client := NewHTTPClient("")
+ if client.Timeout != DefaultRequestTimeout {
+ t.Errorf("timeout = %v, want %v", client.Timeout, DefaultRequestTimeout)
+ }
+}
+
+func TestNewHTTPClient_WithProxy(t *testing.T) {
+ client := NewHTTPClient("http://127.0.0.1:8080")
+ transport, ok := client.Transport.(*http.Transport)
+ if !ok || transport == nil {
+ t.Fatalf("expected http.Transport with proxy, got %T", client.Transport)
+ }
+ req := &http.Request{URL: &url.URL{Scheme: "https", Host: "api.example.com"}}
+ gotProxy, err := transport.Proxy(req)
+ if err != nil {
+ t.Fatalf("proxy function error: %v", err)
+ }
+ if gotProxy == nil || gotProxy.String() != "http://127.0.0.1:8080" {
+ t.Errorf("proxy = %v, want http://127.0.0.1:8080", gotProxy)
+ }
+}
+
+func TestNewHTTPClient_NoProxy(t *testing.T) {
+ client := NewHTTPClient("")
+ if client.Transport != nil {
+ t.Errorf("expected nil transport without proxy, got %T", client.Transport)
+ }
+}
+
+func TestNewHTTPClient_InvalidProxy(t *testing.T) {
+ // Should not panic, just log and return client without proxy
+ client := NewHTTPClient("://bad-url")
+ if client == nil {
+ t.Fatal("expected non-nil client even with invalid proxy")
+ }
+}
+
+// --- SerializeMessages tests ---
+
+func TestSerializeMessages_PlainText(t *testing.T) {
+ messages := []Message{
+ {Role: "user", Content: "hello"},
+ {Role: "assistant", Content: "hi", ReasoningContent: "thinking..."},
+ }
+ result := SerializeMessages(messages)
+
+ data, _ := json.Marshal(result)
+ var msgs []map[string]any
+ json.Unmarshal(data, &msgs)
+
+ if msgs[0]["content"] != "hello" {
+ t.Errorf("expected plain string content, got %v", msgs[0]["content"])
+ }
+ if msgs[1]["reasoning_content"] != "thinking..." {
+ t.Errorf("reasoning_content not preserved, got %v", msgs[1]["reasoning_content"])
+ }
+}
+
+func TestSerializeMessages_WithMedia(t *testing.T) {
+ messages := []Message{
+ {Role: "user", Content: "describe this", Media: []string{"data:image/png;base64,abc123"}},
+ }
+ result := SerializeMessages(messages)
+
+ data, _ := json.Marshal(result)
+ var msgs []map[string]any
+ json.Unmarshal(data, &msgs)
+
+ content, ok := msgs[0]["content"].([]any)
+ if !ok {
+ t.Fatalf("expected array content for media message, got %T", msgs[0]["content"])
+ }
+ if len(content) != 2 {
+ t.Fatalf("expected 2 content parts, got %d", len(content))
+ }
+}
+
+func TestSerializeMessages_MediaWithToolCallID(t *testing.T) {
+ messages := []Message{
+ {Role: "tool", Content: "result", Media: []string{"data:image/png;base64,xyz"}, ToolCallID: "call_1"},
+ }
+ result := SerializeMessages(messages)
+
+ data, _ := json.Marshal(result)
+ var msgs []map[string]any
+ json.Unmarshal(data, &msgs)
+
+ if msgs[0]["tool_call_id"] != "call_1" {
+ t.Errorf("tool_call_id not preserved, got %v", msgs[0]["tool_call_id"])
+ }
+}
+
+func TestSerializeMessages_StripsSystemParts(t *testing.T) {
+ messages := []Message{
+ {
+ Role: "system",
+ Content: "you are helpful",
+ SystemParts: []protocoltypes.ContentBlock{
+ {Type: "text", Text: "you are helpful"},
+ },
+ },
+ }
+ result := SerializeMessages(messages)
+
+ data, _ := json.Marshal(result)
+ if strings.Contains(string(data), "system_parts") {
+ t.Error("system_parts should not appear in serialized output")
+ }
+}
+
+// --- ParseResponse tests ---
+
+func TestParseResponse_BasicContent(t *testing.T) {
+ body := `{"choices":[{"message":{"content":"hello world"},"finish_reason":"stop"}]}`
+ out, err := ParseResponse(strings.NewReader(body))
+ if err != nil {
+ t.Fatalf("ParseResponse() error = %v", err)
+ }
+ if out.Content != "hello world" {
+ t.Errorf("Content = %q, want %q", out.Content, "hello world")
+ }
+ if out.FinishReason != "stop" {
+ t.Errorf("FinishReason = %q, want %q", out.FinishReason, "stop")
+ }
+}
+
+func TestParseResponse_EmptyChoices(t *testing.T) {
+ body := `{"choices":[]}`
+ out, err := ParseResponse(strings.NewReader(body))
+ if err != nil {
+ t.Fatalf("ParseResponse() error = %v", err)
+ }
+ if out.Content != "" {
+ t.Errorf("Content = %q, want empty", out.Content)
+ }
+ if out.FinishReason != "stop" {
+ t.Errorf("FinishReason = %q, want %q", out.FinishReason, "stop")
+ }
+}
+
+func TestParseResponse_WithToolCalls(t *testing.T) {
+ body := `{"choices":[{"message":{"content":"","tool_calls":[{"id":"call_1","type":"function","function":{"name":"get_weather","arguments":"{\"city\":\"SF\"}"}}]},"finish_reason":"tool_calls"}]}`
+ out, err := ParseResponse(strings.NewReader(body))
+ if err != nil {
+ t.Fatalf("ParseResponse() error = %v", err)
+ }
+ if len(out.ToolCalls) != 1 {
+ t.Fatalf("len(ToolCalls) = %d, want 1", len(out.ToolCalls))
+ }
+ if out.ToolCalls[0].Name != "get_weather" {
+ t.Errorf("ToolCalls[0].Name = %q, want %q", out.ToolCalls[0].Name, "get_weather")
+ }
+ if out.ToolCalls[0].Arguments["city"] != "SF" {
+ t.Errorf("ToolCalls[0].Arguments[city] = %v, want SF", out.ToolCalls[0].Arguments["city"])
+ }
+}
+
+func TestParseResponse_WithUsage(t *testing.T) {
+ body := `{"choices":[{"message":{"content":"ok"},"finish_reason":"stop"}],"usage":{"prompt_tokens":10,"completion_tokens":5,"total_tokens":15}}`
+ out, err := ParseResponse(strings.NewReader(body))
+ if err != nil {
+ t.Fatalf("ParseResponse() error = %v", err)
+ }
+ if out.Usage == nil {
+ t.Fatal("Usage is nil")
+ }
+ if out.Usage.PromptTokens != 10 {
+ t.Errorf("PromptTokens = %d, want 10", out.Usage.PromptTokens)
+ }
+}
+
+func TestParseResponse_WithReasoningContent(t *testing.T) {
+ body := `{"choices":[{"message":{"content":"2","reasoning_content":"Let me think... 1+1=2"},"finish_reason":"stop"}]}`
+ out, err := ParseResponse(strings.NewReader(body))
+ if err != nil {
+ t.Fatalf("ParseResponse() error = %v", err)
+ }
+ if out.ReasoningContent != "Let me think... 1+1=2" {
+ t.Errorf("ReasoningContent = %q, want %q", out.ReasoningContent, "Let me think... 1+1=2")
+ }
+}
+
+func TestParseResponse_InvalidJSON(t *testing.T) {
+ _, err := ParseResponse(strings.NewReader("not json"))
+ if err == nil {
+ t.Fatal("expected error for invalid JSON")
+ }
+}
+
+// --- DecodeToolCallArguments tests ---
+
+func TestDecodeToolCallArguments_ObjectJSON(t *testing.T) {
+ raw := json.RawMessage(`{"city":"Seattle","units":"metric"}`)
+ args := DecodeToolCallArguments(raw, "test")
+ if args["city"] != "Seattle" {
+ t.Errorf("city = %v, want Seattle", args["city"])
+ }
+ if args["units"] != "metric" {
+ t.Errorf("units = %v, want metric", args["units"])
+ }
+}
+
+func TestDecodeToolCallArguments_StringJSON(t *testing.T) {
+ raw := json.RawMessage(`"{\"city\":\"SF\"}"`)
+ args := DecodeToolCallArguments(raw, "test")
+ if args["city"] != "SF" {
+ t.Errorf("city = %v, want SF", args["city"])
+ }
+}
+
+func TestDecodeToolCallArguments_EmptyInput(t *testing.T) {
+ args := DecodeToolCallArguments(nil, "test")
+ if len(args) != 0 {
+ t.Errorf("expected empty map, got %v", args)
+ }
+}
+
+func TestDecodeToolCallArguments_NullInput(t *testing.T) {
+ args := DecodeToolCallArguments(json.RawMessage(`null`), "test")
+ if len(args) != 0 {
+ t.Errorf("expected empty map, got %v", args)
+ }
+}
+
+func TestDecodeToolCallArguments_InvalidJSON(t *testing.T) {
+ args := DecodeToolCallArguments(json.RawMessage(`not-json`), "test")
+ if _, ok := args["raw"]; !ok {
+ t.Error("expected 'raw' fallback key for invalid JSON")
+ }
+}
+
+func TestDecodeToolCallArguments_EmptyStringJSON(t *testing.T) {
+ args := DecodeToolCallArguments(json.RawMessage(`" "`), "test")
+ if len(args) != 0 {
+ t.Errorf("expected empty map for whitespace string, got %v", args)
+ }
+}
+
+// --- HandleErrorResponse tests ---
+
+func TestHandleErrorResponse_JSONError(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "application/json")
+ w.WriteHeader(http.StatusBadRequest)
+ w.Write([]byte(`{"error":"bad request"}`))
+ }))
+ defer server.Close()
+
+ resp, err := http.Get(server.URL)
+ if err != nil {
+ t.Fatalf("http.Get() error = %v", err)
+ }
+ defer resp.Body.Close()
+ err = HandleErrorResponse(resp, server.URL)
+ if err == nil {
+ t.Fatal("expected error")
+ }
+ if !strings.Contains(err.Error(), "400") {
+ t.Errorf("error should contain status code, got %v", err)
+ }
+ if strings.Contains(err.Error(), "HTML") {
+ t.Errorf("should not mention HTML for JSON error, got %v", err)
+ }
+}
+
+func TestHandleErrorResponse_HTMLError(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "text/html")
+ w.WriteHeader(http.StatusBadGateway)
+ w.Write([]byte("bad gateway"))
+ }))
+ defer server.Close()
+
+ resp, err := http.Get(server.URL)
+ if err != nil {
+ t.Fatalf("http.Get() error = %v", err)
+ }
+ defer resp.Body.Close()
+ err = HandleErrorResponse(resp, server.URL)
+ if err == nil {
+ t.Fatal("expected error")
+ }
+ if !strings.Contains(err.Error(), "HTML instead of JSON") {
+ t.Errorf("expected HTML error message, got %v", err)
+ }
+}
+
+// --- ReadAndParseResponse tests ---
+
+func TestReadAndParseResponse_ValidJSON(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "application/json")
+ w.Write([]byte(`{"choices":[{"message":{"content":"ok"},"finish_reason":"stop"}]}`))
+ }))
+ defer server.Close()
+
+ resp, err := http.Get(server.URL)
+ if err != nil {
+ t.Fatalf("http.Get() error = %v", err)
+ }
+ defer resp.Body.Close()
+ out, err := ReadAndParseResponse(resp, server.URL)
+ if err != nil {
+ t.Fatalf("ReadAndParseResponse() error = %v", err)
+ }
+ if out.Content != "ok" {
+ t.Errorf("Content = %q, want %q", out.Content, "ok")
+ }
+}
+
+func TestReadAndParseResponse_HTMLResponse(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "text/html")
+ w.Write([]byte("login page"))
+ }))
+ defer server.Close()
+
+ resp, err := http.Get(server.URL)
+ if err != nil {
+ t.Fatalf("http.Get() error = %v", err)
+ }
+ defer resp.Body.Close()
+ _, err = ReadAndParseResponse(resp, server.URL)
+ if err == nil {
+ t.Fatal("expected error for HTML response")
+ }
+ if !strings.Contains(err.Error(), "HTML instead of JSON") {
+ t.Errorf("expected HTML error, got %v", err)
+ }
+}
+
+// --- LooksLikeHTML tests ---
+
+func TestLooksLikeHTML_ContentTypeHTML(t *testing.T) {
+ if !LooksLikeHTML(nil, "text/html; charset=utf-8") {
+ t.Error("expected true for text/html content type")
+ }
+}
+
+func TestLooksLikeHTML_ContentTypeXHTML(t *testing.T) {
+ if !LooksLikeHTML(nil, "application/xhtml+xml") {
+ t.Error("expected true for xhtml content type")
+ }
+}
+
+func TestLooksLikeHTML_BodyPrefix(t *testing.T) {
+ tests := []struct {
+ name string
+ body string
+ }{
+ {"doctype", ""},
+ {"html tag", ""},
+ {"head tag", ""},
+ {"body tag", "content"},
+ {"whitespace before", " \n\t"},
+ }
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ if !LooksLikeHTML([]byte(tt.body), "application/json") {
+ t.Errorf("expected true for body %q", tt.body)
+ }
+ })
+ }
+}
+
+func TestLooksLikeHTML_NotHTML(t *testing.T) {
+ if LooksLikeHTML([]byte(`{"error":"bad"}`), "application/json") {
+ t.Error("expected false for JSON body")
+ }
+}
+
+// --- ResponsePreview tests ---
+
+func TestResponsePreview_Short(t *testing.T) {
+ got := ResponsePreview([]byte("hello"), 128)
+ if got != "hello" {
+ t.Errorf("got %q, want %q", got, "hello")
+ }
+}
+
+func TestResponsePreview_Truncated(t *testing.T) {
+ body := strings.Repeat("a", 200)
+ got := ResponsePreview([]byte(body), 128)
+ if len(got) != 131 { // 128 + "..."
+ t.Errorf("len = %d, want 131", len(got))
+ }
+ if !strings.HasSuffix(got, "...") {
+ t.Error("expected ... suffix")
+ }
+}
+
+func TestResponsePreview_Empty(t *testing.T) {
+ got := ResponsePreview([]byte(""), 128)
+ if got != "" {
+ t.Errorf("got %q, want %q", got, "")
+ }
+}
+
+func TestResponsePreview_Whitespace(t *testing.T) {
+ got := ResponsePreview([]byte(" \n\t "), 128)
+ if got != "" {
+ t.Errorf("got %q, want %q for whitespace-only body", got, "")
+ }
+}
+
+// --- AsInt tests ---
+
+func TestAsInt(t *testing.T) {
+ tests := []struct {
+ name string
+ val any
+ want int
+ ok bool
+ }{
+ {"int", 42, 42, true},
+ {"int64", int64(99), 99, true},
+ {"float64", float64(512), 512, true},
+ {"float32", float32(256), 256, true},
+ {"string", "nope", 0, false},
+ {"nil", nil, 0, false},
+ }
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ got, ok := AsInt(tt.val)
+ if ok != tt.ok || got != tt.want {
+ t.Errorf("AsInt(%v) = (%d, %v), want (%d, %v)", tt.val, got, ok, tt.want, tt.ok)
+ }
+ })
+ }
+}
+
+// --- AsFloat tests ---
+
+func TestAsFloat(t *testing.T) {
+ tests := []struct {
+ name string
+ val any
+ want float64
+ ok bool
+ }{
+ {"float64", float64(0.7), 0.7, true},
+ {"float32", float32(0.5), float64(float32(0.5)), true},
+ {"int", 1, 1.0, true},
+ {"int64", int64(100), 100.0, true},
+ {"string", "nope", 0, false},
+ {"nil", nil, 0, false},
+ }
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ got, ok := AsFloat(tt.val)
+ if ok != tt.ok || got != tt.want {
+ t.Errorf("AsFloat(%v) = (%f, %v), want (%f, %v)", tt.val, got, ok, tt.want, tt.ok)
+ }
+ })
+ }
+}
+
+// --- WrapHTMLResponseError tests ---
+
+func TestWrapHTMLResponseError(t *testing.T) {
+ err := WrapHTMLResponseError(502, []byte("bad"), "text/html", "https://api.example.com")
+ if err == nil {
+ t.Fatal("expected error")
+ }
+ msg := err.Error()
+ if !strings.Contains(msg, "502") {
+ t.Errorf("expected status code in error, got %v", msg)
+ }
+ if !strings.Contains(msg, "https://api.example.com") {
+ t.Errorf("expected api base in error, got %v", msg)
+ }
+ if !strings.Contains(msg, "HTML instead of JSON") {
+ t.Errorf("expected HTML mention in error, got %v", msg)
+ }
+}
+
+// --- HandleErrorResponse with read failure ---
+
+func TestHandleErrorResponse_EmptyBody(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "application/json")
+ w.WriteHeader(http.StatusInternalServerError)
+ // empty body
+ }))
+ defer server.Close()
+
+ resp, err := http.Get(server.URL)
+ if err != nil {
+ t.Fatalf("http.Get() error = %v", err)
+ }
+ defer resp.Body.Close()
+ err = HandleErrorResponse(resp, server.URL)
+ if err == nil {
+ t.Fatal("expected error")
+ }
+ if !strings.Contains(err.Error(), "500") {
+ t.Errorf("expected status code, got %v", err)
+ }
+}
+
+// --- ReadAndParseResponse with invalid JSON ---
+
+func TestReadAndParseResponse_InvalidJSON(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "application/json")
+ w.Write([]byte("not valid json"))
+ }))
+ defer server.Close()
+
+ resp, err := http.Get(server.URL)
+ if err != nil {
+ t.Fatalf("http.Get() error = %v", err)
+ }
+ defer resp.Body.Close()
+ _, err = ReadAndParseResponse(resp, server.URL)
+ if err == nil {
+ t.Fatal("expected error for invalid JSON")
+ }
+}
+
+// --- ParseResponse with thought_signature (Google/Gemini) ---
+
+func TestParseResponse_WithThoughtSignature(t *testing.T) {
+ body := `{"choices":[{"message":{"content":"","tool_calls":[{"id":"call_1","type":"function","function":{"name":"test_tool","arguments":"{}"},"extra_content":{"google":{"thought_signature":"sig123"}}}]},"finish_reason":"tool_calls"}]}`
+ out, err := ParseResponse(strings.NewReader(body))
+ if err != nil {
+ t.Fatalf("ParseResponse() error = %v", err)
+ }
+ if len(out.ToolCalls) != 1 {
+ t.Fatalf("len(ToolCalls) = %d, want 1", len(out.ToolCalls))
+ }
+ if out.ToolCalls[0].ThoughtSignature != "sig123" {
+ t.Errorf("ThoughtSignature = %q, want %q", out.ToolCalls[0].ThoughtSignature, "sig123")
+ }
+ if out.ToolCalls[0].ExtraContent == nil || out.ToolCalls[0].ExtraContent.Google == nil {
+ t.Fatal("ExtraContent.Google is nil")
+ }
+ if out.ToolCalls[0].ExtraContent.Google.ThoughtSignature != "sig123" {
+ t.Errorf("ExtraContent.Google.ThoughtSignature = %q, want %q",
+ out.ToolCalls[0].ExtraContent.Google.ThoughtSignature, "sig123")
+ }
+}
diff --git a/pkg/providers/factory_provider.go b/pkg/providers/factory_provider.go
index e99e07bc2..a7fef8f5b 100644
--- a/pkg/providers/factory_provider.go
+++ b/pkg/providers/factory_provider.go
@@ -11,6 +11,7 @@ import (
"github.com/sipeed/picoclaw/pkg/config"
anthropicmessages "github.com/sipeed/picoclaw/pkg/providers/anthropic_messages"
+ "github.com/sipeed/picoclaw/pkg/providers/azure"
)
// createClaudeAuthProvider creates a Claude provider using OAuth credentials from auth store.
@@ -54,8 +55,8 @@ func ExtractProtocol(model string) (protocol, modelID string) {
// CreateProviderFromConfig creates a provider based on the ModelConfig.
// It uses the protocol prefix in the Model field to determine which provider to create.
-// Supported protocols: openai, litellm, anthropic, anthropic-messages, antigravity,
-// claude-cli, codex-cli, github-copilot
+// Supported protocols: openai, litellm, novita, anthropic, anthropic-messages,
+// antigravity, claude-cli, codex-cli, github-copilot
// Returns the provider, the model ID (without protocol prefix), and any error.
func CreateProviderFromConfig(cfg *config.ModelConfig) (LLMProvider, string, error) {
if cfg == nil {
@@ -94,10 +95,29 @@ func CreateProviderFromConfig(cfg *config.ModelConfig) (LLMProvider, string, err
cfg.RequestTimeout,
), modelID, nil
+ case "azure", "azure-openai":
+ // Azure OpenAI uses deployment-based URLs, api-key header auth,
+ // and always sends max_completion_tokens.
+ if cfg.APIKey == "" {
+ return nil, "", fmt.Errorf("api_key is required for azure protocol")
+ }
+ if cfg.APIBase == "" {
+ return nil, "", fmt.Errorf(
+ "api_base is required for azure protocol (e.g., https://your-resource.openai.azure.com)",
+ )
+ }
+ return azure.NewProviderWithTimeout(
+ cfg.APIKey,
+ cfg.APIBase,
+ cfg.Proxy,
+ cfg.RequestTimeout,
+ ), modelID, nil
+
case "litellm", "openrouter", "groq", "zhipu", "gemini", "nvidia",
"ollama", "moonshot", "shengsuanyun", "deepseek", "cerebras",
- "vivgrid", "volcengine", "vllm", "qwen", "mistral", "avian",
- "minimax", "longcat", "modelscope":
+ "vivgrid", "volcengine", "vllm", "qwen", "qwen-intl", "qwen-international", "dashscope-intl",
+ "qwen-us", "dashscope-us", "mistral", "avian", "minimax", "longcat", "modelscope", "novita",
+ "coding-plan", "alibaba-coding", "qwen-coding":
// All other OpenAI-compatible HTTP providers
if cfg.APIKey == "" && cfg.APIBase == "" {
return nil, "", fmt.Errorf("api_key or api_base is required for HTTP-based protocol %q", protocol)
@@ -154,6 +174,21 @@ func CreateProviderFromConfig(cfg *config.ModelConfig) (LLMProvider, string, err
cfg.RequestTimeout,
), modelID, nil
+ case "coding-plan-anthropic", "alibaba-coding-anthropic":
+ // Alibaba Coding Plan with Anthropic-compatible API
+ apiBase := cfg.APIBase
+ if apiBase == "" {
+ apiBase = getDefaultAPIBase(protocol)
+ }
+ if cfg.APIKey == "" {
+ return nil, "", fmt.Errorf("api_key is required for %q protocol (model: %s)", protocol, cfg.Model)
+ }
+ return anthropicmessages.NewProviderWithTimeout(
+ cfg.APIKey,
+ apiBase,
+ cfg.RequestTimeout,
+ ), modelID, nil
+
case "antigravity":
return NewAntigravityProvider(), modelID, nil
@@ -200,6 +235,8 @@ func getDefaultAPIBase(protocol string) string {
return "https://openrouter.ai/api/v1"
case "litellm":
return "http://localhost:4000/v1"
+ case "novita":
+ return "https://api.novita.ai/openai"
case "groq":
return "https://api.groq.com/openai/v1"
case "zhipu":
@@ -224,6 +261,14 @@ func getDefaultAPIBase(protocol string) string {
return "https://ark.cn-beijing.volces.com/api/v3"
case "qwen":
return "https://dashscope.aliyuncs.com/compatible-mode/v1"
+ case "qwen-intl", "qwen-international", "dashscope-intl":
+ return "https://dashscope-intl.aliyuncs.com/compatible-mode/v1"
+ case "qwen-us", "dashscope-us":
+ return "https://dashscope-us.aliyuncs.com/compatible-mode/v1"
+ case "coding-plan", "alibaba-coding", "qwen-coding":
+ return "https://coding-intl.dashscope.aliyuncs.com/v1"
+ case "coding-plan-anthropic", "alibaba-coding-anthropic":
+ return "https://coding-intl.dashscope.aliyuncs.com/apps/anthropic"
case "vllm":
return "http://localhost:8000/v1"
case "mistral":
diff --git a/pkg/providers/factory_provider_test.go b/pkg/providers/factory_provider_test.go
index 00676ebf9..8b9ddeecd 100644
--- a/pkg/providers/factory_provider_test.go
+++ b/pkg/providers/factory_provider_test.go
@@ -64,6 +64,12 @@ func TestExtractProtocol(t *testing.T) {
wantProtocol: "nvidia",
wantModelID: "meta/llama-3.1-8b",
},
+ {
+ name: "azure with prefix",
+ model: "azure/my-gpt5-deployment",
+ wantProtocol: "azure",
+ wantModelID: "my-gpt5-deployment",
+ },
}
for _, tt := range tests {
@@ -106,6 +112,7 @@ func TestCreateProviderFromConfig_DefaultAPIBase(t *testing.T) {
}{
{"openai", "openai"},
{"groq", "groq"},
+ {"novita", "novita"},
{"openrouter", "openrouter"},
{"cerebras", "cerebras"},
{"vivgrid", "vivgrid"},
@@ -216,6 +223,34 @@ func TestGetDefaultAPIBase_ModelScope(t *testing.T) {
}
}
+func TestCreateProviderFromConfig_Novita(t *testing.T) {
+ cfg := &config.ModelConfig{
+ ModelName: "test-novita",
+ Model: "novita/deepseek/deepseek-v3.2",
+ APIKey: "test-key",
+ }
+
+ provider, modelID, err := CreateProviderFromConfig(cfg)
+ if err != nil {
+ t.Fatalf("CreateProviderFromConfig() error = %v", err)
+ }
+ if provider == nil {
+ t.Fatal("CreateProviderFromConfig() returned nil provider")
+ }
+ if modelID != "deepseek/deepseek-v3.2" {
+ t.Errorf("modelID = %q, want %q", modelID, "deepseek/deepseek-v3.2")
+ }
+ if _, ok := provider.(*HTTPProvider); !ok {
+ t.Fatalf("expected *HTTPProvider, got %T", provider)
+ }
+}
+
+func TestGetDefaultAPIBase_Novita(t *testing.T) {
+ if got := getDefaultAPIBase("novita"); got != "https://api.novita.ai/openai" {
+ t.Fatalf("getDefaultAPIBase(%q) = %q, want %q", "novita", got, "https://api.novita.ai/openai")
+ }
+}
+
func TestCreateProviderFromConfig_Anthropic(t *testing.T) {
cfg := &config.ModelConfig{
ModelName: "test-anthropic",
@@ -371,3 +406,200 @@ func TestCreateProviderFromConfig_RequestTimeoutPropagation(t *testing.T) {
t.Fatalf("Chat() error = %q, want timeout-related error", errMsg)
}
}
+
+func TestCreateProviderFromConfig_Azure(t *testing.T) {
+ cfg := &config.ModelConfig{
+ ModelName: "azure-gpt5",
+ Model: "azure/my-gpt5-deployment",
+ APIKey: "test-azure-key",
+ APIBase: "https://my-resource.openai.azure.com",
+ }
+
+ provider, modelID, err := CreateProviderFromConfig(cfg)
+ if err != nil {
+ t.Fatalf("CreateProviderFromConfig() error = %v", err)
+ }
+ if provider == nil {
+ t.Fatal("CreateProviderFromConfig() returned nil provider")
+ }
+ if modelID != "my-gpt5-deployment" {
+ t.Errorf("modelID = %q, want %q", modelID, "my-gpt5-deployment")
+ }
+}
+
+func TestCreateProviderFromConfig_AzureOpenAIAlias(t *testing.T) {
+ cfg := &config.ModelConfig{
+ ModelName: "azure-gpt4",
+ Model: "azure-openai/my-deployment",
+ APIKey: "test-azure-key",
+ APIBase: "https://my-resource.openai.azure.com",
+ }
+
+ provider, modelID, err := CreateProviderFromConfig(cfg)
+ if err != nil {
+ t.Fatalf("CreateProviderFromConfig() error = %v", err)
+ }
+ if provider == nil {
+ t.Fatal("CreateProviderFromConfig() returned nil provider")
+ }
+ if modelID != "my-deployment" {
+ t.Errorf("modelID = %q, want %q", modelID, "my-deployment")
+ }
+}
+
+func TestCreateProviderFromConfig_AzureMissingAPIKey(t *testing.T) {
+ cfg := &config.ModelConfig{
+ ModelName: "azure-gpt5",
+ Model: "azure/my-gpt5-deployment",
+ APIBase: "https://my-resource.openai.azure.com",
+ }
+
+ _, _, err := CreateProviderFromConfig(cfg)
+ if err == nil {
+ t.Fatal("CreateProviderFromConfig() expected error for missing API key")
+ }
+}
+
+func TestCreateProviderFromConfig_AzureMissingAPIBase(t *testing.T) {
+ cfg := &config.ModelConfig{
+ ModelName: "azure-gpt5",
+ Model: "azure/my-gpt5-deployment",
+ APIKey: "test-azure-key",
+ }
+
+ _, _, err := CreateProviderFromConfig(cfg)
+ if err == nil {
+ t.Fatal("CreateProviderFromConfig() expected error for missing API base")
+ }
+}
+
+func TestCreateProviderFromConfig_QwenInternationalAlias(t *testing.T) {
+ tests := []struct {
+ name string
+ protocol string
+ }{
+ {"qwen-international", "qwen-international"},
+ {"dashscope-intl", "dashscope-intl"},
+ {"qwen-intl", "qwen-intl"},
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ cfg := &config.ModelConfig{
+ ModelName: "test-" + tt.protocol,
+ Model: tt.protocol + "/qwen-max",
+ APIKey: "test-key",
+ }
+
+ provider, modelID, err := CreateProviderFromConfig(cfg)
+ if err != nil {
+ t.Fatalf("CreateProviderFromConfig() error = %v", err)
+ }
+ if provider == nil {
+ t.Fatal("CreateProviderFromConfig() returned nil provider")
+ }
+ if modelID != "qwen-max" {
+ t.Errorf("modelID = %q, want %q", modelID, "qwen-max")
+ }
+ if _, ok := provider.(*HTTPProvider); !ok {
+ t.Fatalf("expected *HTTPProvider, got %T", provider)
+ }
+ })
+ }
+}
+
+func TestCreateProviderFromConfig_QwenUSAlias(t *testing.T) {
+ tests := []struct {
+ name string
+ protocol string
+ }{
+ {"qwen-us", "qwen-us"},
+ {"dashscope-us", "dashscope-us"},
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ cfg := &config.ModelConfig{
+ ModelName: "test-" + tt.protocol,
+ Model: tt.protocol + "/qwen-max",
+ APIKey: "test-key",
+ }
+
+ provider, modelID, err := CreateProviderFromConfig(cfg)
+ if err != nil {
+ t.Fatalf("CreateProviderFromConfig() error = %v", err)
+ }
+ if provider == nil {
+ t.Fatal("CreateProviderFromConfig() returned nil provider")
+ }
+ if modelID != "qwen-max" {
+ t.Errorf("modelID = %q, want %q", modelID, "qwen-max")
+ }
+ if _, ok := provider.(*HTTPProvider); !ok {
+ t.Fatalf("expected *HTTPProvider, got %T", provider)
+ }
+ })
+ }
+}
+
+func TestCreateProviderFromConfig_CodingPlanAnthropic(t *testing.T) {
+ tests := []struct {
+ name string
+ protocol string
+ }{
+ {"coding-plan-anthropic", "coding-plan-anthropic"},
+ {"alibaba-coding-anthropic", "alibaba-coding-anthropic"},
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ cfg := &config.ModelConfig{
+ ModelName: "test-" + tt.protocol,
+ Model: tt.protocol + "/claude-sonnet-4-20250514",
+ APIKey: "test-key",
+ }
+
+ provider, modelID, err := CreateProviderFromConfig(cfg)
+ if err != nil {
+ t.Fatalf("CreateProviderFromConfig() error = %v", err)
+ }
+ if provider == nil {
+ t.Fatal("CreateProviderFromConfig() returned nil provider")
+ }
+ if modelID != "claude-sonnet-4-20250514" {
+ t.Errorf("modelID = %q, want %q", modelID, "claude-sonnet-4-20250514")
+ }
+ // coding-plan-anthropic uses Anthropic Messages provider
+ // Verify it's the anthropic messages provider by checking interface
+ var _ LLMProvider = provider
+ })
+ }
+}
+
+func TestGetDefaultAPIBase_CodingPlanAnthropic(t *testing.T) {
+ expectedURL := "https://coding-intl.dashscope.aliyuncs.com/apps/anthropic"
+ if got := getDefaultAPIBase("coding-plan-anthropic"); got != expectedURL {
+ t.Fatalf("getDefaultAPIBase(%q) = %q, want %q", "coding-plan-anthropic", got, expectedURL)
+ }
+ if got := getDefaultAPIBase("alibaba-coding-anthropic"); got != expectedURL {
+ t.Fatalf("getDefaultAPIBase(%q) = %q, want %q", "alibaba-coding-anthropic", got, expectedURL)
+ }
+}
+
+func TestGetDefaultAPIBase_QwenIntlAliases(t *testing.T) {
+ expectedURL := "https://dashscope-intl.aliyuncs.com/compatible-mode/v1"
+ for _, protocol := range []string{"qwen-intl", "qwen-international", "dashscope-intl"} {
+ if got := getDefaultAPIBase(protocol); got != expectedURL {
+ t.Fatalf("getDefaultAPIBase(%q) = %q, want %q", protocol, got, expectedURL)
+ }
+ }
+}
+
+func TestGetDefaultAPIBase_QwenUSAliases(t *testing.T) {
+ expectedURL := "https://dashscope-us.aliyuncs.com/compatible-mode/v1"
+ for _, protocol := range []string{"qwen-us", "dashscope-us"} {
+ if got := getDefaultAPIBase(protocol); got != expectedURL {
+ t.Fatalf("getDefaultAPIBase(%q) = %q, want %q", protocol, got, expectedURL)
+ }
+ }
+}
diff --git a/pkg/providers/fallback.go b/pkg/providers/fallback.go
index 7ba563b66..549ec7837 100644
--- a/pkg/providers/fallback.go
+++ b/pkg/providers/fallback.go
@@ -117,17 +117,19 @@ func (fc *FallbackChain) Execute(
return nil, context.Canceled
}
- // Check cooldown.
- if !fc.cooldown.IsAvailable(candidate.Provider) {
- remaining := fc.cooldown.CooldownRemaining(candidate.Provider)
+ // Check cooldown (per provider/model, not just provider).
+ // This allows multi-key failover where different keys use different model names.
+ cooldownKey := ModelKey(candidate.Provider, candidate.Model)
+ if !fc.cooldown.IsAvailable(cooldownKey) {
+ remaining := fc.cooldown.CooldownRemaining(cooldownKey)
result.Attempts = append(result.Attempts, FallbackAttempt{
Provider: candidate.Provider,
Model: candidate.Model,
Skipped: true,
Reason: FailoverRateLimit,
Error: fmt.Errorf(
- "provider %s in cooldown (%s remaining)",
- candidate.Provider,
+ "%s in cooldown (%s remaining)",
+ cooldownKey,
remaining.Round(time.Second),
),
})
@@ -141,7 +143,7 @@ func (fc *FallbackChain) Execute(
if err == nil {
// Success.
- fc.cooldown.MarkSuccess(candidate.Provider)
+ fc.cooldown.MarkSuccess(cooldownKey)
result.Response = resp
result.Provider = candidate.Provider
result.Model = candidate.Model
@@ -187,7 +189,7 @@ func (fc *FallbackChain) Execute(
}
// Retriable error: mark failure and continue to next candidate.
- fc.cooldown.MarkFailure(candidate.Provider, failErr.Reason)
+ fc.cooldown.MarkFailure(cooldownKey, failErr.Reason)
result.Attempts = append(result.Attempts, FallbackAttempt{
Provider: candidate.Provider,
Model: candidate.Model,
diff --git a/pkg/providers/fallback_multikey_test.go b/pkg/providers/fallback_multikey_test.go
new file mode 100644
index 000000000..9ed8fa73c
--- /dev/null
+++ b/pkg/providers/fallback_multikey_test.go
@@ -0,0 +1,384 @@
+package providers
+
+import (
+ "context"
+ "errors"
+ "testing"
+)
+
+// TestMultiKeyFailover tests the complete failover flow with multiple API keys.
+// This simulates the config expansion scenario where api_keys: ["key1", "key2", "key3"]
+// is expanded into primary + fallbacks.
+func TestMultiKeyFailover(t *testing.T) {
+ // Simulate expanded config: primary with 2 fallbacks
+ // This is what ExpandMultiKeyModels would produce for api_keys: ["key1", "key2", "key3"]
+ cfg := ModelConfig{
+ Primary: "glm-4.7",
+ Fallbacks: []string{"glm-4.7__key_1", "glm-4.7__key_2"},
+ }
+
+ candidates := ResolveCandidates(cfg, "zhipu")
+
+ if len(candidates) != 3 {
+ t.Fatalf("expected 3 candidates, got %d: %v", len(candidates), candidates)
+ }
+
+ // Create fallback chain
+ cooldown := NewCooldownTracker()
+ chain := NewFallbackChain(cooldown)
+
+ // Mock run function: first call fails with 429, second succeeds
+ callCount := 0
+ mockRun := func(ctx context.Context, provider, model string) (*LLMResponse, error) {
+ callCount++
+ if callCount == 1 {
+ // First call: simulate rate limit
+ return nil, errors.New("http error: status 429 - rate limit exceeded")
+ }
+ // Second call: success
+ return &LLMResponse{
+ Content: "Hello from key2!",
+ }, nil
+ }
+
+ // Execute fallback chain
+ result, err := chain.Execute(context.Background(), candidates, mockRun)
+ if err != nil {
+ t.Fatalf("expected success after failover, got error: %v", err)
+ }
+
+ if result == nil {
+ t.Fatal("expected result, got nil")
+ }
+
+ if result.Response.Content != "Hello from key2!" {
+ t.Errorf("expected response from key2, got: %s", result.Response.Content)
+ }
+
+ if callCount != 2 {
+ t.Errorf("expected 2 calls (1 fail + 1 success), got %d", callCount)
+ }
+
+ // Verify first attempt was recorded
+ if len(result.Attempts) != 1 {
+ t.Errorf("expected 1 failed attempt recorded, got %d", len(result.Attempts))
+ }
+
+ if result.Attempts[0].Reason != FailoverRateLimit {
+ t.Errorf(
+ "expected first attempt reason to be rate_limit, got: %s",
+ result.Attempts[0].Reason,
+ )
+ }
+}
+
+// TestMultiKeyFailoverAllFail tests when all keys hit rate limit
+func TestMultiKeyFailoverAllFail(t *testing.T) {
+ cfg := ModelConfig{
+ Primary: "glm-4.7",
+ Fallbacks: []string{"glm-4.7__key_1", "glm-4.7__key_2"},
+ }
+
+ candidates := ResolveCandidates(cfg, "zhipu")
+
+ cooldown := NewCooldownTracker()
+ chain := NewFallbackChain(cooldown)
+
+ // Mock run function: all calls fail with rate limit
+ callCount := 0
+ mockRun := func(ctx context.Context, provider, model string) (*LLMResponse, error) {
+ callCount++
+ return nil, errors.New("status: 429 - too many requests")
+ }
+
+ // Execute fallback chain
+ result, err := chain.Execute(context.Background(), candidates, mockRun)
+
+ if err == nil {
+ t.Fatal("expected error when all keys fail, got nil")
+ }
+
+ if result != nil {
+ t.Errorf("expected nil result on failure, got: %v", result)
+ }
+
+ if callCount != 3 {
+ t.Errorf("expected 3 calls (all fail), got %d", callCount)
+ }
+
+ // Verify error type
+ var exhausted *FallbackExhaustedError
+ if !errors.As(err, &exhausted) {
+ t.Errorf("expected FallbackExhaustedError, got: %T - %v", err, err)
+ }
+
+ if len(exhausted.Attempts) != 3 {
+ t.Errorf("expected 3 attempts in exhausted error, got %d", len(exhausted.Attempts))
+ }
+}
+
+// TestMultiKeyFailoverCooldown tests that a key in cooldown is skipped
+func TestMultiKeyFailoverCooldown(t *testing.T) {
+ cfg := ModelConfig{
+ Primary: "glm-4.7",
+ Fallbacks: []string{"glm-4.7__key_1"},
+ }
+
+ candidates := ResolveCandidates(cfg, "zhipu")
+
+ cooldown := NewCooldownTracker()
+ chain := NewFallbackChain(cooldown)
+
+ // Put the first model in cooldown (using ModelKey now, not just provider)
+ cooldownKey := ModelKey(candidates[0].Provider, candidates[0].Model)
+ cooldown.MarkFailure(cooldownKey, FailoverRateLimit)
+
+ // Verify it's not available
+ if cooldown.IsAvailable(cooldownKey) {
+ t.Fatal("expected first model to be in cooldown")
+ }
+
+ // Mock run function: only second should be called
+ callCount := 0
+ calledProviders := []string{}
+ mockRun := func(ctx context.Context, provider, model string) (*LLMResponse, error) {
+ callCount++
+ calledProviders = append(calledProviders, provider+"/"+model)
+ return &LLMResponse{Content: "success"}, nil
+ }
+
+ result, err := chain.Execute(context.Background(), candidates, mockRun)
+ if err != nil {
+ t.Fatalf("expected success, got error: %v", err)
+ }
+
+ // First provider should have been skipped
+ if callCount != 1 {
+ t.Errorf("expected 1 call (first skipped due to cooldown), got %d", callCount)
+ }
+
+ // Should have called the second provider/model
+ if len(calledProviders) != 1 ||
+ calledProviders[0] != candidates[1].Provider+"/"+candidates[1].Model {
+ t.Errorf("expected second model to be called, got: %v", calledProviders)
+ }
+
+ // Verify first attempt was recorded as skipped
+ if len(result.Attempts) != 1 {
+ t.Fatalf("expected 1 attempt (skipped), got %d", len(result.Attempts))
+ }
+
+ if !result.Attempts[0].Skipped {
+ t.Error("expected first attempt to be marked as skipped")
+ }
+}
+
+// TestMultiKeyFailoverWithFormatError tests that format errors are non-retriable
+func TestMultiKeyFailoverWithFormatError(t *testing.T) {
+ cfg := ModelConfig{
+ Primary: "glm-4.7",
+ Fallbacks: []string{"glm-4.7__key_1"},
+ }
+
+ candidates := ResolveCandidates(cfg, "zhipu")
+
+ cooldown := NewCooldownTracker()
+ chain := NewFallbackChain(cooldown)
+
+ // Mock run function: first call fails with format error (bad request)
+ callCount := 0
+ mockRun := func(ctx context.Context, provider, model string) (*LLMResponse, error) {
+ callCount++
+ return nil, errors.New("invalid request format: tool_use.id missing")
+ }
+
+ // Execute fallback chain
+ result, err := chain.Execute(context.Background(), candidates, mockRun)
+
+ if err == nil {
+ t.Fatal("expected error for format failure, got nil")
+ }
+
+ // Format errors should NOT trigger failover (non-retriable)
+ // So we should only have 1 call
+ if callCount != 1 {
+ t.Errorf("expected 1 call (format error is non-retriable), got %d", callCount)
+ }
+
+ // Verify the error is a FailoverError with format reason
+ var failoverErr *FailoverError
+ if !errors.As(err, &failoverErr) {
+ t.Errorf("expected FailoverError, got: %T - %v", err, err)
+ }
+
+ if failoverErr.Reason != FailoverFormat {
+ t.Errorf("expected FailoverFormat reason, got: %s", failoverErr.Reason)
+ }
+
+ _ = result // result should be nil
+}
+
+// TestMultiKeyWithModelFallback tests multi-key failover combined with model fallback.
+// This simulates the scenario: api_keys: ["k1", "k2"] + fallbacks: ["minimax"]
+// Expected failover order: glm-4.7 (k1) → glm-4.7__key_1 (k2) → minimax
+func TestMultiKeyWithModelFallback(t *testing.T) {
+ // Simulate expanded config from:
+ // { "model_name": "glm-4.7", "api_keys": ["k1", "k2"], "fallbacks": ["minimax"] }
+ // After ExpandMultiKeyModels, primaryEntry.Fallbacks = ["glm-4.7__key_1", "minimax"]
+ // Note: In production, "minimax" would be resolved via model lookup to "minimax/minimax"
+ // In this test, we use the full format to avoid needing a lookup function.
+ cfg := ModelConfig{
+ Primary: "glm-4.7",
+ Fallbacks: []string{"glm-4.7__key_1", "minimax/minimax"},
+ }
+
+ candidates := ResolveCandidates(cfg, "zhipu")
+
+ // Should have 3 candidates: glm-4.7 (zhipu), glm-4.7__key_1 (zhipu), minimax (minimax)
+ if len(candidates) != 3 {
+ t.Fatalf("expected 3 candidates, got %d: %v", len(candidates), candidates)
+ }
+
+ // Verify candidate order
+ if candidates[0].Model != "glm-4.7" || candidates[0].Provider != "zhipu" {
+ t.Errorf(
+ "expected first candidate to be zhipu/glm-4.7, got: %s/%s",
+ candidates[0].Provider,
+ candidates[0].Model,
+ )
+ }
+ if candidates[1].Model != "glm-4.7__key_1" || candidates[1].Provider != "zhipu" {
+ t.Errorf(
+ "expected second candidate to be zhipu/glm-4.7__key_1, got: %s/%s",
+ candidates[1].Provider,
+ candidates[1].Model,
+ )
+ }
+ if candidates[2].Model != "minimax" || candidates[2].Provider != "minimax" {
+ t.Errorf(
+ "expected third candidate to be minimax/minimax, got: %s/%s",
+ candidates[2].Provider,
+ candidates[2].Model,
+ )
+ }
+
+ cooldown := NewCooldownTracker()
+ chain := NewFallbackChain(cooldown)
+
+ // Mock run function: first two fail, third succeeds (model fallback)
+ callCount := 0
+ calledModels := []string{}
+ mockRun := func(ctx context.Context, provider, model string) (*LLMResponse, error) {
+ callCount++
+ calledModels = append(calledModels, provider+"/"+model)
+
+ switch callCount {
+ case 1:
+ // k1: rate limit
+ return nil, errors.New("status: 429 - rate limit")
+ case 2:
+ // k2: also rate limit (all zhipu keys exhausted)
+ return nil, errors.New("status: 429 - rate limit")
+ case 3:
+ // minimax: success
+ return &LLMResponse{Content: "success from minimax"}, nil
+ default:
+ return nil, errors.New("unexpected call")
+ }
+ }
+
+ result, err := chain.Execute(context.Background(), candidates, mockRun)
+ if err != nil {
+ t.Fatalf("expected success after failover to model fallback, got error: %v", err)
+ }
+
+ if callCount != 3 {
+ t.Errorf("expected 3 calls (k1 fail + k2 fail + minimax success), got %d", callCount)
+ }
+
+ if result.Response.Content != "success from minimax" {
+ t.Errorf("expected response from minimax, got: %s", result.Response.Content)
+ }
+
+ // Verify call order
+ if len(calledModels) != 3 {
+ t.Fatalf("expected 3 called models, got %d", len(calledModels))
+ }
+ if calledModels[0] != "zhipu/glm-4.7" {
+ t.Errorf("expected first call to zhipu/glm-4.7, got: %s", calledModels[0])
+ }
+ if calledModels[1] != "zhipu/glm-4.7__key_1" {
+ t.Errorf("expected second call to zhipu/glm-4.7__key_1, got: %s", calledModels[1])
+ }
+ if calledModels[2] != "minimax/minimax" {
+ t.Errorf("expected third call to minimax/minimax, got: %s", calledModels[2])
+ }
+
+ // Verify 2 failed attempts recorded
+ if len(result.Attempts) != 2 {
+ t.Errorf("expected 2 failed attempts, got %d", len(result.Attempts))
+ }
+
+ // Both should be rate limit
+ for i, attempt := range result.Attempts {
+ if attempt.Reason != FailoverRateLimit {
+ t.Errorf("expected attempt %d to be rate_limit, got: %s", i, attempt.Reason)
+ }
+ }
+}
+
+// TestMultiKeyFailoverMixedErrors tests failover with different error types
+func TestMultiKeyFailoverMixedErrors(t *testing.T) {
+ cfg := ModelConfig{
+ Primary: "glm-4.7",
+ Fallbacks: []string{"glm-4.7__key_1", "glm-4.7__key_2"},
+ }
+
+ candidates := ResolveCandidates(cfg, "zhipu")
+
+ cooldown := NewCooldownTracker()
+ chain := NewFallbackChain(cooldown)
+
+ // Mock run function: different errors for each key
+ callCount := 0
+ mockRun := func(ctx context.Context, provider, model string) (*LLMResponse, error) {
+ callCount++
+ switch callCount {
+ case 1:
+ // First: rate limit (retriable)
+ return nil, errors.New("status: 429 - rate limit")
+ case 2:
+ // Second: timeout (retriable)
+ return nil, errors.New("context deadline exceeded")
+ case 3:
+ // Third: success
+ return &LLMResponse{Content: "success from key3"}, nil
+ default:
+ return nil, errors.New("unexpected call")
+ }
+ }
+
+ result, err := chain.Execute(context.Background(), candidates, mockRun)
+ if err != nil {
+ t.Fatalf("expected success after 2 failovers, got error: %v", err)
+ }
+
+ if callCount != 3 {
+ t.Errorf("expected 3 calls, got %d", callCount)
+ }
+
+ // Verify both failed attempts were recorded
+ if len(result.Attempts) != 2 {
+ t.Errorf("expected 2 failed attempts, got %d", len(result.Attempts))
+ }
+
+ // First should be rate limit
+ if result.Attempts[0].Reason != FailoverRateLimit {
+ t.Errorf("expected first attempt to be rate_limit, got: %s", result.Attempts[0].Reason)
+ }
+
+ // Second should be timeout
+ if result.Attempts[1].Reason != FailoverTimeout {
+ t.Errorf("expected second attempt to be timeout, got: %s", result.Attempts[1].Reason)
+ }
+}
diff --git a/pkg/providers/fallback_test.go b/pkg/providers/fallback_test.go
index 1783ebcb5..1a1118e33 100644
--- a/pkg/providers/fallback_test.go
+++ b/pkg/providers/fallback_test.go
@@ -157,8 +157,8 @@ func TestFallback_CooldownSkip(t *testing.T) {
ct, _ := newTestTracker(now)
fc := NewFallbackChain(ct)
- // Put openai in cooldown
- ct.MarkFailure("openai", FailoverRateLimit)
+ // Put openai/gpt-4 in cooldown (using ModelKey now)
+ ct.MarkFailure(ModelKey("openai", "gpt-4"), FailoverRateLimit)
candidates := []FallbackCandidate{
makeCandidate("openai", "gpt-4"),
@@ -195,9 +195,9 @@ func TestFallback_AllInCooldown(t *testing.T) {
ct := NewCooldownTracker()
fc := NewFallbackChain(ct)
- // Put all providers in cooldown
- ct.MarkFailure("openai", FailoverRateLimit)
- ct.MarkFailure("anthropic", FailoverBilling)
+ // Put all models in cooldown (using ModelKey now)
+ ct.MarkFailure(ModelKey("openai", "gpt-4"), FailoverRateLimit)
+ ct.MarkFailure(ModelKey("anthropic", "claude"), FailoverBilling)
candidates := []FallbackCandidate{
makeCandidate("openai", "gpt-4"),
@@ -273,12 +273,13 @@ func TestFallback_SuccessResetsCooldown(t *testing.T) {
fc := NewFallbackChain(ct)
candidates := []FallbackCandidate{makeCandidate("openai", "gpt-4")}
+ modelKey := ModelKey("openai", "gpt-4")
attempt := 0
run := func(ctx context.Context, provider, model string) (*LLMResponse, error) {
attempt++
if attempt == 1 {
- ct.MarkFailure("openai", FailoverRateLimit) // simulate failure tracked elsewhere
+ ct.MarkFailure(modelKey, FailoverRateLimit) // simulate failure tracked elsewhere
}
return &LLMResponse{Content: "ok", FinishReason: "stop"}, nil
}
@@ -287,7 +288,7 @@ func TestFallback_SuccessResetsCooldown(t *testing.T) {
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
- if !ct.IsAvailable("openai") {
+ if !ct.IsAvailable(modelKey) {
t.Error("success should reset cooldown")
}
}
diff --git a/pkg/providers/http_provider.go b/pkg/providers/http_provider.go
index 5c328f418..803165edb 100644
--- a/pkg/providers/http_provider.go
+++ b/pkg/providers/http_provider.go
@@ -52,6 +52,23 @@ func (p *HTTPProvider) Chat(
return p.delegate.Chat(ctx, messages, tools, model, options)
}
+// ChatStream implements providers.StreamingProvider by delegating to the
+// OpenAI-compatible streaming endpoint (SSE with stream: true).
+func (p *HTTPProvider) ChatStream(
+ ctx context.Context,
+ messages []Message,
+ tools []ToolDefinition,
+ model string,
+ options map[string]any,
+ onChunk func(accumulated string),
+) (*LLMResponse, error) {
+ return p.delegate.ChatStream(ctx, messages, tools, model, options, onChunk)
+}
+
func (p *HTTPProvider) GetDefaultModel() string {
return ""
}
+
+func (p *HTTPProvider) SupportsNativeSearch() bool {
+ return p.delegate.SupportsNativeSearch()
+}
diff --git a/pkg/providers/model_ref.go b/pkg/providers/model_ref.go
index 0d1b02d16..be9f63bc6 100644
--- a/pkg/providers/model_ref.go
+++ b/pkg/providers/model_ref.go
@@ -53,6 +53,14 @@ func NormalizeProvider(provider string) string {
return "zhipu"
case "google":
return "gemini"
+ case "alibaba-coding", "qwen-coding":
+ return "coding-plan"
+ case "alibaba-coding-anthropic":
+ return "coding-plan-anthropic"
+ case "qwen-international", "dashscope-intl":
+ return "qwen-intl"
+ case "dashscope-us":
+ return "qwen-us"
}
return p
diff --git a/pkg/providers/model_ref_test.go b/pkg/providers/model_ref_test.go
index 6dd25167f..040c511ba 100644
--- a/pkg/providers/model_ref_test.go
+++ b/pkg/providers/model_ref_test.go
@@ -73,6 +73,14 @@ func TestNormalizeProvider(t *testing.T) {
{"glm", "zhipu"},
{"google", "gemini"},
{"groq", "groq"},
+ // Alibaba Coding Plan aliases
+ {"alibaba-coding", "coding-plan"},
+ {"qwen-coding", "coding-plan"},
+ {"alibaba-coding-anthropic", "coding-plan-anthropic"},
+ // Qwen international aliases
+ {"qwen-international", "qwen-intl"},
+ {"dashscope-intl", "qwen-intl"},
+ {"dashscope-us", "qwen-us"},
{"", ""},
}
diff --git a/pkg/providers/openai_compat/provider.go b/pkg/providers/openai_compat/provider.go
index f97bf3acd..938e4ea8b 100644
--- a/pkg/providers/openai_compat/provider.go
+++ b/pkg/providers/openai_compat/provider.go
@@ -13,6 +13,7 @@ import (
"strings"
"time"
+ "github.com/sipeed/picoclaw/pkg/providers/common"
"github.com/sipeed/picoclaw/pkg/providers/protocoltypes"
)
@@ -38,7 +39,7 @@ type Provider struct {
type Option func(*Provider)
-const defaultRequestTimeout = 120 * time.Second
+const defaultRequestTimeout = common.DefaultRequestTimeout
func WithMaxTokensField(maxTokensField string) Option {
return func(p *Provider) {
@@ -55,25 +56,10 @@ func WithRequestTimeout(timeout time.Duration) Option {
}
func NewProvider(apiKey, apiBase, proxy string, opts ...Option) *Provider {
- client := &http.Client{
- Timeout: defaultRequestTimeout,
- }
-
- if proxy != "" {
- parsed, err := url.Parse(proxy)
- if err == nil {
- client.Transport = &http.Transport{
- Proxy: http.ProxyURL(parsed),
- }
- } else {
- log.Printf("openai_compat: invalid proxy URL %q: %v", proxy, err)
- }
- }
-
p := &Provider{
apiKey: apiKey,
apiBase: strings.TrimRight(apiBase, "/"),
- httpClient: client,
+ httpClient: common.NewHTTPClient(proxy),
}
for _, opt := range opts {
@@ -102,34 +88,28 @@ func NewProviderWithMaxTokensFieldAndTimeout(
)
}
-func (p *Provider) Chat(
- ctx context.Context,
- messages []Message,
- tools []ToolDefinition,
- model string,
- options map[string]any,
-) (*LLMResponse, error) {
- if p.apiBase == "" {
- return nil, fmt.Errorf("API base not configured")
- }
-
+// buildRequestBody constructs the common request body for Chat and ChatStream.
+func (p *Provider) buildRequestBody(
+ messages []Message, tools []ToolDefinition, model string, options map[string]any,
+) map[string]any {
model = normalizeModel(model, p.apiBase)
requestBody := map[string]any{
"model": model,
- "messages": serializeMessages(messages),
+ "messages": common.SerializeMessages(messages),
}
- if len(tools) > 0 {
- requestBody["tools"] = tools
+ // When fallback uses a different provider (e.g. DeepSeek), that provider must not inject web_search_preview.
+ nativeSearch, _ := options["native_search"].(bool)
+ nativeSearch = nativeSearch && isNativeSearchHost(p.apiBase)
+ if len(tools) > 0 || nativeSearch {
+ requestBody["tools"] = buildToolsList(tools, nativeSearch)
requestBody["tool_choice"] = "auto"
}
- if maxTokens, ok := asInt(options["max_tokens"]); ok {
- // Use configured maxTokensField if specified, otherwise fallback to model-based detection
+ if maxTokens, ok := common.AsInt(options["max_tokens"]); ok {
fieldName := p.maxTokensField
if fieldName == "" {
- // Fallback: detect from model name for backward compatibility
lowerModel := strings.ToLower(model)
if strings.Contains(lowerModel, "glm") || strings.Contains(lowerModel, "o1") ||
strings.Contains(lowerModel, "gpt-5") {
@@ -141,9 +121,8 @@ func (p *Provider) Chat(
requestBody[fieldName] = maxTokens
}
- if temperature, ok := asFloat(options["temperature"]); ok {
+ if temperature, ok := common.AsFloat(options["temperature"]); ok {
lowerModel := strings.ToLower(model)
- // Kimi k2 models only support temperature=1.
if strings.Contains(lowerModel, "kimi") && strings.Contains(lowerModel, "k2") {
requestBody["temperature"] = 1.0
} else {
@@ -153,17 +132,30 @@ func (p *Provider) Chat(
// Prompt caching: pass a stable cache key so OpenAI can bucket requests
// with the same key and reuse prefix KV cache across calls.
- // The key is typically the agent ID — stable per agent, shared across requests.
- // See: https://platform.openai.com/docs/guides/prompt-caching
// Prompt caching is only supported by OpenAI-native endpoints.
- // Non-OpenAI providers (Mistral, Gemini, DeepSeek, etc.) reject unknown
- // fields with 422 errors, so only include it for OpenAI APIs.
+ // Non-OpenAI providers reject unknown fields with 422 errors.
if cacheKey, ok := options["prompt_cache_key"].(string); ok && cacheKey != "" {
if supportsPromptCacheKey(p.apiBase) {
requestBody["prompt_cache_key"] = cacheKey
}
}
+ return requestBody
+}
+
+func (p *Provider) Chat(
+ ctx context.Context,
+ messages []Message,
+ tools []ToolDefinition,
+ model string,
+ options map[string]any,
+) (*LLMResponse, error) {
+ if p.apiBase == "" {
+ return nil, fmt.Errorf("API base not configured")
+ }
+
+ requestBody := p.buildRequestBody(messages, tools, model, options)
+
jsonData, err := json.Marshal(requestBody)
if err != nil {
return nil, fmt.Errorf("failed to marshal request: %w", err)
@@ -185,275 +177,200 @@ func (p *Provider) Chat(
}
defer resp.Body.Close()
- contentType := resp.Header.Get("Content-Type")
-
- // Non-200: read a prefix to tell HTML error page apart from JSON error body.
if resp.StatusCode != http.StatusOK {
- body, readErr := io.ReadAll(io.LimitReader(resp.Body, 256))
- if readErr != nil {
- return nil, fmt.Errorf("failed to read response: %w", readErr)
- }
- if looksLikeHTML(body, contentType) {
- return nil, wrapHTMLResponseError(resp.StatusCode, body, contentType, p.apiBase)
- }
- return nil, fmt.Errorf(
- "API request failed:\n Status: %d\n Body: %s",
- resp.StatusCode,
- responsePreview(body, 128),
- )
+ return nil, common.HandleErrorResponse(resp, p.apiBase)
}
- // Peek without consuming so the full stream reaches the JSON decoder.
- reader := bufio.NewReader(resp.Body)
- prefix, err := reader.Peek(256) // io.EOF/ErrBufferFull are normal; only real errors abort
- if err != nil && err != io.EOF && err != bufio.ErrBufferFull {
- return nil, fmt.Errorf("failed to inspect response: %w", err)
- }
- if looksLikeHTML(prefix, contentType) {
- return nil, wrapHTMLResponseError(resp.StatusCode, prefix, contentType, p.apiBase)
+ return common.ReadAndParseResponse(resp, p.apiBase)
+}
+
+// ChatStream implements streaming via OpenAI-compatible SSE (stream: true).
+// onChunk receives the accumulated text so far on each text delta.
+func (p *Provider) ChatStream(
+ ctx context.Context,
+ messages []Message,
+ tools []ToolDefinition,
+ model string,
+ options map[string]any,
+ onChunk func(accumulated string),
+) (*LLMResponse, error) {
+ if p.apiBase == "" {
+ return nil, fmt.Errorf("API base not configured")
}
- out, err := parseResponse(reader)
+ requestBody := p.buildRequestBody(messages, tools, model, options)
+ requestBody["stream"] = true
+
+ jsonData, err := json.Marshal(requestBody)
if err != nil {
- return nil, fmt.Errorf("failed to parse JSON response: %w", err)
+ return nil, fmt.Errorf("failed to marshal request: %w", err)
}
- return out, nil
+ req, err := http.NewRequestWithContext(ctx, "POST", p.apiBase+"/chat/completions", bytes.NewReader(jsonData))
+ if err != nil {
+ return nil, fmt.Errorf("failed to create request: %w", err)
+ }
+
+ req.Header.Set("Content-Type", "application/json")
+ req.Header.Set("Accept", "text/event-stream")
+ if p.apiKey != "" {
+ req.Header.Set("Authorization", "Bearer "+p.apiKey)
+ }
+
+ // Use a client without Timeout for streaming — the http.Client.Timeout covers
+ // the entire request lifecycle including body reads, which would kill long streams.
+ // Context cancellation still provides the safety net.
+ streamClient := &http.Client{Transport: p.httpClient.Transport}
+ resp, err := streamClient.Do(req)
+ if err != nil {
+ return nil, fmt.Errorf("failed to send request: %w", err)
+ }
+ defer resp.Body.Close()
+
+ if resp.StatusCode != http.StatusOK {
+ return nil, common.HandleErrorResponse(resp, p.apiBase)
+ }
+
+ return parseStreamResponse(ctx, resp.Body, onChunk)
}
-func wrapHTMLResponseError(statusCode int, body []byte, contentType, apiBase string) error {
- respPreview := responsePreview(body, 128)
- return fmt.Errorf(
- "API request failed: %s returned HTML instead of JSON (content-type: %s); check api_base or proxy configuration.\n Status: %d\n Body: %s",
- apiBase,
- contentType,
- statusCode,
- respPreview,
- )
-}
+// parseStreamResponse parses an OpenAI-compatible SSE stream.
+func parseStreamResponse(
+ ctx context.Context,
+ reader io.Reader,
+ onChunk func(accumulated string),
+) (*LLMResponse, error) {
+ var textContent strings.Builder
+ var finishReason string
+ var usage *UsageInfo
-func looksLikeHTML(body []byte, contentType string) bool {
- contentType = strings.ToLower(strings.TrimSpace(contentType))
- if strings.Contains(contentType, "text/html") || strings.Contains(contentType, "application/xhtml+xml") {
- return true
+ // Tool call assembly: OpenAI streams tool calls as incremental deltas
+ type toolAccum struct {
+ id string
+ name string
+ argsJSON strings.Builder
}
- prefix := bytes.ToLower(leadingTrimmedPrefix(body, 128))
- return bytes.HasPrefix(prefix, []byte(" len(body) {
- end = len(body)
- }
- return body[i:end]
- }
- }
- return nil
-}
-
-func responsePreview(body []byte, maxLen int) string {
- trimmed := bytes.TrimSpace(body)
- if len(trimmed) == 0 {
- return ""
- }
- if len(trimmed) <= maxLen {
- return string(trimmed)
- }
- return string(trimmed[:maxLen]) + "..."
-}
-
-func parseResponse(body io.Reader) (*LLMResponse, error) {
- var apiResponse struct {
- Choices []struct {
- Message struct {
- Content string `json:"content"`
- ReasoningContent string `json:"reasoning_content"`
- Reasoning string `json:"reasoning"`
- ReasoningDetails []ReasoningDetail `json:"reasoning_details"`
- ToolCalls []struct {
- ID string `json:"id"`
- Type string `json:"type"`
- Function *struct {
- Name string `json:"name"`
- Arguments json.RawMessage `json:"arguments"`
- } `json:"function"`
- ExtraContent *struct {
- Google *struct {
- ThoughtSignature string `json:"thought_signature"`
- } `json:"google"`
- } `json:"extra_content"`
- } `json:"tool_calls"`
- } `json:"message"`
- FinishReason string `json:"finish_reason"`
- } `json:"choices"`
- Usage *UsageInfo `json:"usage"`
- }
-
- if err := json.NewDecoder(body).Decode(&apiResponse); err != nil {
- return nil, fmt.Errorf("failed to decode response: %w", err)
- }
-
- if len(apiResponse.Choices) == 0 {
- return &LLMResponse{
- Content: "",
- FinishReason: "stop",
- }, nil
- }
-
- choice := apiResponse.Choices[0]
- toolCalls := make([]ToolCall, 0, len(choice.Message.ToolCalls))
- for _, tc := range choice.Message.ToolCalls {
- arguments := make(map[string]any)
- name := ""
-
- // Extract thought_signature from Gemini/Google-specific extra content
- thoughtSignature := ""
- if tc.ExtraContent != nil && tc.ExtraContent.Google != nil {
- thoughtSignature = tc.ExtraContent.Google.ThoughtSignature
+ scanner := bufio.NewScanner(reader)
+ scanner.Buffer(make([]byte, 0, 1024*1024), 10*1024*1024) // 1MB initial, 10MB max
+ for scanner.Scan() {
+ // Check for context cancellation between chunks
+ if err := ctx.Err(); err != nil {
+ return nil, err
}
- if tc.Function != nil {
- name = tc.Function.Name
- arguments = decodeToolCallArguments(tc.Function.Arguments, name)
+ line := scanner.Text()
+
+ if !strings.HasPrefix(line, "data: ") {
+ continue
+ }
+ data := strings.TrimPrefix(line, "data: ")
+ if data == "[DONE]" {
+ break
}
- // Build ToolCall with ExtraContent for Gemini 3 thought_signature persistence
- toolCall := ToolCall{
- ID: tc.ID,
- Name: name,
- Arguments: arguments,
- ThoughtSignature: thoughtSignature,
+ var chunk struct {
+ Choices []struct {
+ Delta struct {
+ Content string `json:"content"`
+ ToolCalls []struct {
+ Index int `json:"index"`
+ ID string `json:"id"`
+ Function *struct {
+ Name string `json:"name"`
+ Arguments string `json:"arguments"`
+ } `json:"function"`
+ } `json:"tool_calls"`
+ } `json:"delta"`
+ FinishReason *string `json:"finish_reason"`
+ } `json:"choices"`
+ Usage *UsageInfo `json:"usage"`
}
- if thoughtSignature != "" {
- toolCall.ExtraContent = &ExtraContent{
- Google: &GoogleExtra{
- ThoughtSignature: thoughtSignature,
- },
- }
+ if err := json.Unmarshal([]byte(data), &chunk); err != nil {
+ continue // skip malformed chunks
}
- toolCalls = append(toolCalls, toolCall)
- }
-
- return &LLMResponse{
- Content: choice.Message.Content,
- ReasoningContent: choice.Message.ReasoningContent,
- Reasoning: choice.Message.Reasoning,
- ReasoningDetails: choice.Message.ReasoningDetails,
- ToolCalls: toolCalls,
- FinishReason: choice.FinishReason,
- Usage: apiResponse.Usage,
- }, nil
-}
-
-func decodeToolCallArguments(raw json.RawMessage, name string) map[string]any {
- arguments := make(map[string]any)
- raw = bytes.TrimSpace(raw)
- if len(raw) == 0 || bytes.Equal(raw, []byte("null")) {
- return arguments
- }
-
- var decoded any
- if err := json.Unmarshal(raw, &decoded); err != nil {
- log.Printf("openai_compat: failed to decode tool call arguments payload for %q: %v", name, err)
- arguments["raw"] = string(raw)
- return arguments
- }
-
- switch v := decoded.(type) {
- case string:
- if strings.TrimSpace(v) == "" {
- return arguments
+ if chunk.Usage != nil {
+ usage = chunk.Usage
}
- if err := json.Unmarshal([]byte(v), &arguments); err != nil {
- log.Printf("openai_compat: failed to decode tool call arguments for %q: %v", name, err)
- arguments["raw"] = v
- }
- return arguments
- case map[string]any:
- return v
- default:
- log.Printf("openai_compat: unsupported tool call arguments type for %q: %T", name, decoded)
- arguments["raw"] = string(raw)
- return arguments
- }
-}
-// openaiMessage is the wire-format message for OpenAI-compatible APIs.
-// It mirrors protocoltypes.Message but omits SystemParts, which is an
-// internal field that would be unknown to third-party endpoints.
-type openaiMessage struct {
- Role string `json:"role"`
- Content string `json:"content"`
- ReasoningContent string `json:"reasoning_content,omitempty"`
- ToolCalls []ToolCall `json:"tool_calls,omitempty"`
- ToolCallID string `json:"tool_call_id,omitempty"`
-}
-
-// serializeMessages converts internal Message structs to the OpenAI wire format.
-// - Strips SystemParts (unknown to third-party endpoints)
-// - Converts messages with Media to multipart content format (text + image_url parts)
-// - Preserves ToolCallID, ToolCalls, and ReasoningContent for all messages
-func serializeMessages(messages []Message) []any {
- out := make([]any, 0, len(messages))
- for _, m := range messages {
- if len(m.Media) == 0 {
- out = append(out, openaiMessage{
- Role: m.Role,
- Content: m.Content,
- ReasoningContent: m.ReasoningContent,
- ToolCalls: m.ToolCalls,
- ToolCallID: m.ToolCallID,
- })
+ if len(chunk.Choices) == 0 {
continue
}
- // Multipart content format for messages with media
- parts := make([]map[string]any, 0, 1+len(m.Media))
- if m.Content != "" {
- parts = append(parts, map[string]any{
- "type": "text",
- "text": m.Content,
- })
- }
- for _, mediaURL := range m.Media {
- if strings.HasPrefix(mediaURL, "data:image/") {
- parts = append(parts, map[string]any{
- "type": "image_url",
- "image_url": map[string]any{
- "url": mediaURL,
- },
- })
+ choice := chunk.Choices[0]
+
+ // Accumulate text content
+ if choice.Delta.Content != "" {
+ textContent.WriteString(choice.Delta.Content)
+ if onChunk != nil {
+ onChunk(textContent.String())
}
}
- msg := map[string]any{
- "role": m.Role,
- "content": parts,
+ // Accumulate tool call deltas
+ for _, tc := range choice.Delta.ToolCalls {
+ acc, ok := activeTools[tc.Index]
+ if !ok {
+ acc = &toolAccum{}
+ activeTools[tc.Index] = acc
+ }
+ if tc.ID != "" {
+ acc.id = tc.ID
+ }
+ if tc.Function != nil {
+ if tc.Function.Name != "" {
+ acc.name = tc.Function.Name
+ }
+ if tc.Function.Arguments != "" {
+ acc.argsJSON.WriteString(tc.Function.Arguments)
+ }
+ }
}
- if m.ToolCallID != "" {
- msg["tool_call_id"] = m.ToolCallID
+
+ if choice.FinishReason != nil {
+ finishReason = *choice.FinishReason
}
- if len(m.ToolCalls) > 0 {
- msg["tool_calls"] = m.ToolCalls
- }
- if m.ReasoningContent != "" {
- msg["reasoning_content"] = m.ReasoningContent
- }
- out = append(out, msg)
}
- return out
+
+ if err := scanner.Err(); err != nil {
+ return nil, fmt.Errorf("streaming read error: %w", err)
+ }
+
+ // Assemble tool calls from accumulated deltas
+ var toolCalls []ToolCall
+ for i := 0; i < len(activeTools); i++ {
+ acc, ok := activeTools[i]
+ if !ok {
+ continue
+ }
+ args := make(map[string]any)
+ raw := acc.argsJSON.String()
+ if raw != "" {
+ if err := json.Unmarshal([]byte(raw), &args); err != nil {
+ log.Printf("openai_compat stream: failed to decode tool call arguments for %q: %v", acc.name, err)
+ args["raw"] = raw
+ }
+ }
+ toolCalls = append(toolCalls, ToolCall{
+ ID: acc.id,
+ Name: acc.name,
+ Arguments: args,
+ })
+ }
+
+ if finishReason == "" {
+ finishReason = "stop"
+ }
+
+ return &LLMResponse{
+ Content: textContent.String(),
+ ToolCalls: toolCalls,
+ FinishReason: finishReason,
+ Usage: usage,
+ }, nil
}
func normalizeModel(model, apiBase string) string {
@@ -469,41 +386,38 @@ func normalizeModel(model, apiBase string) string {
prefix := strings.ToLower(before)
switch prefix {
case "litellm", "moonshot", "nvidia", "groq", "ollama", "deepseek", "google",
- "openrouter", "zhipu", "mistral", "vivgrid", "minimax":
+ "openrouter", "zhipu", "mistral", "vivgrid", "minimax", "novita":
return after
default:
return model
}
}
-func asInt(v any) (int, bool) {
- switch val := v.(type) {
- case int:
- return val, true
- case int64:
- return int(val), true
- case float64:
- return int(val), true
- case float32:
- return int(val), true
- default:
- return 0, false
+func buildToolsList(tools []ToolDefinition, nativeSearch bool) []any {
+ result := make([]any, 0, len(tools)+1)
+ for _, t := range tools {
+ if nativeSearch && strings.EqualFold(t.Function.Name, "web_search") {
+ continue
+ }
+ result = append(result, t)
}
+ if nativeSearch {
+ result = append(result, map[string]any{"type": "web_search_preview"})
+ }
+ return result
}
-func asFloat(v any) (float64, bool) {
- switch val := v.(type) {
- case float64:
- return val, true
- case float32:
- return float64(val), true
- case int:
- return float64(val), true
- case int64:
- return float64(val), true
- default:
- return 0, false
+func (p *Provider) SupportsNativeSearch() bool {
+ return isNativeSearchHost(p.apiBase)
+}
+
+func isNativeSearchHost(apiBase string) bool {
+ u, err := url.Parse(apiBase)
+ if err != nil {
+ return false
}
+ host := u.Hostname()
+ return host == "api.openai.com" || strings.HasSuffix(host, ".openai.azure.com")
}
// supportsPromptCacheKey reports whether the given API base is known to
diff --git a/pkg/providers/openai_compat/provider_test.go b/pkg/providers/openai_compat/provider_test.go
index 41f278a1b..efb03ccb8 100644
--- a/pkg/providers/openai_compat/provider_test.go
+++ b/pkg/providers/openai_compat/provider_test.go
@@ -12,6 +12,7 @@ import (
"testing"
"time"
+ "github.com/sipeed/picoclaw/pkg/providers/common"
"github.com/sipeed/picoclaw/pkg/providers/protocoltypes"
)
@@ -431,7 +432,28 @@ func TestProviderChat_StripsMoonshotPrefixAndNormalizesKimiTemperature(t *testin
}
}
-func TestProviderChat_StripsGroqOllamaDeepseekVivgridPrefixes(t *testing.T) {
+func TestProviderChat_StripsGroqOllamaDeepseekVivgridNovitaPrefixes(t *testing.T) {
+ var requestBody map[string]any
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if err := json.NewDecoder(r.Body).Decode(&requestBody); err != nil {
+ http.Error(w, err.Error(), http.StatusBadRequest)
+ return
+ }
+ resp := map[string]any{
+ "choices": []map[string]any{
+ {
+ "message": map[string]any{"content": "ok"},
+ "finish_reason": "stop",
+ },
+ },
+ }
+ w.Header().Set("Content-Type", "application/json")
+ json.NewEncoder(w).Encode(resp)
+ }))
+ defer server.Close()
+
+ p := NewProvider("key", server.URL, "")
tests := []struct {
name string
input string
@@ -462,31 +484,25 @@ func TestProviderChat_StripsGroqOllamaDeepseekVivgridPrefixes(t *testing.T) {
input: "vivgrid/auto",
wantModel: "auto",
},
+ {
+ name: "strips novita prefix deepseek model",
+ input: "novita/deepseek/deepseek-v3.2",
+ wantModel: "deepseek/deepseek-v3.2",
+ },
+ {
+ name: "strips novita prefix zai model",
+ input: "novita/zai-org/glm-5",
+ wantModel: "zai-org/glm-5",
+ },
+ {
+ name: "strips novita prefix minimax model",
+ input: "novita/minimax/minimax-m2.5",
+ wantModel: "minimax/minimax-m2.5",
+ },
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
- var requestBody map[string]any
-
- server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
- if err := json.NewDecoder(r.Body).Decode(&requestBody); err != nil {
- http.Error(w, err.Error(), http.StatusBadRequest)
- return
- }
- resp := map[string]any{
- "choices": []map[string]any{
- {
- "message": map[string]any{"content": "ok"},
- "finish_reason": "stop",
- },
- },
- }
- w.Header().Set("Content-Type", "application/json")
- json.NewEncoder(w).Encode(resp)
- }))
- defer server.Close()
-
- p := NewProvider("key", server.URL, "")
_, err := p.Chat(t.Context(), []Message{{Role: "user", Content: "hi"}}, nil, tt.input, nil)
if err != nil {
t.Fatalf("Chat() error = %v", err)
@@ -572,6 +588,12 @@ func TestNormalizeModel_UsesAPIBase(t *testing.T) {
if got := normalizeModel("vivgrid/auto", "https://api.vivgrid.com/v1"); got != "auto" {
t.Fatalf("normalizeModel(vivgrid auto) = %q, want %q", got, "auto")
}
+ if got := normalizeModel(
+ "novita/deepseek/deepseek-v3.2",
+ "https://api.novita.ai/openai",
+ ); got != "deepseek/deepseek-v3.2" {
+ t.Fatalf("normalizeModel(novita) = %q, want %q", got, "deepseek/deepseek-v3.2")
+ }
}
func TestProvider_RequestTimeoutDefault(t *testing.T) {
@@ -648,7 +670,7 @@ func TestSerializeMessages_PlainText(t *testing.T) {
{Role: "user", Content: "hello"},
{Role: "assistant", Content: "hi", ReasoningContent: "thinking..."},
}
- result := serializeMessages(messages)
+ result := common.SerializeMessages(messages)
data, err := json.Marshal(result)
if err != nil {
@@ -670,7 +692,7 @@ func TestSerializeMessages_WithMedia(t *testing.T) {
messages := []protocoltypes.Message{
{Role: "user", Content: "describe this", Media: []string{"data:image/png;base64,abc123"}},
}
- result := serializeMessages(messages)
+ result := common.SerializeMessages(messages)
data, _ := json.Marshal(result)
var msgs []map[string]any
@@ -703,7 +725,7 @@ func TestSerializeMessages_MediaWithToolCallID(t *testing.T) {
messages := []protocoltypes.Message{
{Role: "tool", Content: "image result", Media: []string{"data:image/png;base64,xyz"}, ToolCallID: "call_1"},
}
- result := serializeMessages(messages)
+ result := common.SerializeMessages(messages)
data, _ := json.Marshal(result)
var msgs []map[string]any
@@ -823,6 +845,232 @@ func TestSupportsPromptCacheKey(t *testing.T) {
}
}
+func TestBuildToolsList_NativeSearchAddsWebSearchPreview(t *testing.T) {
+ tools := []ToolDefinition{
+ {Type: "function", Function: ToolFunctionDefinition{Name: "read_file", Description: "read"}},
+ }
+ result := buildToolsList(tools, true)
+ if len(result) != 2 {
+ t.Fatalf("len(result) = %d, want 2", len(result))
+ }
+ wsEntry, ok := result[1].(map[string]any)
+ if !ok {
+ t.Fatalf("web search entry is %T, want map[string]any", result[1])
+ }
+ if wsEntry["type"] != "web_search_preview" {
+ t.Fatalf("type = %v, want web_search_preview", wsEntry["type"])
+ }
+}
+
+func TestBuildToolsList_NativeSearchFiltersClientWebSearch(t *testing.T) {
+ tools := []ToolDefinition{
+ {Type: "function", Function: ToolFunctionDefinition{Name: "web_search", Description: "search"}},
+ {Type: "function", Function: ToolFunctionDefinition{Name: "read_file", Description: "read"}},
+ }
+ result := buildToolsList(tools, true)
+ for _, entry := range result {
+ if td, ok := entry.(ToolDefinition); ok && strings.EqualFold(td.Function.Name, "web_search") {
+ t.Fatal("client-side web_search should be filtered out when native search is enabled")
+ }
+ }
+ if len(result) != 2 { // read_file + web_search_preview
+ t.Fatalf("len(result) = %d, want 2 (read_file + web_search_preview)", len(result))
+ }
+}
+
+func TestBuildToolsList_NoNativeSearchPassesThrough(t *testing.T) {
+ tools := []ToolDefinition{
+ {Type: "function", Function: ToolFunctionDefinition{Name: "web_search", Description: "search"}},
+ {Type: "function", Function: ToolFunctionDefinition{Name: "read_file", Description: "read"}},
+ }
+ result := buildToolsList(tools, false)
+ if len(result) != 2 {
+ t.Fatalf("len(result) = %d, want 2", len(result))
+ }
+}
+
+func TestIsNativeSearchHost(t *testing.T) {
+ tests := []struct {
+ apiBase string
+ want bool
+ }{
+ {"https://api.openai.com/v1", true},
+ {"https://myresource.openai.azure.com/openai/deployments/gpt-4", true},
+ {"https://api.mistral.ai/v1", false},
+ {"https://api.deepseek.com/v1", false},
+ {"https://api.groq.com/openai/v1", false},
+ {"http://localhost:11434/v1", false},
+ {"", false},
+ }
+ for _, tt := range tests {
+ if got := isNativeSearchHost(tt.apiBase); got != tt.want {
+ t.Errorf("isNativeSearchHost(%q) = %v, want %v", tt.apiBase, got, tt.want)
+ }
+ }
+}
+
+func TestSupportsNativeSearch_OpenAI(t *testing.T) {
+ p := NewProvider("key", "https://api.openai.com/v1", "")
+ if !p.SupportsNativeSearch() {
+ t.Fatal("OpenAI provider should support native search")
+ }
+}
+
+func TestSupportsNativeSearch_NonOpenAI(t *testing.T) {
+ p := NewProvider("key", "https://api.deepseek.com/v1", "")
+ if p.SupportsNativeSearch() {
+ t.Fatal("DeepSeek provider should not support native search")
+ }
+}
+
+func TestProviderChat_NativeSearchToolInjected(t *testing.T) {
+ var requestBody map[string]any
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if err := json.NewDecoder(r.Body).Decode(&requestBody); err != nil {
+ http.Error(w, err.Error(), http.StatusBadRequest)
+ return
+ }
+ resp := map[string]any{
+ "choices": []map[string]any{
+ {
+ "message": map[string]any{"content": "ok"},
+ "finish_reason": "stop",
+ },
+ },
+ }
+ w.Header().Set("Content-Type", "application/json")
+ json.NewEncoder(w).Encode(resp)
+ }))
+ defer server.Close()
+
+ p := NewProvider("key", server.URL, "")
+ p.apiBase = "https://api.openai.com/v1"
+ p.httpClient = &http.Client{
+ Transport: roundTripperFunc(func(r *http.Request) (*http.Response, error) {
+ r.URL, _ = url.Parse(server.URL + r.URL.Path)
+ return http.DefaultTransport.RoundTrip(r)
+ }),
+ }
+ tools := []ToolDefinition{
+ {Type: "function", Function: ToolFunctionDefinition{Name: "read_file", Description: "read"}},
+ }
+ _, err := p.Chat(
+ t.Context(),
+ []Message{{Role: "user", Content: "hi"}},
+ tools,
+ "gpt-5.4",
+ map[string]any{"native_search": true},
+ )
+ if err != nil {
+ t.Fatalf("Chat() error = %v", err)
+ }
+
+ toolsRaw, ok := requestBody["tools"].([]any)
+ if !ok {
+ t.Fatalf("tools is %T, want []any", requestBody["tools"])
+ }
+ if len(toolsRaw) != 2 {
+ t.Fatalf("len(tools) = %d, want 2 (read_file + web_search_preview)", len(toolsRaw))
+ }
+
+ lastTool, ok := toolsRaw[1].(map[string]any)
+ if !ok {
+ t.Fatalf("last tool is %T, want map[string]any", toolsRaw[1])
+ }
+ if lastTool["type"] != "web_search_preview" {
+ t.Fatalf("last tool type = %v, want web_search_preview", lastTool["type"])
+ }
+}
+
+func TestProviderChat_NativeSearchNotInjectedWithoutOption(t *testing.T) {
+ var requestBody map[string]any
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if err := json.NewDecoder(r.Body).Decode(&requestBody); err != nil {
+ http.Error(w, err.Error(), http.StatusBadRequest)
+ return
+ }
+ resp := map[string]any{
+ "choices": []map[string]any{
+ {
+ "message": map[string]any{"content": "ok"},
+ "finish_reason": "stop",
+ },
+ },
+ }
+ w.Header().Set("Content-Type", "application/json")
+ json.NewEncoder(w).Encode(resp)
+ }))
+ defer server.Close()
+
+ p := NewProvider("key", server.URL, "")
+ tools := []ToolDefinition{
+ {Type: "function", Function: ToolFunctionDefinition{Name: "web_search", Description: "search"}},
+ }
+ _, err := p.Chat(
+ t.Context(),
+ []Message{{Role: "user", Content: "hi"}},
+ tools,
+ "gpt-5.4",
+ map[string]any{},
+ )
+ if err != nil {
+ t.Fatalf("Chat() error = %v", err)
+ }
+
+ toolsRaw, ok := requestBody["tools"].([]any)
+ if !ok {
+ t.Fatalf("tools is %T, want []any", requestBody["tools"])
+ }
+ if len(toolsRaw) != 1 {
+ t.Fatalf("len(tools) = %d, want 1 (web_search only)", len(toolsRaw))
+ }
+}
+
+// TestProviderChat_NativeSearchIgnoredOnNonOpenAI verifies that when native_search
+// is true in options but the provider's apiBase is not OpenAI (e.g. fallback to DeepSeek),
+// we do not inject web_search_preview to avoid API errors.
+func TestProviderChat_NativeSearchIgnoredOnNonOpenAI(t *testing.T) {
+ var requestBody map[string]any
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if err := json.NewDecoder(r.Body).Decode(&requestBody); err != nil {
+ http.Error(w, err.Error(), http.StatusBadRequest)
+ return
+ }
+ resp := map[string]any{
+ "choices": []map[string]any{
+ {
+ "message": map[string]any{"content": "ok"},
+ "finish_reason": "stop",
+ },
+ },
+ }
+ w.Header().Set("Content-Type", "application/json")
+ json.NewEncoder(w).Encode(resp)
+ }))
+ defer server.Close()
+
+ // Use server.URL so host is not api.openai.com — simulates DeepSeek/other provider
+ p := NewProvider("key", server.URL, "")
+ _, err := p.Chat(
+ t.Context(),
+ []Message{{Role: "user", Content: "hi"}},
+ nil,
+ "deepseek-chat",
+ map[string]any{"native_search": true},
+ )
+ if err != nil {
+ t.Fatalf("Chat() error = %v", err)
+ }
+
+ // Should not have tools at all (no tools passed, and we must not add web_search_preview)
+ if toolsRaw, ok := requestBody["tools"]; ok {
+ t.Fatalf("tools should be omitted for non-OpenAI when only native_search was requested, got %v", toolsRaw)
+ }
+}
+
func TestSerializeMessages_StripsSystemParts(t *testing.T) {
messages := []protocoltypes.Message{
{
@@ -833,7 +1081,7 @@ func TestSerializeMessages_StripsSystemParts(t *testing.T) {
},
},
}
- result := serializeMessages(messages)
+ result := common.SerializeMessages(messages)
data, _ := json.Marshal(result)
raw := string(data)
diff --git a/pkg/providers/types.go b/pkg/providers/types.go
index 68bbd1e65..9a4d126a7 100644
--- a/pkg/providers/types.go
+++ b/pkg/providers/types.go
@@ -37,6 +37,20 @@ type StatefulProvider interface {
Close()
}
+// StreamingProvider is an optional interface for providers that support token streaming.
+// onChunk receives the accumulated text so far (not individual deltas).
+// The returned LLMResponse is the same complete response for compatibility with tool-call handling.
+type StreamingProvider interface {
+ ChatStream(
+ ctx context.Context,
+ messages []Message,
+ tools []ToolDefinition,
+ model string,
+ options map[string]any,
+ onChunk func(accumulated string),
+ ) (*LLMResponse, error)
+}
+
// ThinkingCapable is an optional interface for providers that support
// extended thinking (e.g. Anthropic). Used by the agent loop to warn
// when thinking_level is configured but the active provider cannot use it.
@@ -44,6 +58,15 @@ type ThinkingCapable interface {
SupportsThinking() bool
}
+// NativeSearchCapable is an optional interface for providers that support
+// built-in web search during LLM inference (e.g. OpenAI web_search_preview,
+// xAI Grok search). When the active provider implements this interface and
+// returns true, the agent loop can hide the client-side web_search tool to
+// avoid duplicate search surfaces and use the provider's native search instead.
+type NativeSearchCapable interface {
+ SupportsNativeSearch() bool
+}
+
// FailoverReason classifies why an LLM request failed for fallback decisions.
type FailoverReason string
diff --git a/pkg/tools/cron.go b/pkg/tools/cron.go
index 648cc3c6c..154ec75f0 100644
--- a/pkg/tools/cron.go
+++ b/pkg/tools/cron.go
@@ -20,10 +20,12 @@ type JobExecutor interface {
// CronTool provides scheduling capabilities for the agent
type CronTool struct {
- cronService *cron.CronService
- executor JobExecutor
- msgBus *bus.MessageBus
- execTool *ExecTool
+ cronService *cron.CronService
+ executor JobExecutor
+ msgBus *bus.MessageBus
+ execTool *ExecTool
+ allowCommand bool
+ execEnabled bool
}
// NewCronTool creates a new CronTool
@@ -32,17 +34,32 @@ func NewCronTool(
cronService *cron.CronService, executor JobExecutor, msgBus *bus.MessageBus, workspace string, restrict bool,
execTimeout time.Duration, config *config.Config,
) (*CronTool, error) {
- execTool, err := NewExecToolWithConfig(workspace, restrict, config)
- if err != nil {
- return nil, fmt.Errorf("unable to configure exec tool: %w", err)
+ allowCommand := true
+ execEnabled := true
+ if config != nil {
+ allowCommand = config.Tools.Cron.AllowCommand
+ execEnabled = config.Tools.Exec.Enabled
}
- execTool.SetTimeout(execTimeout)
+ var execTool *ExecTool
+ if execEnabled {
+ var err error
+ execTool, err = NewExecToolWithConfig(workspace, restrict, config)
+ if err != nil {
+ return nil, fmt.Errorf("unable to configure exec tool: %w", err)
+ }
+ }
+
+ if execTool != nil {
+ execTool.SetTimeout(execTimeout)
+ }
return &CronTool{
- cronService: cronService,
- executor: executor,
- msgBus: msgBus,
- execTool: execTool,
+ cronService: cronService,
+ executor: executor,
+ msgBus: msgBus,
+ execTool: execTool,
+ allowCommand: allowCommand,
+ execEnabled: execEnabled,
}, nil
}
@@ -76,7 +93,7 @@ func (t *CronTool) Parameters() map[string]any {
},
"command_confirm": map[string]any{
"type": "boolean",
- "description": "Required when using command=true. Must be true to explicitly confirm scheduling a shell command.",
+ "description": "Optional explicit confirmation flag for scheduling a shell command. Command execution must also be enabled via tools.cron.allow_command.",
},
"at_seconds": map[string]any{
"type": "integer",
@@ -96,7 +113,7 @@ func (t *CronTool) Parameters() map[string]any {
},
"deliver": map[string]any{
"type": "boolean",
- "description": "If true, send message directly to channel. If false, let agent process message (for complex tasks). Default: true",
+ "description": "If true, send message directly to channel. If false, let agent process message (for complex tasks). Default: false",
},
},
"required": []string{"action"},
@@ -174,22 +191,26 @@ func (t *CronTool) addJob(ctx context.Context, args map[string]any) *ToolResult
return ErrorResult("one of at_seconds, every_seconds, or cron_expr is required")
}
- // Read deliver parameter, default to true
- deliver := true
+ // Read deliver parameter, default to false so scheduled tasks execute through the agent
+ deliver := false
if d, ok := args["deliver"].(bool); ok {
deliver = d
}
- // GHSA-pv8c-p6jf-3fpp: command scheduling requires internal channel + explicit confirm.
- // Non-command reminders (plain messages) remain open to all channels.
+ // GHSA-pv8c-p6jf-3fpp: command scheduling requires internal channel. When
+ // allow_command is disabled, explicit confirmation is required as an override.
+ // Non-command reminders remain open to all channels.
command, _ := args["command"].(string)
commandConfirm, _ := args["command_confirm"].(bool)
if command != "" {
+ if !t.execEnabled {
+ return ErrorResult("command execution is disabled")
+ }
if !constants.IsInternalChannel(channel) {
return ErrorResult("scheduling command execution is restricted to internal channels")
}
- if !commandConfirm {
- return ErrorResult("command_confirm=true is required to schedule command execution")
+ if !t.allowCommand && !commandConfirm {
+ return ErrorResult("command_confirm=true is required when allow_command is disabled")
}
deliver = false
}
@@ -290,6 +311,18 @@ func (t *CronTool) ExecuteJob(ctx context.Context, job *cron.CronJob) string {
// Execute command if present
if job.Payload.Command != "" {
+ if !t.execEnabled || t.execTool == nil {
+ output := "Error executing scheduled command: command execution is disabled"
+ pubCtx, pubCancel := context.WithTimeout(context.Background(), 5*time.Second)
+ defer pubCancel()
+ t.msgBus.PublishOutbound(pubCtx, bus.OutboundMessage{
+ Channel: channel,
+ ChatID: chatID,
+ Content: output,
+ })
+ return "ok"
+ }
+
args := map[string]any{
"command": job.Payload.Command,
"__channel": channel,
diff --git a/pkg/tools/cron_test.go b/pkg/tools/cron_test.go
index 1776abc65..cd7d39860 100644
--- a/pkg/tools/cron_test.go
+++ b/pkg/tools/cron_test.go
@@ -5,18 +5,18 @@ import (
"path/filepath"
"strings"
"testing"
+ "time"
"github.com/sipeed/picoclaw/pkg/bus"
"github.com/sipeed/picoclaw/pkg/config"
"github.com/sipeed/picoclaw/pkg/cron"
)
-func newTestCronTool(t *testing.T) *CronTool {
+func newTestCronToolWithConfig(t *testing.T, cfg *config.Config) *CronTool {
t.Helper()
storePath := filepath.Join(t.TempDir(), "cron.json")
cronService := cron.NewCronService(storePath, nil)
msgBus := bus.NewMessageBus()
- cfg := config.DefaultConfig()
tool, err := NewCronTool(cronService, nil, msgBus, t.TempDir(), true, 0, cfg)
if err != nil {
t.Fatalf("NewCronTool() error: %v", err)
@@ -24,6 +24,11 @@ func newTestCronTool(t *testing.T) *CronTool {
return tool
}
+func newTestCronTool(t *testing.T) *CronTool {
+ t.Helper()
+ return newTestCronToolWithConfig(t, config.DefaultConfig())
+}
+
// TestCronTool_CommandBlockedFromRemoteChannel verifies command scheduling is restricted to internal channels
func TestCronTool_CommandBlockedFromRemoteChannel(t *testing.T) {
tool := newTestCronTool(t)
@@ -44,8 +49,7 @@ func TestCronTool_CommandBlockedFromRemoteChannel(t *testing.T) {
}
}
-// TestCronTool_CommandRequiresConfirm verifies command_confirm=true is required
-func TestCronTool_CommandRequiresConfirm(t *testing.T) {
+func TestCronTool_CommandDoesNotRequireConfirmByDefault(t *testing.T) {
tool := newTestCronTool(t)
ctx := WithToolContext(context.Background(), "cli", "direct")
result := tool.Execute(ctx, map[string]any{
@@ -55,11 +59,79 @@ func TestCronTool_CommandRequiresConfirm(t *testing.T) {
"at_seconds": float64(60),
})
+ if result.IsError {
+ t.Fatalf("expected command scheduling without confirm to succeed by default, got: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "Cron job added") {
+ t.Errorf("expected 'Cron job added', got: %s", result.ForLLM)
+ }
+}
+
+func TestCronTool_CommandRequiresConfirmWhenAllowCommandDisabled(t *testing.T) {
+ cfg := config.DefaultConfig()
+ cfg.Tools.Cron.AllowCommand = false
+
+ tool := newTestCronToolWithConfig(t, cfg)
+ ctx := WithToolContext(context.Background(), "cli", "direct")
+ result := tool.Execute(ctx, map[string]any{
+ "action": "add",
+ "message": "check disk",
+ "command": "df -h",
+ "at_seconds": float64(60),
+ })
+
if !result.IsError {
- t.Fatal("expected error when command_confirm is missing")
+ t.Fatal("expected command scheduling to require confirm when allow_command is disabled")
}
if !strings.Contains(result.ForLLM, "command_confirm=true") {
- t.Errorf("expected 'command_confirm=true' message, got: %s", result.ForLLM)
+ t.Errorf("expected command_confirm requirement message, got: %s", result.ForLLM)
+ }
+}
+
+func TestCronTool_CommandAllowedWithConfirmWhenAllowCommandDisabled(t *testing.T) {
+ cfg := config.DefaultConfig()
+ cfg.Tools.Cron.AllowCommand = false
+
+ tool := newTestCronToolWithConfig(t, cfg)
+ ctx := WithToolContext(context.Background(), "cli", "direct")
+ result := tool.Execute(ctx, map[string]any{
+ "action": "add",
+ "message": "check disk",
+ "command": "df -h",
+ "command_confirm": true,
+ "at_seconds": float64(60),
+ })
+
+ if result.IsError {
+ t.Fatalf(
+ "expected command scheduling with confirm to succeed when allow_command is disabled, got: %s",
+ result.ForLLM,
+ )
+ }
+ if !strings.Contains(result.ForLLM, "Cron job added") {
+ t.Errorf("expected 'Cron job added', got: %s", result.ForLLM)
+ }
+}
+
+func TestCronTool_CommandBlockedWhenExecDisabled(t *testing.T) {
+ cfg := config.DefaultConfig()
+ cfg.Tools.Exec.Enabled = false
+
+ tool := newTestCronToolWithConfig(t, cfg)
+ ctx := WithToolContext(context.Background(), "cli", "direct")
+ result := tool.Execute(ctx, map[string]any{
+ "action": "add",
+ "message": "check disk",
+ "command": "df -h",
+ "command_confirm": true,
+ "at_seconds": float64(60),
+ })
+
+ if !result.IsError {
+ t.Fatal("expected command scheduling to be blocked when exec is disabled")
+ }
+ if !strings.Contains(result.ForLLM, "command execution is disabled") {
+ t.Errorf("expected exec disabled message, got: %s", result.ForLLM)
}
}
@@ -114,3 +186,54 @@ func TestCronTool_NonCommandJobAllowedFromRemoteChannel(t *testing.T) {
t.Fatalf("expected non-command reminder to succeed from remote channel, got: %s", result.ForLLM)
}
}
+
+func TestCronTool_NonCommandJobDefaultsDeliverToFalse(t *testing.T) {
+ tool := newTestCronTool(t)
+ ctx := WithToolContext(context.Background(), "telegram", "chat-1")
+ result := tool.Execute(ctx, map[string]any{
+ "action": "add",
+ "message": "send me a poem",
+ "at_seconds": float64(600),
+ })
+
+ if result.IsError {
+ t.Fatalf("expected non-command reminder to succeed, got: %s", result.ForLLM)
+ }
+
+ jobs := tool.cronService.ListJobs(false)
+ if len(jobs) != 1 {
+ t.Fatalf("expected 1 job, got %d", len(jobs))
+ }
+ if jobs[0].Payload.Deliver {
+ t.Fatal("expected deliver=false by default for non-command jobs")
+ }
+}
+
+func TestCronTool_ExecuteJobPublishesErrorWhenExecDisabled(t *testing.T) {
+ cfg := config.DefaultConfig()
+ cfg.Tools.Exec.Enabled = false
+
+ tool := newTestCronToolWithConfig(t, cfg)
+ job := &cron.CronJob{}
+ job.Payload.Channel = "cli"
+ job.Payload.To = "direct"
+ job.Payload.Command = "df -h"
+
+ if got := tool.ExecuteJob(context.Background(), job); got != "ok" {
+ t.Fatalf("ExecuteJob() = %q, want ok", got)
+ }
+
+ ctx, cancel := context.WithTimeout(context.Background(), time.Second)
+ defer cancel()
+
+ var msg bus.OutboundMessage
+ select {
+ case msg = <-tool.msgBus.OutboundChan():
+ // got message
+ case <-ctx.Done():
+ t.Fatal("timeout waiting for outbound message")
+ }
+ if !strings.Contains(msg.Content, "command execution is disabled") {
+ t.Fatalf("expected exec disabled message, got: %s", msg.Content)
+ }
+}
diff --git a/pkg/tools/filesystem.go b/pkg/tools/filesystem.go
index 6b1cb1475..39d45013d 100644
--- a/pkg/tools/filesystem.go
+++ b/pkg/tools/filesystem.go
@@ -20,8 +20,7 @@ import (
const MaxReadFileSize = 64 * 1024 // 64KB limit to avoid context overflow
-// validatePath ensures the given path is within the workspace if restrict is true.
-func validatePath(path, workspace string, restrict bool) (string, error) {
+func validatePathWithAllowPaths(path, workspace string, restrict bool, patterns []*regexp.Regexp) (string, error) {
if workspace == "" {
return path, fmt.Errorf("workspace is not defined")
}
@@ -42,6 +41,10 @@ func validatePath(path, workspace string, restrict bool) (string, error) {
}
if restrict {
+ if isAllowedPath(absPath, patterns) {
+ return absPath, nil
+ }
+
if !isWithinWorkspace(absPath, absWorkspace) {
return "", fmt.Errorf("access denied: path is outside the workspace")
}
@@ -73,6 +76,137 @@ func validatePath(path, workspace string, restrict bool) (string, error) {
return absPath, nil
}
+func isAllowedPath(path string, patterns []*regexp.Regexp) bool {
+ if len(patterns) == 0 {
+ return false
+ }
+
+ cleaned := filepath.Clean(path)
+ if !filepath.IsAbs(cleaned) {
+ return false
+ }
+ if !matchesAllowedPath(cleaned, patterns) {
+ return false
+ }
+
+ resolved, err := resolvePathAgainstExistingAncestor(cleaned)
+ if err != nil {
+ return false
+ }
+
+ return matchesAllowedPath(resolved, patterns)
+}
+
+func matchesAllowedPath(path string, patterns []*regexp.Regexp) bool {
+ cleaned := filepath.Clean(path)
+ for _, pattern := range patterns {
+ if pattern.MatchString(cleaned) {
+ return true
+ }
+ if root, ok := extractAllowedPathRoot(pattern); ok && isWithinAllowedRoot(cleaned, root) {
+ return true
+ }
+ }
+ return false
+}
+
+func extractAllowedPathRoot(pattern *regexp.Regexp) (string, bool) {
+ raw := pattern.String()
+ if !strings.HasPrefix(raw, "^") {
+ return "", false
+ }
+
+ literal := strings.TrimPrefix(raw, "^")
+
+ // Recognize the common "directory prefix" form: ^(?:/|$)
+ literal = strings.TrimSuffix(literal, "(?:/|$)")
+ literal = strings.TrimSuffix(literal, `(?:\\|$)`)
+
+ // Reject patterns that still contain regex operators after removing the
+ // optional anchored-directory suffix. That keeps arbitrary regex behavior
+ // unchanged and only enables normalized prefix matching for literal paths.
+ if containsUnescapedRegexMeta(literal) {
+ return "", false
+ }
+
+ unescaped, ok := unescapeRegexLiteral(literal)
+ if !ok || unescaped == "" {
+ return "", false
+ }
+
+ return filepath.Clean(unescaped), filepath.IsAbs(unescaped)
+}
+
+func appendUniquePath(paths []string, path string) []string {
+ for _, existing := range paths {
+ if existing == path {
+ return paths
+ }
+ }
+ return append(paths, path)
+}
+
+func containsUnescapedRegexMeta(s string) bool {
+ escaped := false
+ for _, r := range s {
+ if escaped {
+ escaped = false
+ continue
+ }
+ if r == '\\' {
+ escaped = true
+ continue
+ }
+ switch r {
+ case '.', '+', '*', '?', '(', ')', '[', ']', '{', '}', '|':
+ return true
+ }
+ }
+ return escaped
+}
+
+func unescapeRegexLiteral(s string) (string, bool) {
+ var b strings.Builder
+ b.Grow(len(s))
+
+ escaped := false
+ for _, r := range s {
+ if escaped {
+ b.WriteRune(r)
+ escaped = false
+ continue
+ }
+ if r == '\\' {
+ escaped = true
+ continue
+ }
+ b.WriteRune(r)
+ }
+
+ if escaped {
+ return "", false
+ }
+
+ return b.String(), true
+}
+
+func isWithinAllowedRoot(path, root string) bool {
+ candidate := filepath.Clean(path)
+ allowedVariants := []string{filepath.Clean(root)}
+
+ if resolvedRoot, err := resolvePathAgainstExistingAncestor(root); err == nil {
+ allowedVariants = appendUniquePath(allowedVariants, filepath.Clean(resolvedRoot))
+ }
+
+ for _, allowedRoot := range allowedVariants {
+ if isWithinWorkspace(candidate, allowedRoot) {
+ return true
+ }
+ }
+
+ return false
+}
+
func resolveExistingAncestor(path string) (string, error) {
for current := filepath.Clean(path); ; current = filepath.Dir(current) {
if resolved, err := filepath.EvalSymlinks(current); err == nil {
@@ -86,9 +220,32 @@ func resolveExistingAncestor(path string) (string, error) {
}
}
+func resolvePathAgainstExistingAncestor(path string) (string, error) {
+ cleaned := filepath.Clean(path)
+ for current := cleaned; ; current = filepath.Dir(current) {
+ resolved, err := filepath.EvalSymlinks(current)
+ if err == nil {
+ suffix, relErr := filepath.Rel(current, cleaned)
+ if relErr != nil {
+ return "", relErr
+ }
+ if suffix == "." {
+ return filepath.Clean(resolved), nil
+ }
+ return filepath.Clean(filepath.Join(resolved, suffix)), nil
+ }
+ if !os.IsNotExist(err) {
+ return "", err
+ }
+ if filepath.Dir(current) == current {
+ return "", os.ErrNotExist
+ }
+ }
+}
+
func isWithinWorkspace(candidate, workspace string) bool {
rel, err := filepath.Rel(filepath.Clean(workspace), filepath.Clean(candidate))
- return err == nil && filepath.IsLocal(rel)
+ return err == nil && (rel == "." || filepath.IsLocal(rel))
}
type ReadFileTool struct {
@@ -339,7 +496,7 @@ func (t *WriteFileTool) Name() string {
}
func (t *WriteFileTool) Description() string {
- return "Write content to a file"
+ return "Write content to a file. If the file already exists, you must set overwrite=true to replace it."
}
func (t *WriteFileTool) Parameters() map[string]any {
@@ -354,6 +511,11 @@ func (t *WriteFileTool) Parameters() map[string]any {
"type": "string",
"description": "Content to write to the file",
},
+ "overwrite": map[string]any{
+ "type": "boolean",
+ "description": "Must be set to true to overwrite an existing file.",
+ "default": false,
+ },
},
"required": []string{"path", "content"},
}
@@ -370,6 +532,14 @@ func (t *WriteFileTool) Execute(ctx context.Context, args map[string]any) *ToolR
return ErrorResult("content is required")
}
+ overwrite, _ := args["overwrite"].(bool)
+
+ if !overwrite {
+ if _, err := t.fs.Open(path); err == nil {
+ return ErrorResult(fmt.Sprintf("file: %s already exists. Set overwrite=true to replace.", path))
+ }
+ }
+
if err := t.fs.WriteFile(path, []byte(content)); err != nil {
return ErrorResult(err.Error())
}
@@ -625,12 +795,7 @@ type whitelistFs struct {
}
func (w *whitelistFs) matches(path string) bool {
- for _, p := range w.patterns {
- if p.MatchString(path) {
- return true
- }
- }
- return false
+ return isAllowedPath(path, w.patterns)
}
func (w *whitelistFs) ReadFile(path string) ([]byte, error) {
diff --git a/pkg/tools/filesystem_test.go b/pkg/tools/filesystem_test.go
index 0bbf6caf0..0b4dd310b 100644
--- a/pkg/tools/filesystem_test.go
+++ b/pkg/tools/filesystem_test.go
@@ -189,6 +189,121 @@ func TestFilesystemTool_WriteFile_MissingContent(t *testing.T) {
}
}
+// TestFilesystemTool_WriteFile_OverwriteDefaultBlocked verifies that writing to an
+// existing file without overwrite=true returns an error.
+func TestFilesystemTool_WriteFile_OverwriteDefaultBlocked(t *testing.T) {
+ tmpDir := t.TempDir()
+ testFile := filepath.Join(tmpDir, "existing.txt")
+ os.WriteFile(testFile, []byte("original"), 0o644)
+
+ tool := NewWriteFileTool("", false)
+ result := tool.Execute(context.Background(), map[string]any{
+ "path": testFile,
+ "content": "new content",
+ })
+
+ assert.True(t, result.IsError, "expected error when overwriting without overwrite=true")
+ assert.Contains(t, result.ForLLM, "already exists")
+ assert.Contains(t, result.ForLLM, "overwrite=true")
+
+ // Original content must be untouched
+ data, err := os.ReadFile(testFile)
+ assert.NoError(t, err)
+ assert.Equal(t, "original", string(data))
+}
+
+// TestFilesystemTool_WriteFile_OverwriteExplicitAllowed verifies that setting
+// overwrite=true replaces the existing file.
+func TestFilesystemTool_WriteFile_OverwriteExplicitAllowed(t *testing.T) {
+ tmpDir := t.TempDir()
+ testFile := filepath.Join(tmpDir, "existing.txt")
+ os.WriteFile(testFile, []byte("original"), 0o644)
+
+ tool := NewWriteFileTool("", false)
+ result := tool.Execute(context.Background(), map[string]any{
+ "path": testFile,
+ "content": "replaced",
+ "overwrite": true,
+ })
+
+ assert.False(t, result.IsError, "expected success with overwrite=true, got: %s", result.ForLLM)
+
+ data, err := os.ReadFile(testFile)
+ assert.NoError(t, err)
+ assert.Equal(t, "replaced", string(data))
+}
+
+// TestFilesystemTool_WriteFile_NewFileNoOverwriteFlag verifies that a new (non-existing)
+// file can be written without setting overwrite=true.
+func TestFilesystemTool_WriteFile_NewFileNoOverwriteFlag(t *testing.T) {
+ tmpDir := t.TempDir()
+ testFile := filepath.Join(tmpDir, "newfile.txt")
+
+ tool := NewWriteFileTool("", false)
+ result := tool.Execute(context.Background(), map[string]any{
+ "path": testFile,
+ "content": "brand new",
+ })
+
+ assert.False(t, result.IsError, "expected success for new file, got: %s", result.ForLLM)
+
+ data, err := os.ReadFile(testFile)
+ assert.NoError(t, err)
+ assert.Equal(t, "brand new", string(data))
+}
+
+// TestFilesystemTool_WriteFile_OverwriteFalseExplicitBlocked verifies that
+// explicitly passing overwrite=false also blocks overwriting.
+func TestFilesystemTool_WriteFile_OverwriteFalseExplicitBlocked(t *testing.T) {
+ tmpDir := t.TempDir()
+ testFile := filepath.Join(tmpDir, "existing.txt")
+ os.WriteFile(testFile, []byte("original"), 0o644)
+
+ tool := NewWriteFileTool("", false)
+ result := tool.Execute(context.Background(), map[string]any{
+ "path": testFile,
+ "content": "new content",
+ "overwrite": false,
+ })
+
+ assert.True(t, result.IsError, "expected error when overwrite=false")
+ assert.Contains(t, result.ForLLM, "already exists")
+
+ data, err := os.ReadFile(testFile)
+ assert.NoError(t, err)
+ assert.Equal(t, "original", string(data))
+}
+
+// TestFilesystemTool_WriteFile_OverwriteSandboxed verifies the overwrite guard
+// works correctly in restricted (sandbox) mode.
+func TestFilesystemTool_WriteFile_OverwriteSandboxed(t *testing.T) {
+ workspace := t.TempDir()
+ testFile := "file.txt"
+ os.WriteFile(filepath.Join(workspace, testFile), []byte("original"), 0o644)
+
+ tool := NewWriteFileTool(workspace, true)
+
+ // Without overwrite=true → blocked
+ result := tool.Execute(context.Background(), map[string]any{
+ "path": testFile,
+ "content": "new content",
+ })
+ assert.True(t, result.IsError, "expected error in sandbox mode without overwrite=true")
+ assert.Contains(t, result.ForLLM, "already exists")
+
+ // With overwrite=true → allowed
+ result = tool.Execute(context.Background(), map[string]any{
+ "path": testFile,
+ "content": "replaced in sandbox",
+ "overwrite": true,
+ })
+ assert.False(t, result.IsError, "expected success in sandbox mode with overwrite=true, got: %s", result.ForLLM)
+
+ data, err := os.ReadFile(filepath.Join(workspace, testFile))
+ assert.NoError(t, err)
+ assert.Equal(t, "replaced in sandbox", string(data))
+}
+
// TestFilesystemTool_ListDir_Success verifies successful directory listing
func TestFilesystemTool_ListDir_Success(t *testing.T) {
tmpDir := t.TempDir()
@@ -521,6 +636,90 @@ func TestWhitelistFs_AllowsMatchingPaths(t *testing.T) {
}
}
+func TestWhitelistFs_BlocksSymlinkEscapeInAllowedDir(t *testing.T) {
+ workspace := t.TempDir()
+ allowedDir := t.TempDir()
+ secretDir := t.TempDir()
+ secretFile := filepath.Join(secretDir, "secret.txt")
+ if err := os.WriteFile(secretFile, []byte("top secret"), 0o644); err != nil {
+ t.Fatalf("WriteFile(secretFile) error = %v", err)
+ }
+
+ linkPath := filepath.Join(allowedDir, "link_out")
+ if err := os.Symlink(secretDir, linkPath); err != nil {
+ t.Skipf("symlink not supported in this environment: %v", err)
+ }
+
+ patterns := []*regexp.Regexp{regexp.MustCompile(`^` + regexp.QuoteMeta(allowedDir))}
+ tool := NewReadFileTool(workspace, true, MaxReadFileSize, patterns)
+
+ result := tool.Execute(context.Background(), map[string]any{"path": filepath.Join(linkPath, "secret.txt")})
+ if !result.IsError {
+ t.Fatalf("expected symlink escape from allowed dir to be blocked, got: %s", result.ForLLM)
+ }
+}
+
+func TestWhitelistFs_WriteAllowsNewFileUnderAllowedDir(t *testing.T) {
+ workspace := t.TempDir()
+ rootDir := t.TempDir()
+ allowedDir := filepath.Join(rootDir, "allowed")
+ targetFile := filepath.Join(allowedDir, "nested", "file.txt")
+
+ patterns := []*regexp.Regexp{regexp.MustCompile(`^` + regexp.QuoteMeta(allowedDir))}
+ tool := NewWriteFileTool(workspace, true, patterns)
+
+ result := tool.Execute(context.Background(), map[string]any{
+ "path": targetFile,
+ "content": "outside write",
+ })
+ if result.IsError {
+ t.Fatalf("expected whitelisted write to succeed, got: %s", result.ForLLM)
+ }
+
+ data, err := os.ReadFile(targetFile)
+ if err != nil {
+ t.Fatalf("ReadFile(targetFile) error = %v", err)
+ }
+ if string(data) != "outside write" {
+ t.Fatalf("target file content = %q, want %q", string(data), "outside write")
+ }
+}
+
+func TestWhitelistFs_AllowsResolvedAllowedRootAlias(t *testing.T) {
+ workspace := t.TempDir()
+ realDir := t.TempDir()
+ linkParent := t.TempDir()
+ allowedAlias := filepath.Join(linkParent, "allowed-link")
+
+ if err := os.Symlink(realDir, allowedAlias); err != nil {
+ t.Skipf("symlink not supported in this environment: %v", err)
+ }
+
+ targetFile := filepath.Join(allowedAlias, "nested", "alias.txt")
+ if err := os.MkdirAll(filepath.Dir(targetFile), 0o755); err != nil {
+ t.Fatalf("MkdirAll(targetFile dir) error = %v", err)
+ }
+ if err := os.WriteFile(targetFile, []byte("through alias"), 0o644); err != nil {
+ t.Fatalf("WriteFile(targetFile) error = %v", err)
+ }
+
+ patterns := []*regexp.Regexp{
+ regexp.MustCompile(
+ "^" + regexp.QuoteMeta(filepath.Clean(allowedAlias)) +
+ "(?:" + regexp.QuoteMeta(string(os.PathSeparator)) + "|$)",
+ ),
+ }
+ tool := NewReadFileTool(workspace, true, MaxReadFileSize, patterns)
+
+ result := tool.Execute(context.Background(), map[string]any{"path": targetFile})
+ if result.IsError {
+ t.Fatalf("expected symlink-backed allowed root to be readable, got: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "through alias") {
+ t.Fatalf("expected file content, got: %s", result.ForLLM)
+ }
+}
+
// TestReadFileTool_ChunkedReading verifies the pagination logic of the tool
// by reading a file in multiple chunks using 'offset' and 'length'.
func TestReadFileTool_ChunkedReading(t *testing.T) {
diff --git a/pkg/tools/registry.go b/pkg/tools/registry.go
index 0635f47d7..ed373a28f 100644
--- a/pkg/tools/registry.go
+++ b/pkg/tools/registry.go
@@ -188,15 +188,48 @@ func (r *ToolRegistry) ExecuteWithContext(
// The callback is a call parameter, not mutable state on the tool instance.
var result *ToolResult
start := time.Now()
- if asyncExec, ok := tool.(AsyncExecutor); ok && asyncCallback != nil {
- logger.DebugCF("tool", "Executing async tool via ExecuteAsync",
- map[string]any{
- "tool": name,
- })
- result = asyncExec.ExecuteAsync(ctx, args, asyncCallback)
- } else {
- result = tool.Execute(ctx, args)
+
+ // Use recover to catch any panics during tool execution
+ // This prevents tool crashes from killing the entire agent
+ func() {
+ defer func() {
+ if re := recover(); re != nil {
+ errMsg := fmt.Sprintf("Tool '%s' crashed with panic: %v", name, re)
+ logger.ErrorCF("tool", "Tool execution panic recovered",
+ map[string]any{
+ "tool": name,
+ "panic": fmt.Sprintf("%v", re),
+ })
+ result = &ToolResult{
+ ForLLM: errMsg,
+ ForUser: errMsg,
+ IsError: true,
+ Err: fmt.Errorf("panic: %v", re),
+ }
+ }
+ }()
+
+ if asyncExec, ok := tool.(AsyncExecutor); ok && asyncCallback != nil {
+ logger.DebugCF("tool", "Executing async tool via ExecuteAsync",
+ map[string]any{
+ "tool": name,
+ })
+ result = asyncExec.ExecuteAsync(ctx, args, asyncCallback)
+ } else {
+ result = tool.Execute(ctx, args)
+ }
+ }()
+
+ // Handle nil result (should not happen, but defensive)
+ if result == nil {
+ result = &ToolResult{
+ ForLLM: fmt.Sprintf("Tool '%s' returned nil result unexpectedly", name),
+ ForUser: fmt.Sprintf("Tool '%s' returned nil result unexpectedly", name),
+ IsError: true,
+ Err: fmt.Errorf("nil result from tool"),
+ }
}
+
duration := time.Since(start)
// Log based on result type
@@ -303,6 +336,28 @@ func (r *ToolRegistry) List() []string {
return r.sortedToolNames()
}
+// Clone creates an independent copy of the registry containing the same tool
+// entries (shallow copy of each ToolEntry). This is used to give subagents a
+// snapshot of the parent agent's tools without sharing the same registry —
+// tools registered on the parent after cloning (e.g. spawn, spawn_status)
+// will NOT be visible to the clone, preventing recursive subagent spawning.
+// The version counter is reset to 0 in the clone as it's a new independent registry.
+func (r *ToolRegistry) Clone() *ToolRegistry {
+ r.mu.RLock()
+ defer r.mu.RUnlock()
+ clone := &ToolRegistry{
+ tools: make(map[string]*ToolEntry, len(r.tools)),
+ }
+ for name, entry := range r.tools {
+ clone.tools[name] = &ToolEntry{
+ Tool: entry.Tool,
+ IsCore: entry.IsCore,
+ TTL: entry.TTL,
+ }
+ }
+ return clone
+}
+
// Count returns the number of registered tools.
func (r *ToolRegistry) Count() int {
r.mu.RLock()
@@ -329,3 +384,22 @@ func (r *ToolRegistry) GetSummaries() []string {
}
return summaries
}
+
+// GetAll returns all registered tools (both core and non-core with TTL > 0).
+// Used by SubTurn to inherit parent's tool set.
+func (r *ToolRegistry) GetAll() []Tool {
+ r.mu.RLock()
+ defer r.mu.RUnlock()
+
+ sorted := r.sortedToolNames()
+ tools := make([]Tool, 0, len(sorted))
+ for _, name := range sorted {
+ entry := r.tools[name]
+
+ // Include core tools and non-core tools with active TTL
+ if entry.IsCore || entry.TTL > 0 {
+ tools = append(tools, entry.Tool)
+ }
+ }
+ return tools
+}
diff --git a/pkg/tools/registry_test.go b/pkg/tools/registry_test.go
index 92d7d5abd..967758dfa 100644
--- a/pkg/tools/registry_test.go
+++ b/pkg/tools/registry_test.go
@@ -2,6 +2,7 @@ package tools
import (
"context"
+ "errors"
"strings"
"sync"
"testing"
@@ -335,6 +336,96 @@ func TestToolToSchema(t *testing.T) {
}
}
+func TestToolRegistry_Clone(t *testing.T) {
+ r := NewToolRegistry()
+ r.Register(newMockTool("read_file", "reads files"))
+ r.Register(newMockTool("exec", "runs commands"))
+ r.Register(newMockTool("web_search", "searches the web"))
+
+ clone := r.Clone()
+
+ // Clone should have the same tools
+ if clone.Count() != 3 {
+ t.Errorf("expected clone to have 3 tools, got %d", clone.Count())
+ }
+ for _, name := range []string{"read_file", "exec", "web_search"} {
+ if _, ok := clone.Get(name); !ok {
+ t.Errorf("expected clone to have tool %q", name)
+ }
+ }
+
+ // Registering on parent should NOT affect clone
+ r.Register(newMockTool("spawn", "spawns subagent"))
+ if r.Count() != 4 {
+ t.Errorf("expected parent to have 4 tools, got %d", r.Count())
+ }
+ if clone.Count() != 3 {
+ t.Errorf("expected clone to still have 3 tools after parent mutation, got %d", clone.Count())
+ }
+ if _, ok := clone.Get("spawn"); ok {
+ t.Error("expected clone NOT to have 'spawn' tool registered on parent after cloning")
+ }
+
+ // Registering on clone should NOT affect parent
+ clone.Register(newMockTool("custom", "custom tool"))
+ if clone.Count() != 4 {
+ t.Errorf("expected clone to have 4 tools, got %d", clone.Count())
+ }
+ if _, ok := r.Get("custom"); ok {
+ t.Error("expected parent NOT to have 'custom' tool registered on clone")
+ }
+}
+
+func TestToolRegistry_Clone_Empty(t *testing.T) {
+ r := NewToolRegistry()
+ clone := r.Clone()
+ if clone.Count() != 0 {
+ t.Errorf("expected empty clone, got count %d", clone.Count())
+ }
+}
+
+func TestToolRegistry_Clone_PreservesHiddenToolState(t *testing.T) {
+ r := NewToolRegistry()
+ r.RegisterHidden(newMockTool("mcp_tool", "dynamic MCP tool"))
+
+ clone := r.Clone()
+
+ // Hidden tools with TTL=0 should not be gettable (same behavior as parent)
+ if _, ok := clone.Get("mcp_tool"); ok {
+ t.Error("expected hidden tool with TTL=0 to be invisible in clone")
+ }
+
+ // But the entry should exist (count includes hidden tools)
+ if clone.Count() != 1 {
+ t.Errorf("expected clone count 1 (hidden entry exists), got %d", clone.Count())
+ }
+}
+
+func TestToolRegistry_Clone_PreservesTTLValue(t *testing.T) {
+ r := NewToolRegistry()
+ r.RegisterHidden(newMockTool("ttl_tool", "tool with TTL"))
+
+ // Manually set a non-zero TTL on the entry
+ r.mu.RLock()
+ if entry, ok := r.tools["ttl_tool"]; ok {
+ entry.TTL = 5
+ }
+ r.mu.RUnlock()
+
+ clone := r.Clone()
+
+ // Verify TTL value is preserved in the clone
+ clone.mu.RLock()
+ defer clone.mu.RUnlock()
+ entry, ok := clone.tools["ttl_tool"]
+ if !ok {
+ t.Fatal("expected ttl_tool to exist in clone")
+ }
+ if entry.TTL != 5 {
+ t.Errorf("expected TTL=5 in clone, got %d", entry.TTL)
+ }
+}
+
func TestToolRegistry_ConcurrentAccess(t *testing.T) {
r := NewToolRegistry()
var wg sync.WaitGroup
@@ -358,3 +449,175 @@ func TestToolRegistry_ConcurrentAccess(t *testing.T) {
t.Error("expected tools to be registered after concurrent access")
}
}
+
+// --- Panic and abnormal exit tests ---
+
+// mockPanicTool is a tool that panics during execution
+type mockPanicTool struct {
+ name string
+ panicValue any
+}
+
+func (m *mockPanicTool) Name() string { return m.name }
+func (m *mockPanicTool) Description() string { return "a tool that panics" }
+func (m *mockPanicTool) Parameters() map[string]any { return map[string]any{"type": "object"} }
+func (m *mockPanicTool) Execute(_ context.Context, _ map[string]any) *ToolResult {
+ panic(m.panicValue)
+}
+
+// mockNilResultTool is a tool that returns nil
+type mockNilResultTool struct {
+ name string
+}
+
+func (m *mockNilResultTool) Name() string { return m.name }
+func (m *mockNilResultTool) Description() string { return "a tool that returns nil" }
+func (m *mockNilResultTool) Parameters() map[string]any { return map[string]any{"type": "object"} }
+func (m *mockNilResultTool) Execute(_ context.Context, _ map[string]any) *ToolResult {
+ return nil
+}
+
+func TestToolRegistry_Execute_PanicRecovery(t *testing.T) {
+ r := NewToolRegistry()
+ r.Register(&mockPanicTool{
+ name: "panic_tool",
+ panicValue: "something went terribly wrong",
+ })
+
+ // Should not panic, should return error result
+ result := r.Execute(context.Background(), "panic_tool", nil)
+
+ if result == nil {
+ t.Fatal("expected non-nil result after panic recovery")
+ }
+ if !result.IsError {
+ t.Error("expected IsError=true after panic")
+ }
+ if !strings.Contains(result.ForLLM, "panic") {
+ t.Errorf("expected 'panic' in error message, got %q", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "panic_tool") {
+ t.Errorf("expected tool name in error message, got %q", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "something went terribly wrong") {
+ t.Errorf("expected panic value in error message, got %q", result.ForLLM)
+ }
+ if result.Err == nil {
+ t.Error("expected Err to be set")
+ }
+}
+
+func TestToolRegistry_Execute_PanicRecovery_ErrorType(t *testing.T) {
+ r := NewToolRegistry()
+
+ // Test with error type panic
+ r.Register(&mockPanicTool{
+ name: "error_panic_tool",
+ panicValue: errors.New("custom error panic"),
+ })
+
+ result := r.Execute(context.Background(), "error_panic_tool", nil)
+
+ if !result.IsError {
+ t.Error("expected IsError=true")
+ }
+ if !strings.Contains(result.ForLLM, "custom error panic") {
+ t.Errorf("expected error message in ForLLM, got %q", result.ForLLM)
+ }
+}
+
+func TestToolRegistry_Execute_PanicRecovery_IntType(t *testing.T) {
+ r := NewToolRegistry()
+
+ // Test with int type panic
+ r.Register(&mockPanicTool{
+ name: "int_panic_tool",
+ panicValue: 42,
+ })
+
+ result := r.Execute(context.Background(), "int_panic_tool", nil)
+
+ if !result.IsError {
+ t.Error("expected IsError=true")
+ }
+ if !strings.Contains(result.ForLLM, "42") {
+ t.Errorf("expected panic value '42' in ForLLM, got %q", result.ForLLM)
+ }
+}
+
+func TestToolRegistry_Execute_NilResultHandling(t *testing.T) {
+ r := NewToolRegistry()
+ r.Register(&mockNilResultTool{name: "nil_tool"})
+
+ result := r.Execute(context.Background(), "nil_tool", nil)
+
+ if result == nil {
+ t.Fatal("expected non-nil result when tool returns nil")
+ }
+ if !result.IsError {
+ t.Error("expected IsError=true for nil result")
+ }
+ if !strings.Contains(result.ForLLM, "nil_tool") {
+ t.Errorf("expected tool name in error message, got %q", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "nil result") {
+ t.Errorf("expected 'nil result' in error message, got %q", result.ForLLM)
+ }
+ if result.Err == nil {
+ t.Error("expected Err to be set")
+ }
+}
+
+func TestToolRegistry_ExecuteWithContext_PanicRecovery(t *testing.T) {
+ r := NewToolRegistry()
+ r.Register(&mockPanicTool{
+ name: "ctx_panic_tool",
+ panicValue: "context panic test",
+ })
+
+ // Should not panic even with context
+ result := r.ExecuteWithContext(
+ context.Background(),
+ "ctx_panic_tool",
+ map[string]any{"key": "value"},
+ "telegram",
+ "chat-123",
+ nil,
+ )
+
+ if result == nil {
+ t.Fatal("expected non-nil result")
+ }
+ if !result.IsError {
+ t.Error("expected IsError=true")
+ }
+ if !strings.Contains(result.ForLLM, "context panic test") {
+ t.Errorf("expected panic message, got %q", result.ForLLM)
+ }
+}
+
+func TestToolRegistry_Execute_PanicDoesNotAffectOtherTools(t *testing.T) {
+ r := NewToolRegistry()
+ r.Register(&mockPanicTool{name: "bad_tool", panicValue: "boom"})
+ r.Register(&mockRegistryTool{
+ name: "good_tool",
+ desc: "works fine",
+ params: map[string]any{},
+ result: SilentResult("success"),
+ })
+
+ // First, trigger the panic
+ result1 := r.Execute(context.Background(), "bad_tool", nil)
+ if !result1.IsError {
+ t.Error("expected error from panic tool")
+ }
+
+ // Then, verify the good tool still works
+ result2 := r.Execute(context.Background(), "good_tool", nil)
+ if result2.IsError {
+ t.Errorf("expected success from good tool, got error: %s", result2.ForLLM)
+ }
+ if result2.ForLLM != "success" {
+ t.Errorf("expected 'success', got %q", result2.ForLLM)
+ }
+}
diff --git a/pkg/tools/result.go b/pkg/tools/result.go
index cab833284..bf34b7bc6 100644
--- a/pkg/tools/result.go
+++ b/pkg/tools/result.go
@@ -1,6 +1,10 @@
package tools
-import "encoding/json"
+import (
+ "encoding/json"
+
+ "github.com/sipeed/picoclaw/pkg/providers"
+)
// ToolResult represents the structured return value from tool execution.
// It provides clear semantics for different types of results and supports
@@ -34,6 +38,11 @@ type ToolResult struct {
// Media contains media store refs produced by this tool.
// When non-empty, the agent will publish these as OutboundMediaMessage.
Media []string `json:"media,omitempty"`
+
+ // Messages holds the ephemeral session history after execution.
+ // Only populated by SubTurn executions; used by evaluator_optimizer
+ // to carry stateful worker context across evaluation iterations.
+ Messages []providers.Message `json:"-"`
}
// NewToolResult creates a basic ToolResult with content for the LLM.
diff --git a/pkg/tools/send_file.go b/pkg/tools/send_file.go
index 1a03e58ed..a67bd4210 100644
--- a/pkg/tools/send_file.go
+++ b/pkg/tools/send_file.go
@@ -6,6 +6,7 @@ import (
"mime"
"os"
"path/filepath"
+ "regexp"
"strings"
"github.com/h2non/filetype"
@@ -21,20 +22,32 @@ type SendFileTool struct {
restrict bool
maxFileSize int
mediaStore media.MediaStore
+ allowPaths []*regexp.Regexp
defaultChannel string
defaultChatID string
}
-func NewSendFileTool(workspace string, restrict bool, maxFileSize int, store media.MediaStore) *SendFileTool {
+func NewSendFileTool(
+ workspace string,
+ restrict bool,
+ maxFileSize int,
+ store media.MediaStore,
+ allowPaths ...[]*regexp.Regexp,
+) *SendFileTool {
if maxFileSize <= 0 {
maxFileSize = config.DefaultMaxMediaSize
}
+ var patterns []*regexp.Regexp
+ if len(allowPaths) > 0 {
+ patterns = allowPaths[0]
+ }
return &SendFileTool{
workspace: workspace,
restrict: restrict,
maxFileSize: maxFileSize,
mediaStore: store,
+ allowPaths: patterns,
}
}
@@ -92,7 +105,7 @@ func (t *SendFileTool) Execute(ctx context.Context, args map[string]any) *ToolRe
return ErrorResult("media store not configured")
}
- resolved, err := validatePath(path, t.workspace, t.restrict)
+ resolved, err := validatePathWithAllowPaths(path, t.workspace, t.restrict, t.allowPaths)
if err != nil {
return ErrorResult(fmt.Sprintf("invalid path: %v", err))
}
diff --git a/pkg/tools/send_file_test.go b/pkg/tools/send_file_test.go
index 08d129674..6daaab31c 100644
--- a/pkg/tools/send_file_test.go
+++ b/pkg/tools/send_file_test.go
@@ -4,6 +4,7 @@ import (
"context"
"os"
"path/filepath"
+ "regexp"
"strings"
"testing"
@@ -128,6 +129,44 @@ func TestSendFileTool_CustomFilename(t *testing.T) {
}
}
+func TestSendFileTool_AllowsWhitelistedMediaTempPath(t *testing.T) {
+ workspace := t.TempDir()
+ mediaDir := media.TempDir()
+ if err := os.MkdirAll(mediaDir, 0o700); err != nil {
+ t.Fatalf("MkdirAll(mediaDir) error = %v", err)
+ }
+
+ testFile, err := os.CreateTemp(mediaDir, "send-file-*.txt")
+ if err != nil {
+ t.Fatalf("CreateTemp(mediaDir) error = %v", err)
+ }
+ testPath := testFile.Name()
+ if _, err := testFile.WriteString("forward me"); err != nil {
+ testFile.Close()
+ t.Fatalf("WriteString(testFile) error = %v", err)
+ }
+ if err := testFile.Close(); err != nil {
+ t.Fatalf("Close(testFile) error = %v", err)
+ }
+ t.Cleanup(func() { _ = os.Remove(testPath) })
+
+ pattern := regexp.MustCompile(
+ "^" + regexp.QuoteMeta(filepath.Clean(mediaDir)) + "(?:" + regexp.QuoteMeta(string(os.PathSeparator)) + "|$)",
+ )
+
+ store := media.NewFileMediaStore()
+ tool := NewSendFileTool(workspace, true, 0, store, []*regexp.Regexp{pattern})
+ tool.SetContext("feishu", "chat123")
+
+ result := tool.Execute(context.Background(), map[string]any{"path": testPath})
+ if result.IsError {
+ t.Fatalf("expected whitelisted temp media file to be sendable, got: %s", result.ForLLM)
+ }
+ if len(result.Media) != 1 {
+ t.Fatalf("expected 1 media ref, got %d", len(result.Media))
+ }
+}
+
func TestDetectMediaType_MagicBytes(t *testing.T) {
dir := t.TempDir()
diff --git a/pkg/tools/shell.go b/pkg/tools/shell.go
index 9ea05bb12..78ad2b26d 100644
--- a/pkg/tools/shell.go
+++ b/pkg/tools/shell.go
@@ -23,6 +23,7 @@ type ExecTool struct {
denyPatterns []*regexp.Regexp
allowPatterns []*regexp.Regexp
customAllowPatterns []*regexp.Regexp
+ allowedPathPatterns []*regexp.Regexp
restrictToWorkspace bool
allowRemote bool
}
@@ -95,14 +96,23 @@ var (
}
)
-func NewExecTool(workingDir string, restrict bool) (*ExecTool, error) {
- return NewExecToolWithConfig(workingDir, restrict, nil)
+func NewExecTool(workingDir string, restrict bool, allowPaths ...[]*regexp.Regexp) (*ExecTool, error) {
+ return NewExecToolWithConfig(workingDir, restrict, nil, allowPaths...)
}
-func NewExecToolWithConfig(workingDir string, restrict bool, config *config.Config) (*ExecTool, error) {
+func NewExecToolWithConfig(
+ workingDir string,
+ restrict bool,
+ config *config.Config,
+ allowPaths ...[]*regexp.Regexp,
+) (*ExecTool, error) {
denyPatterns := make([]*regexp.Regexp, 0)
customAllowPatterns := make([]*regexp.Regexp, 0)
+ var allowedPathPatterns []*regexp.Regexp
allowRemote := true
+ if len(allowPaths) > 0 {
+ allowedPathPatterns = allowPaths[0]
+ }
if config != nil {
execConfig := config.Tools.Exec
@@ -146,6 +156,7 @@ func NewExecToolWithConfig(workingDir string, restrict bool, config *config.Conf
denyPatterns: denyPatterns,
allowPatterns: nil,
customAllowPatterns: customAllowPatterns,
+ allowedPathPatterns: allowedPathPatterns,
restrictToWorkspace: restrict,
allowRemote: allowRemote,
}, nil
@@ -198,7 +209,7 @@ func (t *ExecTool) Execute(ctx context.Context, args map[string]any) *ToolResult
cwd := t.workingDir
if wd, ok := args["working_dir"].(string); ok && wd != "" {
if t.restrictToWorkspace && t.workingDir != "" {
- resolvedWD, err := validatePath(wd, t.workingDir, true)
+ resolvedWD, err := validatePathWithAllowPaths(wd, t.workingDir, true, t.allowedPathPatterns)
if err != nil {
return ErrorResult("Command blocked by safety guard (" + err.Error() + ")")
}
@@ -226,16 +237,20 @@ func (t *ExecTool) Execute(ctx context.Context, args map[string]any) *ToolResult
if err != nil {
return ErrorResult(fmt.Sprintf("Command blocked by safety guard (path resolution failed: %v)", err))
}
- absWorkspace, _ := filepath.Abs(t.workingDir)
- wsResolved, _ := filepath.EvalSymlinks(absWorkspace)
- if wsResolved == "" {
- wsResolved = absWorkspace
+ if isAllowedPath(resolved, t.allowedPathPatterns) {
+ cwd = resolved
+ } else {
+ absWorkspace, _ := filepath.Abs(t.workingDir)
+ wsResolved, _ := filepath.EvalSymlinks(absWorkspace)
+ if wsResolved == "" {
+ wsResolved = absWorkspace
+ }
+ rel, err := filepath.Rel(wsResolved, resolved)
+ if err != nil || !filepath.IsLocal(rel) {
+ return ErrorResult("Command blocked by safety guard (working directory escaped workspace)")
+ }
+ cwd = resolved
}
- rel, err := filepath.Rel(wsResolved, resolved)
- if err != nil || !filepath.IsLocal(rel) {
- return ErrorResult("Command blocked by safety guard (working directory escaped workspace)")
- }
- cwd = resolved
}
// timeout == 0 means no timeout
@@ -296,13 +311,30 @@ func (t *ExecTool) Execute(ctx context.Context, args map[string]any) *ToolResult
if err != nil {
if errors.Is(cmdCtx.Err(), context.DeadlineExceeded) {
msg := fmt.Sprintf("Command timed out after %v", t.timeout)
+ if output != "" {
+ msg += "\n\nPartial output before timeout:\n" + output
+ }
return &ToolResult{
ForLLM: msg,
ForUser: msg,
IsError: true,
+ Err: fmt.Errorf("command timeout: %w", err),
}
}
- output += fmt.Sprintf("\nExit code: %v", err)
+
+ // Extract detailed exit information
+ var exitErr *exec.ExitError
+ if errors.As(err, &exitErr) {
+ exitCode := exitErr.ExitCode()
+ output += fmt.Sprintf("\n\n[Command exited with code %d]", exitCode)
+
+ // Add signal information if killed by signal (Unix)
+ if exitCode == -1 {
+ output += " (killed by signal)"
+ }
+ } else {
+ output += fmt.Sprintf("\n\n[Command failed: %v]", err)
+ }
}
if output == "" {
@@ -412,6 +444,9 @@ func (t *ExecTool) guardCommand(command, cwd string) string {
if safePaths[p] {
continue
}
+ if isAllowedPath(p, t.allowedPathPatterns) {
+ continue
+ }
rel, err := filepath.Rel(cwdPath, p)
if err != nil {
diff --git a/pkg/tools/shell_test.go b/pkg/tools/shell_test.go
index c4553020f..f8f83ea74 100644
--- a/pkg/tools/shell_test.go
+++ b/pkg/tools/shell_test.go
@@ -489,6 +489,69 @@ func TestShellTool_SafePathsInWorkspaceRestriction(t *testing.T) {
}
}
+// TestShellTool_ExitCodeDetails verifies that exit codes are captured with details
+func TestShellTool_ExitCodeDetails(t *testing.T) {
+ tool, err := NewExecTool("", false)
+ if err != nil {
+ t.Fatalf("unable to configure exec tool: %s", err)
+ }
+
+ ctx := context.Background()
+ args := map[string]any{
+ "command": "sh -c 'exit 42'",
+ }
+
+ result := tool.Execute(ctx, args)
+
+ if !result.IsError {
+ t.Error("expected error for non-zero exit code")
+ }
+
+ // Should contain the exit code in the message (new format: "exited with code 42")
+ if !strings.Contains(result.ForLLM, "42") {
+ t.Errorf("expected exit code 42 in error message, got: %s", result.ForLLM)
+ }
+
+ // Verify the new detailed message format
+ if !strings.Contains(result.ForLLM, "exited with code") {
+ t.Errorf("expected 'exited with code' in message, got: %s", result.ForLLM)
+ }
+
+ // Err field is set by the exec system (may or may not be set depending on implementation)
+ // The important thing is that IsError=true
+ t.Logf("Exit code result: %s", result.ForLLM)
+}
+
+// TestShellTool_TimeoutWithPartialOutput verifies timeout includes partial output
+func TestShellTool_TimeoutWithPartialOutput(t *testing.T) {
+ tool, err := NewExecTool("", false)
+ if err != nil {
+ t.Fatalf("unable to configure exec tool: %s", err)
+ }
+
+ tool.SetTimeout(1 * time.Second) // Give more time for echo to complete
+
+ ctx := context.Background()
+ // Use a command that outputs immediately then sleeps
+ args := map[string]any{
+ "command": "echo 'partial output before timeout' && sleep 30",
+ }
+
+ result := tool.Execute(ctx, args)
+
+ if !result.IsError {
+ t.Error("expected error for timeout")
+ }
+
+ // Should mention timeout
+ if !strings.Contains(result.ForLLM, "timed out") {
+ t.Errorf("expected 'timed out' in message, got: %s", result.ForLLM)
+ }
+
+ // Log the result for debugging (partial output depends on shell behavior)
+ t.Logf("Timeout result: %s", result.ForLLM)
+}
+
// TestShellTool_CustomAllowPatterns verifies that custom allow patterns exempt
// commands from deny pattern checks.
func TestShellTool_CustomAllowPatterns(t *testing.T) {
diff --git a/pkg/tools/spawn.go b/pkg/tools/spawn.go
index 34ccc80e4..d019d511a 100644
--- a/pkg/tools/spawn.go
+++ b/pkg/tools/spawn.go
@@ -7,7 +7,10 @@ import (
)
type SpawnTool struct {
- manager *SubagentManager
+ spawner SubTurnSpawner
+ defaultModel string
+ maxTokens int
+ temperature float64
allowlistCheck func(targetAgentID string) bool
}
@@ -15,9 +18,19 @@ type SpawnTool struct {
var _ AsyncExecutor = (*SpawnTool)(nil)
func NewSpawnTool(manager *SubagentManager) *SpawnTool {
- return &SpawnTool{
- manager: manager,
+ if manager == nil {
+ return &SpawnTool{}
}
+ return &SpawnTool{
+ defaultModel: manager.defaultModel,
+ maxTokens: manager.maxTokens,
+ temperature: manager.temperature,
+ }
+}
+
+// SetSpawner sets the SubTurnSpawner for direct sub-turn execution.
+func (t *SpawnTool) SetSpawner(spawner SubTurnSpawner) {
+ t.spawner = spawner
}
func (t *SpawnTool) Name() string {
@@ -59,11 +72,19 @@ func (t *SpawnTool) Execute(ctx context.Context, args map[string]any) *ToolResul
// ExecuteAsync implements AsyncExecutor. The callback is passed through to the
// subagent manager as a call parameter — never stored on the SpawnTool instance.
-func (t *SpawnTool) ExecuteAsync(ctx context.Context, args map[string]any, cb AsyncCallback) *ToolResult {
+func (t *SpawnTool) ExecuteAsync(
+ ctx context.Context,
+ args map[string]any,
+ cb AsyncCallback,
+) *ToolResult {
return t.execute(ctx, args, cb)
}
-func (t *SpawnTool) execute(ctx context.Context, args map[string]any, cb AsyncCallback) *ToolResult {
+func (t *SpawnTool) execute(
+ ctx context.Context,
+ args map[string]any,
+ cb AsyncCallback,
+) *ToolResult {
task, ok := args["task"].(string)
if !ok || strings.TrimSpace(task) == "" {
return ErrorResult("task is required and must be a non-empty string")
@@ -79,31 +100,53 @@ func (t *SpawnTool) execute(ctx context.Context, args map[string]any, cb AsyncCa
}
}
- if t.manager == nil {
- return ErrorResult("Subagent manager not configured")
+ // Build system prompt for spawned subagent
+ systemPrompt := fmt.Sprintf(
+ `You are a spawned subagent running in the background. Complete the given task independently and report back when done.
+
+Task: %s`,
+ task,
+ )
+
+ if label != "" {
+ systemPrompt = fmt.Sprintf(
+ `You are a spawned subagent labeled "%s" running in the background. Complete the given task independently and report back when done.
+
+Task: %s`,
+ label,
+ task,
+ )
}
- // Read channel/chatID from context (injected by registry).
- // Fall back to "cli"/"direct" for non-conversation callers (e.g., CLI, tests)
- // to preserve the same defaults as the original NewSpawnTool constructor.
- channel := ToolChannel(ctx)
- if channel == "" {
- channel = "cli"
- }
- chatID := ToolChatID(ctx)
- if chatID == "" {
- chatID = "direct"
+ // Use spawner if available (direct SpawnSubTurn call)
+ if t.spawner != nil {
+ // Launch async sub-turn in goroutine
+ go func() {
+ result, err := t.spawner.SpawnSubTurn(ctx, SubTurnConfig{
+ Model: t.defaultModel,
+ Tools: nil, // Will inherit from parent via context
+ SystemPrompt: systemPrompt,
+ MaxTokens: t.maxTokens,
+ Temperature: t.temperature,
+ Async: true, // Async execution
+ })
+ if err != nil {
+ result = ErrorResult(fmt.Sprintf("Spawn failed: %v", err)).WithError(err)
+ }
+
+ // Call callback if provided
+ if cb != nil {
+ cb(ctx, result)
+ }
+ }()
+
+ // Return immediate acknowledgment
+ if label != "" {
+ return AsyncResult(fmt.Sprintf("Spawned subagent '%s' for task: %s", label, task))
+ }
+ return AsyncResult(fmt.Sprintf("Spawned subagent for task: %s", task))
}
- // Pass callback to manager for async completion notification
- // TODO(eventbus): when background subagents are migrated onto the
- // agent package's runTurn/sub-turn tree, emit SubTurnSpawn here and move
- // lifecycle events out of the legacy SubagentManager path.
- result, err := t.manager.Spawn(ctx, task, label, agentID, channel, chatID, cb)
- if err != nil {
- return ErrorResult(fmt.Sprintf("failed to spawn subagent: %v", err))
- }
-
- // Return AsyncResult since the task runs in background
- return AsyncResult(result)
+ // Fallback: spawner not configured
+ return ErrorResult("Subagent manager not configured")
}
diff --git a/pkg/tools/spawn_status.go b/pkg/tools/spawn_status.go
new file mode 100644
index 000000000..416fd2226
--- /dev/null
+++ b/pkg/tools/spawn_status.go
@@ -0,0 +1,178 @@
+package tools
+
+import (
+ "context"
+ "fmt"
+ "sort"
+ "strings"
+ "time"
+)
+
+// SpawnStatusTool reports the status of subagents that were spawned via the
+// spawn tool. It can query a specific task by ID, or list every known task with
+// a summary count broken-down by status.
+type SpawnStatusTool struct {
+ manager *SubagentManager
+}
+
+// NewSpawnStatusTool creates a SpawnStatusTool backed by the given manager.
+func NewSpawnStatusTool(manager *SubagentManager) *SpawnStatusTool {
+ return &SpawnStatusTool{manager: manager}
+}
+
+func (t *SpawnStatusTool) Name() string {
+ return "spawn_status"
+}
+
+func (t *SpawnStatusTool) Description() string {
+ return "Get the status of spawned subagents. " +
+ "Returns a list of all subagents and their current state " +
+ "(running, completed, failed, or canceled), or retrieves details " +
+ "for a specific subagent task when task_id is provided. " +
+ "Results are scoped to the current conversation's channel and chat ID; " +
+ "all tasks are listed only when no channel/chat context is injected " +
+ "(e.g. direct programmatic calls via Execute)."
+}
+
+func (t *SpawnStatusTool) Parameters() map[string]any {
+ return map[string]any{
+ "type": "object",
+ "properties": map[string]any{
+ "task_id": map[string]any{
+ "type": "string",
+ "description": "Optional task ID (e.g. \"subagent-1\") to inspect a specific " +
+ "subagent. When omitted, all visible subagents are listed.",
+ },
+ },
+ "required": []string{},
+ }
+}
+
+func (t *SpawnStatusTool) Execute(ctx context.Context, args map[string]any) *ToolResult {
+ if t.manager == nil {
+ return ErrorResult("Subagent manager not configured")
+ }
+
+ // Derive the calling conversation's identity so we can scope results to the
+ // current chat only — preventing cross-conversation task leakage in
+ // multi-user deployments.
+ callerChannel := ToolChannel(ctx)
+ callerChatID := ToolChatID(ctx)
+
+ var taskID string
+ if rawTaskID, ok := args["task_id"]; ok && rawTaskID != nil {
+ taskIDStr, ok := rawTaskID.(string)
+ if !ok {
+ return ErrorResult("task_id must be a string")
+ }
+ taskID = strings.TrimSpace(taskIDStr)
+ }
+
+ if taskID != "" {
+ // GetTaskCopy returns a consistent snapshot under the manager lock,
+ // eliminating any data race with the concurrent subagent goroutine.
+ taskCopy, ok := t.manager.GetTaskCopy(taskID)
+ if !ok {
+ return ErrorResult(fmt.Sprintf("No subagent found with task ID: %s", taskID))
+ }
+
+ // Restrict lookup to tasks that belong to this conversation.
+ if callerChannel != "" && taskCopy.OriginChannel != "" && taskCopy.OriginChannel != callerChannel {
+ return ErrorResult(fmt.Sprintf("No subagent found with task ID: %s", taskID))
+ }
+ if callerChatID != "" && taskCopy.OriginChatID != "" && taskCopy.OriginChatID != callerChatID {
+ return ErrorResult(fmt.Sprintf("No subagent found with task ID: %s", taskID))
+ }
+
+ return NewToolResult(spawnStatusFormatTask(&taskCopy))
+ }
+
+ // ListTaskCopies returns consistent snapshots under the manager lock.
+ origTasks := t.manager.ListTaskCopies()
+ if len(origTasks) == 0 {
+ return NewToolResult("No subagents have been spawned yet.")
+ }
+
+ tasks := make([]*SubagentTask, 0, len(origTasks))
+ for i := range origTasks {
+ cpy := &origTasks[i]
+
+ // Filter to tasks that originate from the current conversation only.
+ if callerChannel != "" && cpy.OriginChannel != "" && cpy.OriginChannel != callerChannel {
+ continue
+ }
+ if callerChatID != "" && cpy.OriginChatID != "" && cpy.OriginChatID != callerChatID {
+ continue
+ }
+
+ tasks = append(tasks, cpy)
+ }
+
+ if len(tasks) == 0 {
+ return NewToolResult("No subagents found for this conversation.")
+ }
+
+ // Order by creation time (ascending) so spawning order is preserved.
+ // Fall back to ID string for tasks created in the same millisecond.
+ sort.Slice(tasks, func(i, j int) bool {
+ if tasks[i].Created != tasks[j].Created {
+ return tasks[i].Created < tasks[j].Created
+ }
+ return tasks[i].ID < tasks[j].ID
+ })
+
+ counts := map[string]int{}
+ for _, task := range tasks {
+ counts[task.Status]++
+ }
+
+ var sb strings.Builder
+ sb.WriteString(fmt.Sprintf("Subagent status report (%d total):\n", len(tasks)))
+ for _, status := range []string{"running", "completed", "failed", "canceled"} {
+ if n := counts[status]; n > 0 {
+ label := strings.ToUpper(status[:1]) + status[1:] + ":"
+ sb.WriteString(fmt.Sprintf(" %-10s %d\n", label, n))
+ }
+ }
+ sb.WriteString("\n")
+
+ for _, task := range tasks {
+ sb.WriteString(spawnStatusFormatTask(task))
+ sb.WriteString("\n\n")
+ }
+
+ return NewToolResult(strings.TrimRight(sb.String(), "\n"))
+}
+
+// spawnStatusFormatTask renders a single SubagentTask as a human-readable block.
+func spawnStatusFormatTask(task *SubagentTask) string {
+ var sb strings.Builder
+
+ header := fmt.Sprintf("[%s] status=%s", task.ID, task.Status)
+ if task.Label != "" {
+ header += fmt.Sprintf(" label=%q", task.Label)
+ }
+ if task.AgentID != "" {
+ header += fmt.Sprintf(" agent=%s", task.AgentID)
+ }
+ if task.Created > 0 {
+ created := time.UnixMilli(task.Created).UTC().Format("2006-01-02 15:04:05 UTC")
+ header += fmt.Sprintf(" created=%s", created)
+ }
+ sb.WriteString(header)
+
+ if task.Task != "" {
+ sb.WriteString(fmt.Sprintf("\n task: %s", task.Task))
+ }
+ if task.Result != "" {
+ result := task.Result
+ const maxResultLen = 300
+ runes := []rune(result)
+ if len(runes) > maxResultLen {
+ result = string(runes[:maxResultLen]) + "…"
+ }
+ sb.WriteString(fmt.Sprintf("\n result: %s", result))
+ }
+
+ return sb.String()
+}
diff --git a/pkg/tools/spawn_status_test.go b/pkg/tools/spawn_status_test.go
new file mode 100644
index 000000000..9c772d61a
--- /dev/null
+++ b/pkg/tools/spawn_status_test.go
@@ -0,0 +1,406 @@
+package tools
+
+import (
+ "context"
+ "fmt"
+ "strings"
+ "testing"
+ "time"
+)
+
+func TestSpawnStatusTool_Name(t *testing.T) {
+ provider := &MockLLMProvider{}
+ workspace := t.TempDir()
+ manager := NewSubagentManager(provider, "test-model", workspace)
+ tool := NewSpawnStatusTool(manager)
+
+ if tool.Name() != "spawn_status" {
+ t.Errorf("Expected name 'spawn_status', got '%s'", tool.Name())
+ }
+}
+
+func TestSpawnStatusTool_Description(t *testing.T) {
+ provider := &MockLLMProvider{}
+ workspace := t.TempDir()
+ manager := NewSubagentManager(provider, "test-model", workspace)
+ tool := NewSpawnStatusTool(manager)
+
+ desc := tool.Description()
+ if desc == "" {
+ t.Error("Description should not be empty")
+ }
+ if !strings.Contains(strings.ToLower(desc), "subagent") {
+ t.Errorf("Description should mention 'subagent', got: %s", desc)
+ }
+}
+
+func TestSpawnStatusTool_Parameters(t *testing.T) {
+ provider := &MockLLMProvider{}
+ workspace := t.TempDir()
+ manager := NewSubagentManager(provider, "test-model", workspace)
+ tool := NewSpawnStatusTool(manager)
+
+ params := tool.Parameters()
+ if params["type"] != "object" {
+ t.Errorf("Expected type 'object', got: %v", params["type"])
+ }
+ props, ok := params["properties"].(map[string]any)
+ if !ok {
+ t.Fatal("Expected 'properties' to be a map")
+ }
+ if _, hasTaskID := props["task_id"]; !hasTaskID {
+ t.Error("Expected 'task_id' parameter in properties")
+ }
+}
+
+func TestSpawnStatusTool_NilManager(t *testing.T) {
+ tool := &SpawnStatusTool{manager: nil}
+ result := tool.Execute(context.Background(), map[string]any{})
+ if !result.IsError {
+ t.Error("Expected error result when manager is nil")
+ }
+}
+
+func TestSpawnStatusTool_Empty(t *testing.T) {
+ provider := &MockLLMProvider{}
+ workspace := t.TempDir()
+ manager := NewSubagentManager(provider, "test-model", workspace)
+ tool := NewSpawnStatusTool(manager)
+
+ result := tool.Execute(context.Background(), map[string]any{})
+ if result.IsError {
+ t.Fatalf("Expected success, got error: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "No subagents") {
+ t.Errorf("Expected 'No subagents' message, got: %s", result.ForLLM)
+ }
+}
+
+func TestSpawnStatusTool_ListAll(t *testing.T) {
+ provider := &MockLLMProvider{}
+ workspace := t.TempDir()
+ manager := NewSubagentManager(provider, "test-model", workspace)
+
+ now := time.Now().UnixMilli()
+ manager.mu.Lock()
+ manager.tasks["subagent-1"] = &SubagentTask{
+ ID: "subagent-1",
+ Task: "Do task A",
+ Label: "task-a",
+ Status: "running",
+ Created: now,
+ }
+ manager.tasks["subagent-2"] = &SubagentTask{
+ ID: "subagent-2",
+ Task: "Do task B",
+ Label: "task-b",
+ Status: "completed",
+ Result: "Done successfully",
+ Created: now,
+ }
+ manager.tasks["subagent-3"] = &SubagentTask{
+ ID: "subagent-3",
+ Task: "Do task C",
+ Status: "failed",
+ Result: "Error: something went wrong",
+ }
+ manager.mu.Unlock()
+
+ tool := NewSpawnStatusTool(manager)
+ result := tool.Execute(context.Background(), map[string]any{})
+
+ if result.IsError {
+ t.Fatalf("Expected success, got error: %s", result.ForLLM)
+ }
+
+ // Summary header
+ if !strings.Contains(result.ForLLM, "3 total") {
+ t.Errorf("Expected total count in header, got: %s", result.ForLLM)
+ }
+
+ // Individual task IDs
+ for _, id := range []string{"subagent-1", "subagent-2", "subagent-3"} {
+ if !strings.Contains(result.ForLLM, id) {
+ t.Errorf("Expected task %s in output, got:\n%s", id, result.ForLLM)
+ }
+ }
+
+ // Status values
+ for _, status := range []string{"running", "completed", "failed"} {
+ if !strings.Contains(result.ForLLM, status) {
+ t.Errorf("Expected status '%s' in output, got:\n%s", status, result.ForLLM)
+ }
+ }
+
+ // Result content
+ if !strings.Contains(result.ForLLM, "Done successfully") {
+ t.Errorf("Expected result text in output, got:\n%s", result.ForLLM)
+ }
+}
+
+func TestSpawnStatusTool_GetByID(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+
+ manager.mu.Lock()
+ manager.tasks["subagent-42"] = &SubagentTask{
+ ID: "subagent-42",
+ Task: "Specific task",
+ Label: "my-task",
+ Status: "failed",
+ Result: "Something went wrong",
+ Created: time.Now().UnixMilli(),
+ }
+ manager.mu.Unlock()
+
+ tool := NewSpawnStatusTool(manager)
+ result := tool.Execute(context.Background(), map[string]any{"task_id": "subagent-42"})
+
+ if result.IsError {
+ t.Fatalf("Expected success, got error: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "subagent-42") {
+ t.Errorf("Expected task ID in output, got: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "failed") {
+ t.Errorf("Expected status 'failed' in output, got: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "Something went wrong") {
+ t.Errorf("Expected result text in output, got: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "my-task") {
+ t.Errorf("Expected label in output, got: %s", result.ForLLM)
+ }
+}
+
+func TestSpawnStatusTool_GetByID_NotFound(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+ tool := NewSpawnStatusTool(manager)
+
+ result := tool.Execute(context.Background(), map[string]any{"task_id": "nonexistent-999"})
+ if !result.IsError {
+ t.Errorf("Expected error for nonexistent task, got: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "nonexistent-999") {
+ t.Errorf("Expected task ID in error message, got: %s", result.ForLLM)
+ }
+}
+
+func TestSpawnStatusTool_TaskID_NonString(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+ tool := NewSpawnStatusTool(manager)
+
+ for _, badVal := range []any{42, 3.14, true, map[string]any{"x": 1}, []string{"a"}} {
+ result := tool.Execute(context.Background(), map[string]any{"task_id": badVal})
+ if !result.IsError {
+ t.Errorf("Expected error for task_id=%T(%v), got success: %s", badVal, badVal, result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "task_id must be a string") {
+ t.Errorf("Expected type-error message, got: %s", result.ForLLM)
+ }
+ }
+}
+
+func TestSpawnStatusTool_ResultTruncation(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+
+ longResult := strings.Repeat("X", 500)
+ manager.mu.Lock()
+ manager.tasks["subagent-1"] = &SubagentTask{
+ ID: "subagent-1",
+ Task: "Long task",
+ Status: "completed",
+ Result: longResult,
+ }
+ manager.mu.Unlock()
+
+ tool := NewSpawnStatusTool(manager)
+ result := tool.Execute(context.Background(), map[string]any{"task_id": "subagent-1"})
+
+ if result.IsError {
+ t.Fatalf("Unexpected error: %s", result.ForLLM)
+ }
+ // Output should be shorter than the raw result due to truncation
+ if len(result.ForLLM) >= len(longResult) {
+ t.Errorf("Expected result to be truncated, but ForLLM is %d chars", len(result.ForLLM))
+ }
+ if !strings.Contains(result.ForLLM, "…") {
+ t.Errorf("Expected truncation indicator '…' in output, got: %s", result.ForLLM)
+ }
+}
+
+func TestSpawnStatusTool_ResultTruncation_Unicode(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+
+ // Each CJK rune is 3 bytes; 400 runes = 1200 bytes — well over the 300-rune limit.
+ cjkChar := string(rune(0x5b57))
+ longResult := strings.Repeat(cjkChar, 400)
+ manager.mu.Lock()
+ manager.tasks["subagent-1"] = &SubagentTask{
+ ID: "subagent-1",
+ Task: "Unicode task",
+ Status: "completed",
+ Result: longResult,
+ }
+ manager.mu.Unlock()
+
+ tool := NewSpawnStatusTool(manager)
+ result := tool.Execute(context.Background(), map[string]any{"task_id": "subagent-1"})
+
+ if result.IsError {
+ t.Fatalf("Unexpected error: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "…") {
+ t.Errorf("Expected truncation indicator in output")
+ }
+ // The truncated result must be valid UTF-8 (no split rune boundaries).
+ if !strings.Contains(result.ForLLM, cjkChar) {
+ t.Errorf("Expected CJK runes to appear intact in output")
+ }
+}
+
+func TestSpawnStatusTool_StatusCounts(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+
+ manager.mu.Lock()
+ for i, status := range []string{"running", "running", "completed", "failed", "canceled"} {
+ id := fmt.Sprintf("subagent-%d", i+1)
+ manager.tasks[id] = &SubagentTask{ID: id, Task: "t", Status: status}
+ }
+ manager.mu.Unlock()
+
+ tool := NewSpawnStatusTool(manager)
+ result := tool.Execute(context.Background(), map[string]any{})
+
+ if result.IsError {
+ t.Fatalf("Unexpected error: %s", result.ForLLM)
+ }
+ // The summary line should mention all statuses that have counts
+ for _, want := range []string{"Running:", "Completed:", "Failed:", "Canceled:"} {
+ if !strings.Contains(result.ForLLM, want) {
+ t.Errorf("Expected %q in summary, got:\n%s", want, result.ForLLM)
+ }
+ }
+}
+
+func TestSpawnStatusTool_SortByCreatedTimestamp(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+
+ now := time.Now().UnixMilli()
+ manager.mu.Lock()
+ // Intentionally insert with out-of-order IDs and timestamps that reflect
+ // true spawn order: subagent-2 was spawned first, subagent-10 second.
+ manager.tasks["subagent-10"] = &SubagentTask{
+ ID: "subagent-10", Task: "second", Status: "running",
+ Created: now + 1,
+ }
+ manager.tasks["subagent-2"] = &SubagentTask{
+ ID: "subagent-2", Task: "first", Status: "running",
+ Created: now,
+ }
+ manager.mu.Unlock()
+
+ tool := NewSpawnStatusTool(manager)
+ result := tool.Execute(context.Background(), map[string]any{})
+
+ if result.IsError {
+ t.Fatalf("Unexpected error: %s", result.ForLLM)
+ }
+
+ pos2 := strings.Index(result.ForLLM, "subagent-2")
+ pos10 := strings.Index(result.ForLLM, "subagent-10")
+ if pos2 < 0 || pos10 < 0 {
+ t.Fatalf("Both task IDs should appear in output:\n%s", result.ForLLM)
+ }
+ if pos2 > pos10 {
+ t.Errorf("Expected subagent-2 (created first) to appear before subagent-10, but got:\n%s", result.ForLLM)
+ }
+}
+
+func TestSpawnStatusTool_ChannelFiltering_ListAll(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+
+ manager.mu.Lock()
+ manager.tasks["subagent-1"] = &SubagentTask{
+ ID: "subagent-1", Task: "mine", Status: "running",
+ OriginChannel: "telegram", OriginChatID: "chat-A",
+ }
+ manager.tasks["subagent-2"] = &SubagentTask{
+ ID: "subagent-2", Task: "other user", Status: "running",
+ OriginChannel: "telegram", OriginChatID: "chat-B",
+ }
+ manager.mu.Unlock()
+
+ tool := NewSpawnStatusTool(manager)
+
+ // Caller is chat-A — should only see subagent-1.
+ ctx := WithToolContext(context.Background(), "telegram", "chat-A")
+ result := tool.Execute(ctx, map[string]any{})
+
+ if result.IsError {
+ t.Fatalf("Unexpected error: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "subagent-1") {
+ t.Errorf("Expected own task in output, got:\n%s", result.ForLLM)
+ }
+ if strings.Contains(result.ForLLM, "subagent-2") {
+ t.Errorf("Should NOT see other chat's task, got:\n%s", result.ForLLM)
+ }
+}
+
+func TestSpawnStatusTool_ChannelFiltering_GetByID(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+
+ manager.mu.Lock()
+ manager.tasks["subagent-99"] = &SubagentTask{
+ ID: "subagent-99", Task: "secret", Status: "completed", Result: "private data",
+ OriginChannel: "slack", OriginChatID: "room-Z",
+ }
+ manager.mu.Unlock()
+
+ tool := NewSpawnStatusTool(manager)
+
+ // Different chat trying to look up subagent-99 by ID.
+ ctx := WithToolContext(context.Background(), "slack", "room-OTHER")
+ result := tool.Execute(ctx, map[string]any{"task_id": "subagent-99"})
+
+ if !result.IsError {
+ t.Errorf("Expected error (cross-chat lookup blocked), got: %s", result.ForLLM)
+ }
+}
+
+func TestSpawnStatusTool_ChannelFiltering_NoContext(t *testing.T) {
+ provider := &MockLLMProvider{}
+ manager := NewSubagentManager(provider, "test-model", "/tmp/test")
+
+ manager.mu.Lock()
+ manager.tasks["subagent-1"] = &SubagentTask{
+ ID: "subagent-1", Task: "t", Status: "completed",
+ OriginChannel: "telegram", OriginChatID: "chat-A",
+ }
+ manager.mu.Unlock()
+
+ tool := NewSpawnStatusTool(manager)
+
+ // No ToolContext injected (e.g. a direct programmatic call that bypasses
+ // WithToolContext entirely) — callerChannel and callerChatID are both "".
+ // Note: the normal CLI path uses ProcessDirectWithChannel("cli", "direct"),
+ // which *does* inject a non-empty context; this test covers the case where
+ // no context injection happens at all.
+ // The filter conditions require a non-empty caller value, so all tasks pass through.
+ result := tool.Execute(context.Background(), map[string]any{})
+ if result.IsError {
+ t.Fatalf("Unexpected error: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "subagent-1") {
+ t.Errorf("Expected task visible from no-context caller, got:\n%s", result.ForLLM)
+ }
+}
diff --git a/pkg/tools/spawn_test.go b/pkg/tools/spawn_test.go
index 43223b8db..fda6bbd89 100644
--- a/pkg/tools/spawn_test.go
+++ b/pkg/tools/spawn_test.go
@@ -6,6 +6,24 @@ import (
"testing"
)
+// mockSpawner implements SubTurnSpawner for testing
+type mockSpawner struct{}
+
+func (m *mockSpawner) SpawnSubTurn(ctx context.Context, cfg SubTurnConfig) (*ToolResult, error) {
+ // Extract task from system prompt for response
+ task := cfg.SystemPrompt
+ if strings.Contains(task, "Task: ") {
+ parts := strings.Split(task, "Task: ")
+ if len(parts) > 1 {
+ task = parts[1]
+ }
+ }
+ return &ToolResult{
+ ForLLM: "Task completed: " + task,
+ ForUser: "Task completed",
+ }, nil
+}
+
func TestSpawnTool_Execute_EmptyTask(t *testing.T) {
provider := &MockLLMProvider{}
manager := NewSubagentManager(provider, "test-model", "/tmp/test")
@@ -44,6 +62,7 @@ func TestSpawnTool_Execute_ValidTask(t *testing.T) {
provider := &MockLLMProvider{}
manager := NewSubagentManager(provider, "test-model", "/tmp/test")
tool := NewSpawnTool(manager)
+ tool.SetSpawner(&mockSpawner{})
ctx := context.Background()
args := map[string]any{
diff --git a/pkg/tools/subagent.go b/pkg/tools/subagent.go
index 9915c5900..9a1a8b802 100644
--- a/pkg/tools/subagent.go
+++ b/pkg/tools/subagent.go
@@ -4,11 +4,34 @@ import (
"context"
"fmt"
"sync"
+ "sync/atomic"
"time"
"github.com/sipeed/picoclaw/pkg/providers"
)
+// SubTurnSpawner is an interface for spawning sub-turns.
+// This avoids circular dependency between tools and agent packages.
+type SubTurnSpawner interface {
+ SpawnSubTurn(ctx context.Context, cfg SubTurnConfig) (*ToolResult, error)
+}
+
+// SubTurnConfig holds configuration for spawning a sub-turn.
+type SubTurnConfig struct {
+ Model string
+ Tools []Tool
+ SystemPrompt string
+ MaxTokens int
+ Temperature float64
+ Async bool // true for async (spawn), false for sync (subagent)
+ Critical bool // continue running after parent finishes gracefully
+ Timeout time.Duration // 0 = use default (5 minutes)
+ MaxContextRunes int // 0 = auto, -1 = no limit, >0 = explicit limit
+ ActualSystemPrompt string
+ InitialMessages []providers.Message
+ InitialTokenBudget *atomic.Int64 // Shared token budget for team members; nil if no budget
+}
+
type SubagentTask struct {
ID string
Task string
@@ -21,6 +44,15 @@ type SubagentTask struct {
Created int64
}
+type SpawnSubTurnFunc func(
+ ctx context.Context,
+ task, label, agentID string,
+ tools *ToolRegistry,
+ maxTokens int,
+ temperature float64,
+ hasMaxTokens, hasTemperature bool,
+) (*ToolResult, error)
+
type SubagentManager struct {
tasks map[string]*SubagentTask
mu sync.RWMutex
@@ -34,6 +66,7 @@ type SubagentManager struct {
hasMaxTokens bool
hasTemperature bool
nextID int
+ spawner SpawnSubTurnFunc
}
func NewSubagentManager(
@@ -51,6 +84,12 @@ func NewSubagentManager(
}
}
+func (sm *SubagentManager) SetSpawner(spawner SpawnSubTurnFunc) {
+ sm.mu.Lock()
+ defer sm.mu.Unlock()
+ sm.spawner = spawner
+}
+
// SetLLMOptions sets max tokens and temperature for subagent LLM calls.
func (sm *SubagentManager) SetLLMOptions(maxTokens int, temperature float64) {
sm.mu.Lock()
@@ -108,29 +147,17 @@ func (sm *SubagentManager) Spawn(
return fmt.Sprintf("Spawned subagent for task: %s", task), nil
}
-func (sm *SubagentManager) runTask(ctx context.Context, task *SubagentTask, callback AsyncCallback) {
+func (sm *SubagentManager) runTask(
+ ctx context.Context,
+ task *SubagentTask,
+ callback AsyncCallback,
+) {
task.Status = "running"
task.Created = time.Now().UnixMilli()
// TODO(eventbus): once subagents are modeled as child turns inside
// pkg/agent, emit SubTurnEnd and SubTurnResultDelivered from the parent
// AgentLoop instead of this legacy manager.
- // Build system prompt for subagent
- systemPrompt := `You are a subagent. Complete the given task independently and report the result.
-You have access to tools - use them as needed to complete your task.
-After completing the task, provide a clear summary of what was done.`
-
- messages := []providers.Message{
- {
- Role: "system",
- Content: systemPrompt,
- },
- {
- Role: "user",
- Content: task.Task,
- },
- }
-
// Check if context is already canceled before starting
select {
case <-ctx.Done():
@@ -142,8 +169,8 @@ After completing the task, provide a clear summary of what was done.`
default:
}
- // Run tool loop with access to tools
sm.mu.RLock()
+ spawner := sm.spawner
tools := sm.tools
maxIter := sm.maxIterations
maxTokens := sm.maxTokens
@@ -152,27 +179,69 @@ After completing the task, provide a clear summary of what was done.`
hasTemperature := sm.hasTemperature
sm.mu.RUnlock()
- var llmOptions map[string]any
- if hasMaxTokens || hasTemperature {
- llmOptions = map[string]any{}
- if hasMaxTokens {
- llmOptions["max_tokens"] = maxTokens
+ var result *ToolResult
+ var err error
+
+ if spawner != nil {
+ result, err = spawner(
+ ctx,
+ task.Task,
+ task.Label,
+ task.AgentID,
+ tools,
+ maxTokens,
+ temperature,
+ hasMaxTokens,
+ hasTemperature,
+ )
+ } else {
+ // Fallback to legacy RunToolLoop
+ systemPrompt := `You are a subagent. Complete the given task independently and report the result.
+You have access to tools - use them as needed to complete your task.
+After completing the task, provide a clear summary of what was done.`
+
+ messages := []providers.Message{
+ {Role: "system", Content: systemPrompt},
+ {Role: "user", Content: task.Task},
}
- if hasTemperature {
- llmOptions["temperature"] = temperature
+
+ var llmOptions map[string]any
+ if hasMaxTokens || hasTemperature {
+ llmOptions = map[string]any{}
+ if hasMaxTokens {
+ llmOptions["max_tokens"] = maxTokens
+ }
+ if hasTemperature {
+ llmOptions["temperature"] = temperature
+ }
+ }
+
+ var loopResult *ToolLoopResult
+ loopResult, err = RunToolLoop(ctx, ToolLoopConfig{
+ Provider: sm.provider,
+ Model: sm.defaultModel,
+ Tools: tools,
+ MaxIterations: maxIter,
+ LLMOptions: llmOptions,
+ }, messages, task.OriginChannel, task.OriginChatID)
+
+ if err == nil {
+ result = &ToolResult{
+ ForLLM: fmt.Sprintf(
+ "Subagent '%s' completed (iterations: %d): %s",
+ task.Label,
+ loopResult.Iterations,
+ loopResult.Content,
+ ),
+ ForUser: loopResult.Content,
+ Silent: false,
+ IsError: false,
+ Async: false,
+ }
}
}
- loopResult, err := RunToolLoop(ctx, ToolLoopConfig{
- Provider: sm.provider,
- Model: sm.defaultModel,
- Tools: tools,
- MaxIterations: maxIter,
- LLMOptions: llmOptions,
- }, messages, task.OriginChannel, task.OriginChatID)
-
sm.mu.Lock()
- var result *ToolResult
defer func() {
sm.mu.Unlock()
// Call callback if provided and result is set
@@ -199,19 +268,7 @@ After completing the task, provide a clear summary of what was done.`
}
} else {
task.Status = "completed"
- task.Result = loopResult.Content
- result = &ToolResult{
- ForLLM: fmt.Sprintf(
- "Subagent '%s' completed (iterations: %d): %s",
- task.Label,
- loopResult.Iterations,
- loopResult.Content,
- ),
- ForUser: loopResult.Content,
- Silent: false,
- IsError: false,
- Async: false,
- }
+ task.Result = result.ForLLM
}
}
@@ -222,6 +279,18 @@ func (sm *SubagentManager) GetTask(taskID string) (*SubagentTask, bool) {
return task, ok
}
+// GetTaskCopy returns a copy of the task with the given ID, taken under the
+// read lock, so the caller receives a consistent snapshot with no data race.
+func (sm *SubagentManager) GetTaskCopy(taskID string) (SubagentTask, bool) {
+ sm.mu.RLock()
+ defer sm.mu.RUnlock()
+ task, ok := sm.tasks[taskID]
+ if !ok {
+ return SubagentTask{}, false
+ }
+ return *task, true
+}
+
func (sm *SubagentManager) ListTasks() []*SubagentTask {
sm.mu.RLock()
defer sm.mu.RUnlock()
@@ -233,17 +302,42 @@ func (sm *SubagentManager) ListTasks() []*SubagentTask {
return tasks
}
+// ListTaskCopies returns value copies of all tasks, taken under the read lock,
+// so callers receive consistent snapshots with no data race.
+func (sm *SubagentManager) ListTaskCopies() []SubagentTask {
+ sm.mu.RLock()
+ defer sm.mu.RUnlock()
+
+ copies := make([]SubagentTask, 0, len(sm.tasks))
+ for _, task := range sm.tasks {
+ copies = append(copies, *task)
+ }
+ return copies
+}
+
// SubagentTool executes a subagent task synchronously and returns the result.
-// Unlike SpawnTool which runs tasks asynchronously, SubagentTool waits for completion
-// and returns the result directly in the ToolResult.
+// It directly calls SubTurnSpawner with Async=false for synchronous execution.
type SubagentTool struct {
- manager *SubagentManager
+ spawner SubTurnSpawner
+ defaultModel string
+ maxTokens int
+ temperature float64
}
func NewSubagentTool(manager *SubagentManager) *SubagentTool {
- return &SubagentTool{
- manager: manager,
+ if manager == nil {
+ return &SubagentTool{}
}
+ return &SubagentTool{
+ defaultModel: manager.defaultModel,
+ maxTokens: manager.maxTokens,
+ temperature: manager.temperature,
+ }
+}
+
+// SetSpawner sets the SubTurnSpawner for direct sub-turn execution.
+func (t *SubagentTool) SetSpawner(spawner SubTurnSpawner) {
+ t.spawner = spawner
}
func (t *SubagentTool) Name() string {
@@ -279,86 +373,64 @@ func (t *SubagentTool) Execute(ctx context.Context, args map[string]any) *ToolRe
label, _ := args["label"].(string)
- if t.manager == nil {
- return ErrorResult("Subagent manager not configured").WithError(fmt.Errorf("manager is nil"))
+ // Build system prompt for subagent
+ systemPrompt := fmt.Sprintf(
+ `You are a subagent. Complete the given task independently and provide a clear, concise result.
+
+Task: %s`,
+ task,
+ )
+
+ if label != "" {
+ systemPrompt = fmt.Sprintf(
+ `You are a subagent labeled "%s". Complete the given task independently and provide a clear, concise result.
+
+Task: %s`,
+ label,
+ task,
+ )
}
- // Build messages for subagent
- messages := []providers.Message{
- {
- Role: "system",
- Content: "You are a subagent. Complete the given task independently and provide a clear, concise result.",
- },
- {
- Role: "user",
- Content: task,
- },
- }
-
- // Use RunToolLoop to execute with tools (same as async SpawnTool)
- sm := t.manager
- sm.mu.RLock()
- tools := sm.tools
- maxIter := sm.maxIterations
- maxTokens := sm.maxTokens
- temperature := sm.temperature
- hasMaxTokens := sm.hasMaxTokens
- hasTemperature := sm.hasTemperature
- sm.mu.RUnlock()
-
- var llmOptions map[string]any
- if hasMaxTokens || hasTemperature {
- llmOptions = map[string]any{}
- if hasMaxTokens {
- llmOptions["max_tokens"] = maxTokens
+ // Use spawner if available (direct SpawnSubTurn call)
+ if t.spawner != nil {
+ result, err := t.spawner.SpawnSubTurn(ctx, SubTurnConfig{
+ Model: t.defaultModel,
+ Tools: nil, // Will inherit from parent via context
+ SystemPrompt: systemPrompt,
+ MaxTokens: t.maxTokens,
+ Temperature: t.temperature,
+ Async: false, // Synchronous execution
+ })
+ if err != nil {
+ return ErrorResult(fmt.Sprintf("Subagent execution failed: %v", err)).WithError(err)
}
- if hasTemperature {
- llmOptions["temperature"] = temperature
+
+ // Format result for display
+ userContent := result.ForLLM
+ if result.ForUser != "" {
+ userContent = result.ForUser
+ }
+ maxUserLen := 500
+ if len(userContent) > maxUserLen {
+ userContent = userContent[:maxUserLen] + "..."
+ }
+
+ labelStr := label
+ if labelStr == "" {
+ labelStr = "(unnamed)"
+ }
+ llmContent := fmt.Sprintf("Subagent task completed:\nLabel: %s\nResult: %s",
+ labelStr, result.ForLLM)
+
+ return &ToolResult{
+ ForLLM: llmContent,
+ ForUser: userContent,
+ Silent: false,
+ IsError: result.IsError,
+ Async: false,
}
}
- // Fall back to "cli"/"direct" for non-conversation callers (e.g., CLI, tests)
- // to preserve the same defaults as the original NewSubagentTool constructor.
- channel := ToolChannel(ctx)
- if channel == "" {
- channel = "cli"
- }
- chatID := ToolChatID(ctx)
- if chatID == "" {
- chatID = "direct"
- }
-
- loopResult, err := RunToolLoop(ctx, ToolLoopConfig{
- Provider: sm.provider,
- Model: sm.defaultModel,
- Tools: tools,
- MaxIterations: maxIter,
- LLMOptions: llmOptions,
- }, messages, channel, chatID)
- if err != nil {
- return ErrorResult(fmt.Sprintf("Subagent execution failed: %v", err)).WithError(err)
- }
-
- // ForUser: Brief summary for user (truncated if too long)
- userContent := loopResult.Content
- maxUserLen := 500
- if len(userContent) > maxUserLen {
- userContent = userContent[:maxUserLen] + "..."
- }
-
- // ForLLM: Full execution details
- labelStr := label
- if labelStr == "" {
- labelStr = "(unnamed)"
- }
- llmContent := fmt.Sprintf("Subagent task completed:\nLabel: %s\nIterations: %d\nResult: %s",
- labelStr, loopResult.Iterations, loopResult.Content)
-
- return &ToolResult{
- ForLLM: llmContent,
- ForUser: userContent,
- Silent: false,
- IsError: false,
- Async: false,
- }
+ // Fallback: spawner not configured
+ return ErrorResult("Subagent manager not configured").WithError(fmt.Errorf("spawner not set"))
}
diff --git a/pkg/tools/subagent_tool_test.go b/pkg/tools/subagent_tool_test.go
index 4b6f130a5..89ac7d4b5 100644
--- a/pkg/tools/subagent_tool_test.go
+++ b/pkg/tools/subagent_tool_test.go
@@ -48,24 +48,19 @@ func TestSubagentManager_SetLLMOptions_AppliesToRunToolLoop(t *testing.T) {
provider := &MockLLMProvider{}
manager := NewSubagentManager(provider, "test-model", "/tmp/test")
manager.SetLLMOptions(2048, 0.6)
- tool := NewSubagentTool(manager)
- ctx := WithToolContext(context.Background(), "cli", "direct")
- args := map[string]any{"task": "Do something"}
- result := tool.Execute(ctx, args)
-
- if result == nil || result.IsError {
- t.Fatalf("Expected successful result, got: %+v", result)
+ // Verify options are set on manager
+ if manager.maxTokens != 2048 {
+ t.Errorf("manager.maxTokens = %d, want 2048", manager.maxTokens)
}
-
- if provider.lastOptions == nil {
- t.Fatal("Expected LLM options to be passed, got nil")
+ if manager.temperature != 0.6 {
+ t.Errorf("manager.temperature = %f, want 0.6", manager.temperature)
}
- if provider.lastOptions["max_tokens"] != 2048 {
- t.Fatalf("max_tokens = %v, want %d", provider.lastOptions["max_tokens"], 2048)
+ if !manager.hasMaxTokens {
+ t.Error("manager.hasMaxTokens should be true")
}
- if provider.lastOptions["temperature"] != 0.6 {
- t.Fatalf("temperature = %v, want %v", provider.lastOptions["temperature"], 0.6)
+ if !manager.hasTemperature {
+ t.Error("manager.hasTemperature should be true")
}
}
@@ -150,6 +145,7 @@ func TestSubagentTool_Execute_Success(t *testing.T) {
provider := &MockLLMProvider{}
manager := NewSubagentManager(provider, "test-model", "/tmp/test")
tool := NewSubagentTool(manager)
+ tool.SetSpawner(&mockSpawner{})
ctx := WithToolContext(context.Background(), "telegram", "chat-123")
args := map[string]any{
@@ -204,6 +200,7 @@ func TestSubagentTool_Execute_NoLabel(t *testing.T) {
provider := &MockLLMProvider{}
manager := NewSubagentManager(provider, "test-model", "/tmp/test")
tool := NewSubagentTool(manager)
+ tool.SetSpawner(&mockSpawner{})
ctx := context.Background()
args := map[string]any{
@@ -277,6 +274,7 @@ func TestSubagentTool_Execute_ContextPassing(t *testing.T) {
provider := &MockLLMProvider{}
manager := NewSubagentManager(provider, "test-model", "/tmp/test")
tool := NewSubagentTool(manager)
+ tool.SetSpawner(&mockSpawner{})
channel := "test-channel"
chatID := "test-chat"
@@ -302,6 +300,7 @@ func TestSubagentTool_ForUserTruncation(t *testing.T) {
provider := &MockLLMProvider{}
manager := NewSubagentManager(provider, "test-model", "/tmp/test")
tool := NewSubagentTool(manager)
+ tool.SetSpawner(&mockSpawner{})
ctx := context.Background()
diff --git a/pkg/tools/web.go b/pkg/tools/web.go
index e5036d3a8..42cf79578 100644
--- a/pkg/tools/web.go
+++ b/pkg/tools/web.go
@@ -7,6 +7,7 @@ import (
"errors"
"fmt"
"io"
+ "mime"
"net"
"net/http"
"net/url"
@@ -15,11 +16,14 @@ import (
"sync/atomic"
"time"
+ "github.com/sipeed/picoclaw/pkg/config"
+ "github.com/sipeed/picoclaw/pkg/logger"
"github.com/sipeed/picoclaw/pkg/utils"
)
const (
- userAgent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
+ userAgent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
+ userAgentHonest = "picoclaw/%s (+https://github.com/sipeed/picoclaw; AI assistant bot)"
// HTTP client timeouts for web tool providers.
searchTimeout = 10 * time.Second // Brave, Tavily, DuckDuckGo
@@ -776,22 +780,49 @@ type WebFetchTool struct {
maxChars int
proxy string
client *http.Client
+ format string
fetchLimitBytes int64
+ whitelist *privateHostWhitelist
}
-func NewWebFetchTool(maxChars int, fetchLimitBytes int64) (*WebFetchTool, error) {
+type privateHostWhitelist struct {
+ exact map[string]struct{}
+ cidrs []*net.IPNet
+}
+
+func NewWebFetchTool(maxChars int, format string, fetchLimitBytes int64) (*WebFetchTool, error) {
// createHTTPClient cannot fail with an empty proxy string.
- return NewWebFetchToolWithProxy(maxChars, "", fetchLimitBytes)
+ return NewWebFetchToolWithConfig(maxChars, "", format, fetchLimitBytes, nil)
}
// allowPrivateWebFetchHosts controls whether loopback/private hosts are allowed.
// This is false in normal runtime to reduce SSRF exposure, and tests can override it temporarily.
var allowPrivateWebFetchHosts atomic.Bool
-func NewWebFetchToolWithProxy(maxChars int, proxy string, fetchLimitBytes int64) (*WebFetchTool, error) {
+func NewWebFetchToolWithProxy(
+ maxChars int,
+ proxy string,
+ format string,
+ fetchLimitBytes int64,
+ privateHostWhitelist []string,
+) (*WebFetchTool, error) {
+ return NewWebFetchToolWithConfig(maxChars, proxy, format, fetchLimitBytes, privateHostWhitelist)
+}
+
+func NewWebFetchToolWithConfig(
+ maxChars int,
+ proxy string,
+ format string,
+ fetchLimitBytes int64,
+ privateHostWhitelist []string,
+) (*WebFetchTool, error) {
if maxChars <= 0 {
maxChars = defaultMaxChars
}
+ whitelist, err := newPrivateHostWhitelist(privateHostWhitelist)
+ if err != nil {
+ return nil, fmt.Errorf("failed to parse web fetch private host whitelist: %w", err)
+ }
client, err := utils.CreateHTTPClient(proxy, fetchTimeout)
if err != nil {
return nil, fmt.Errorf("failed to create HTTP client for web fetch: %w", err)
@@ -801,13 +832,13 @@ func NewWebFetchToolWithProxy(maxChars int, proxy string, fetchLimitBytes int64)
Timeout: 15 * time.Second,
KeepAlive: 30 * time.Second,
}
- transport.DialContext = newSafeDialContext(dialer)
+ transport.DialContext = newSafeDialContext(dialer, whitelist)
}
client.CheckRedirect = func(req *http.Request, via []*http.Request) error {
if len(via) >= maxRedirects {
return fmt.Errorf("stopped after %d redirects", maxRedirects)
}
- if isObviousPrivateHost(req.URL.Hostname()) {
+ if isObviousPrivateHost(req.URL.Hostname(), whitelist) {
return fmt.Errorf("redirect target is private or local network host")
}
return nil
@@ -819,7 +850,9 @@ func NewWebFetchToolWithProxy(maxChars int, proxy string, fetchLimitBytes int64)
maxChars: maxChars,
proxy: proxy,
client: client,
+ format: format,
fetchLimitBytes: fetchLimitBytes,
+ whitelist: whitelist,
}, nil
}
@@ -871,7 +904,7 @@ func (t *WebFetchTool) Execute(ctx context.Context, args map[string]any) *ToolRe
// Lightweight pre-flight: block obvious localhost/literal-IP without DNS resolution.
// The real SSRF guard is newSafeDialContext at connect time.
hostname := parsedURL.Hostname()
- if isObviousPrivateHost(hostname) {
+ if isObviousPrivateHost(hostname, t.whitelist) {
return ErrorResult("fetching private or local network hosts is not allowed")
}
@@ -882,56 +915,128 @@ func (t *WebFetchTool) Execute(ctx context.Context, args map[string]any) *ToolRe
}
}
- req, err := http.NewRequestWithContext(ctx, "GET", urlStr, nil)
- if err != nil {
- return ErrorResult(fmt.Sprintf("failed to create request: %v", err))
+ doFetch := func(ua string) (*http.Response, []byte, error) {
+ req, reqErr := http.NewRequestWithContext(ctx, "GET", urlStr, nil)
+ if reqErr != nil {
+ return nil, nil, fmt.Errorf("failed to create request: %w", reqErr)
+ }
+ req.Header.Set("User-Agent", ua)
+ resp, doErr := t.client.Do(req)
+ if doErr != nil {
+ return nil, nil, fmt.Errorf("request failed: %w", doErr)
+ }
+ resp.Body = http.MaxBytesReader(nil, resp.Body, t.fetchLimitBytes)
+
+ b, readErr := io.ReadAll(resp.Body)
+ return resp, b, readErr
}
- req.Header.Set("User-Agent", userAgent)
- resp, err := t.client.Do(req)
- if err != nil {
- return ErrorResult(fmt.Sprintf("request failed: %v", err))
+ resp, body, err := doFetch(userAgent)
+ if resp != nil && resp.Body != nil {
+ defer resp.Body.Close()
}
- resp.Body = http.MaxBytesReader(nil, resp.Body, t.fetchLimitBytes)
-
- defer resp.Body.Close()
-
- body, err := io.ReadAll(resp.Body)
if err != nil {
var maxBytesErr *http.MaxBytesError
if errors.As(err, &maxBytesErr) {
return ErrorResult(fmt.Sprintf("failed to read response: size exceeded %d bytes limit", t.fetchLimitBytes))
}
- return ErrorResult(fmt.Sprintf("failed to read response: %v", err))
+ return ErrorResult(err.Error())
}
+ // Cloudflare (and similar WAFs) signal bot challenges with 403 + cf-mitigated: challenge.
+ // Retry once with an honest User-Agent that identifies picoclaw, which some
+ // operators explicitly allow-list for AI assistants.
+ if resp.StatusCode == http.StatusForbidden && resp.Header.Get("Cf-Mitigated") == "challenge" {
+ logger.DebugCF("tool", "Cloudflare challenge detected, retrying with honest User-Agent",
+ map[string]any{"url": urlStr})
+ honestUA := fmt.Sprintf(userAgentHonest, config.Version)
+ resp2, body2, err2 := doFetch(honestUA)
+ if resp2 != nil && resp2.Body != nil {
+ defer resp2.Body.Close()
+ }
+
+ if err2 == nil {
+ resp, body = resp2, body2
+ } else {
+ var maxBytesErr *http.MaxBytesError
+ if errors.As(err2, &maxBytesErr) {
+ return ErrorResult(
+ fmt.Sprintf("failed to read response: size exceeded %d bytes limit", t.fetchLimitBytes),
+ )
+ }
+ return ErrorResult(err2.Error())
+ }
+ }
+
+ bodyStr := string(body)
contentType := resp.Header.Get("Content-Type")
+ mediaType, params, err := mime.ParseMediaType(contentType)
+ if err != nil {
+ // The most common error here is "mime: no media type" if the header is empty.
+ logger.WarnCF("tool", "Failed to parse Content-Type", map[string]any{
+ "raw_header": contentType,
+ "error": err.Error(),
+ })
+
+ // security fallback
+ mediaType = "application/octet-stream"
+ }
+
+ charset, hasCharset := params["charset"]
+ if hasCharset {
+ // If the charset is not utf-8, we might have to convert the bodyStr
+ // before passing it to the HTML/Markdown parser
+ if strings.ToLower(charset) != "utf-8" {
+ logger.WarnCF("tool", "Note: the content is not in UTF-8", map[string]any{"charset": charset})
+ }
+ }
+
var text, extractor string
- if strings.Contains(contentType, "application/json") {
+ switch {
+ case mediaType == "application/json":
var jsonData any
- if err := json.Unmarshal(body, &jsonData); err == nil {
- formatted, _ := json.MarshalIndent(jsonData, "", " ")
- text = string(formatted)
- extractor = "json"
- } else {
- text = string(body)
+ if err := json.Unmarshal(body, &jsonData); err != nil {
+ text = bodyStr
extractor = "raw"
+ break
}
- } else if strings.Contains(contentType, "text/html") || len(body) > 0 &&
- (strings.HasPrefix(string(body), " maxChars
if truncated {
- text = text[:maxChars]
+ text = text[:maxChars] + "\n[Content truncated due to size limit]"
}
result := map[string]any{
@@ -957,6 +1062,17 @@ func (t *WebFetchTool) Execute(ctx context.Context, args map[string]any) *ToolRe
}
}
+func looksLikeHTML(body string) bool {
+ if body == "" {
+ return false
+ }
+
+ lower := strings.ToLower(body)
+
+ return strings.HasPrefix(body, "" + strings.Repeat("b", 500) + "",
+ format: "plaintext",
+ },
+ {
+ name: "html markdown extractor",
+ contentType: "text/html",
+ body: "" + strings.Repeat("c", 500) + "",
+ format: "markdown",
+ },
+ {
+ name: "json",
+ contentType: "application/json",
+ body: `"` + strings.Repeat("d", 500) + `"`,
+ format: "plaintext",
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", tt.contentType)
+ w.WriteHeader(http.StatusOK)
+ w.Write([]byte(tt.body))
+ }))
+ defer server.Close()
+
+ tool, err := NewWebFetchTool(maxChars, tt.format, testFetchLimit)
+ if err != nil {
+ t.Fatalf("NewWebFetchTool() error: %v", err)
+ }
+
+ result := tool.Execute(context.Background(), map[string]any{"url": server.URL})
+ if result.IsError {
+ t.Fatalf("unexpected error: %s", result.ForLLM)
+ }
+
+ var resultMap map[string]any
+ if err := json.Unmarshal([]byte(result.ForLLM), &resultMap); err != nil {
+ t.Fatalf("failed to unmarshal result JSON: %v", err)
+ }
+
+ text, ok := resultMap["text"].(string)
+ if !ok {
+ t.Fatal("missing 'text' field in result")
+ }
+
+ if !strings.HasSuffix(text, truncationNotice) {
+ t.Errorf("expected text to end with %q, got suffix: %q", truncationNotice, text[max(0, len(text)-60):])
+ }
+
+ if truncated, ok := resultMap["truncated"].(bool); !ok || !truncated {
+ t.Errorf("expected truncated=true in result")
+ }
+ })
+ }
+}
+
+// TestWebTool_WebFetch_NoTruncationNoticeWhenFitsInLimit verifies that the notice
+// is NOT appended when the content fits within the limit.
+func TestWebTool_WebFetch_NoTruncationNoticeWhenFitsInLimit(t *testing.T) {
+ withPrivateWebFetchHostsAllowed(t)
+
+ const truncationNotice = "[Content truncated due to size limit]"
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "text/plain")
+ w.WriteHeader(http.StatusOK)
+ w.Write([]byte("short content"))
+ }))
+ defer server.Close()
+
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
+ if err != nil {
+ t.Fatalf("NewWebFetchTool() error: %v", err)
+ }
+
+ result := tool.Execute(context.Background(), map[string]any{"url": server.URL})
+ if result.IsError {
+ t.Fatalf("unexpected error: %s", result.ForLLM)
+ }
+
+ var resultMap map[string]any
+ if err := json.Unmarshal([]byte(result.ForLLM), &resultMap); err != nil {
+ t.Fatalf("failed to unmarshal result JSON: %v", err)
+ }
+
+ text, _ := resultMap["text"].(string)
+ if strings.Contains(text, truncationNotice) {
+ t.Errorf("expected no truncation notice for content within limit, got: %q", text)
+ }
+
+ if truncated, _ := resultMap["truncated"].(bool); truncated {
+ t.Errorf("expected truncated=false for content within limit")
+ }
}
func TestWebFetchTool_PayloadTooLarge(t *testing.T) {
@@ -228,7 +358,7 @@ func TestWebFetchTool_PayloadTooLarge(t *testing.T) {
defer ts.Close()
// Initialize the tool
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
logger.ErrorCF("agent", "Failed to create web fetch tool", map[string]any{"error": err.Error()})
}
@@ -311,7 +441,7 @@ func TestWebTool_WebFetch_HTMLExtraction(t *testing.T) {
}))
defer server.Close()
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
logger.ErrorCF("agent", "Failed to create web fetch tool", map[string]any{"error": err.Error()})
}
@@ -423,8 +553,31 @@ func withPrivateWebFetchHostsAllowed(t *testing.T) {
})
}
+func serverHostAndPort(t *testing.T, rawURL string) (string, string) {
+ t.Helper()
+ hostPort := strings.TrimPrefix(rawURL, "http://")
+ hostPort = strings.TrimPrefix(hostPort, "https://")
+ host, port, err := net.SplitHostPort(hostPort)
+ if err != nil {
+ t.Fatalf("failed to split host/port from %q: %v", rawURL, err)
+ }
+ return host, port
+}
+
+func singleHostCIDR(t *testing.T, host string) string {
+ t.Helper()
+ ip := net.ParseIP(host)
+ if ip == nil {
+ t.Fatalf("failed to parse IP %q", host)
+ }
+ if ip.To4() != nil {
+ return ip.String() + "/32"
+ }
+ return ip.String() + "/128"
+}
+
func TestWebTool_WebFetch_PrivateHostBlocked(t *testing.T) {
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
t.Fatalf("Failed to create web fetch tool: %v", err)
}
@@ -441,6 +594,56 @@ func TestWebTool_WebFetch_PrivateHostBlocked(t *testing.T) {
}
}
+func TestWebTool_WebFetch_PrivateHostAllowedByExactWhitelist(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "text/plain")
+ w.WriteHeader(http.StatusOK)
+ w.Write([]byte("exact whitelist ok"))
+ }))
+ defer server.Close()
+
+ host, _ := serverHostAndPort(t, server.URL)
+ tool, err := NewWebFetchToolWithConfig(50000, "", format, testFetchLimit, []string{host})
+ if err != nil {
+ t.Fatalf("Failed to create web fetch tool: %v", err)
+ }
+
+ result := tool.Execute(context.Background(), map[string]any{
+ "url": server.URL,
+ })
+ if result.IsError {
+ t.Fatalf("expected success for exact whitelisted private IP, got %q", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "exact whitelist ok") {
+ t.Fatalf("expected fetched content, got %q", result.ForLLM)
+ }
+}
+
+func TestWebTool_WebFetch_PrivateHostAllowedByCIDRWhitelist(t *testing.T) {
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "text/plain")
+ w.WriteHeader(http.StatusOK)
+ w.Write([]byte("cidr whitelist ok"))
+ }))
+ defer server.Close()
+
+ host, _ := serverHostAndPort(t, server.URL)
+ tool, err := NewWebFetchToolWithConfig(50000, "", format, testFetchLimit, []string{singleHostCIDR(t, host)})
+ if err != nil {
+ t.Fatalf("Failed to create web fetch tool: %v", err)
+ }
+
+ result := tool.Execute(context.Background(), map[string]any{
+ "url": server.URL,
+ })
+ if result.IsError {
+ t.Fatalf("expected success for CIDR-whitelisted private IP, got %q", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "cidr whitelist ok") {
+ t.Fatalf("expected fetched content, got %q", result.ForLLM)
+ }
+}
+
func TestWebTool_WebFetch_PrivateHostAllowedForTests(t *testing.T) {
withPrivateWebFetchHostsAllowed(t)
@@ -451,7 +654,7 @@ func TestWebTool_WebFetch_PrivateHostAllowedForTests(t *testing.T) {
}))
defer server.Close()
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
t.Fatalf("Failed to create web fetch tool: %v", err)
}
@@ -466,7 +669,7 @@ func TestWebTool_WebFetch_PrivateHostAllowedForTests(t *testing.T) {
// TestWebFetch_BlocksIPv4MappedIPv6Loopback verifies ::ffff:127.0.0.1 is blocked
func TestWebFetch_BlocksIPv4MappedIPv6Loopback(t *testing.T) {
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
t.Fatalf("Failed to create web fetch tool: %v", err)
}
@@ -481,7 +684,7 @@ func TestWebFetch_BlocksIPv4MappedIPv6Loopback(t *testing.T) {
// TestWebFetch_BlocksMetadataIP verifies 169.254.169.254 is blocked
func TestWebFetch_BlocksMetadataIP(t *testing.T) {
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
t.Fatalf("Failed to create web fetch tool: %v", err)
}
@@ -496,7 +699,7 @@ func TestWebFetch_BlocksMetadataIP(t *testing.T) {
// TestWebFetch_BlocksIPv6UniqueLocal verifies fc00::/7 addresses are blocked
func TestWebFetch_BlocksIPv6UniqueLocal(t *testing.T) {
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
t.Fatalf("Failed to create web fetch tool: %v", err)
}
@@ -511,7 +714,7 @@ func TestWebFetch_BlocksIPv6UniqueLocal(t *testing.T) {
// TestWebFetch_Blocks6to4WithPrivateEmbed verifies 6to4 with private embedded IPv4 is blocked
func TestWebFetch_Blocks6to4WithPrivateEmbed(t *testing.T) {
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
t.Fatalf("Failed to create web fetch tool: %v", err)
}
@@ -527,7 +730,7 @@ func TestWebFetch_Blocks6to4WithPrivateEmbed(t *testing.T) {
// TestWebFetch_Allows6to4WithPublicEmbed verifies 6to4 with public embedded IPv4 is NOT blocked
func TestWebFetch_Allows6to4WithPublicEmbed(t *testing.T) {
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
t.Fatalf("Failed to create web fetch tool: %v", err)
}
@@ -557,7 +760,7 @@ func TestWebFetch_RedirectToPrivateBlocked(t *testing.T) {
allowPrivateWebFetchHosts.Store(false)
defer allowPrivateWebFetchHosts.Store(true)
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
t.Fatalf("Failed to create web fetch tool: %v", err)
}
@@ -570,6 +773,69 @@ func TestWebFetch_RedirectToPrivateBlocked(t *testing.T) {
}
}
+func TestNewSafeDialContext_BlocksPrivateDNSResolutionWithoutWhitelist(t *testing.T) {
+ listener, err := net.Listen("tcp", "127.0.0.1:0")
+ if err != nil {
+ t.Fatalf("failed to listen on loopback: %v", err)
+ }
+ defer listener.Close()
+
+ _, port, err := net.SplitHostPort(listener.Addr().String())
+ if err != nil {
+ t.Fatalf("failed to split listener address: %v", err)
+ }
+
+ dialContext := newSafeDialContext(&net.Dialer{Timeout: time.Second}, nil)
+ _, err = dialContext(context.Background(), "tcp", net.JoinHostPort("localhost", port))
+ if err == nil {
+ t.Fatal("expected localhost DNS resolution to be blocked without whitelist")
+ }
+ if !strings.Contains(err.Error(), "private") && !strings.Contains(err.Error(), "whitelisted") {
+ t.Fatalf("unexpected error: %v", err)
+ }
+}
+
+func TestNewSafeDialContext_AllowsWhitelistedPrivateDNSResolution(t *testing.T) {
+ listener, err := net.Listen("tcp", "127.0.0.1:0")
+ if err != nil {
+ t.Fatalf("failed to listen on loopback: %v", err)
+ }
+ defer listener.Close()
+
+ accepted := make(chan struct{}, 1)
+ go func() {
+ conn, acceptErr := listener.Accept()
+ if acceptErr != nil {
+ return
+ }
+ conn.Close()
+ accepted <- struct{}{}
+ }()
+
+ _, port, err := net.SplitHostPort(listener.Addr().String())
+ if err != nil {
+ t.Fatalf("failed to split listener address: %v", err)
+ }
+
+ whitelist, err := newPrivateHostWhitelist([]string{"127.0.0.0/8"})
+ if err != nil {
+ t.Fatalf("failed to parse whitelist: %v", err)
+ }
+
+ dialContext := newSafeDialContext(&net.Dialer{Timeout: time.Second}, whitelist)
+ conn, err := dialContext(context.Background(), "tcp", net.JoinHostPort("localhost", port))
+ if err != nil {
+ t.Fatalf("expected localhost DNS resolution to succeed with whitelist, got %v", err)
+ }
+ conn.Close()
+
+ select {
+ case <-accepted:
+ case <-time.After(time.Second):
+ t.Fatal("expected localhost listener to accept a connection")
+ }
+}
+
// TestIsPrivateOrRestrictedIP_Table tests IP classification logic
func TestIsPrivateOrRestrictedIP_Table(t *testing.T) {
tests := []struct {
@@ -615,7 +881,7 @@ func TestIsPrivateOrRestrictedIP_Table(t *testing.T) {
// TestWebTool_WebFetch_MissingDomain verifies error handling for URL without domain
func TestWebTool_WebFetch_MissingDomain(t *testing.T) {
- tool, err := NewWebFetchTool(50000, testFetchLimit)
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
if err != nil {
logger.ErrorCF("agent", "Failed to create web fetch tool", map[string]any{"error": err.Error()})
}
@@ -639,7 +905,7 @@ func TestWebTool_WebFetch_MissingDomain(t *testing.T) {
}
func TestNewWebFetchToolWithProxy(t *testing.T) {
- tool, err := NewWebFetchToolWithProxy(1024, "http://127.0.0.1:7890", testFetchLimit)
+ tool, err := NewWebFetchToolWithProxy(1024, "http://127.0.0.1:7890", format, testFetchLimit, nil)
if err != nil {
logger.ErrorCF("agent", "Failed to create web fetch tool", map[string]any{"error": err.Error()})
} else if tool.maxChars != 1024 {
@@ -650,7 +916,7 @@ func TestNewWebFetchToolWithProxy(t *testing.T) {
t.Fatalf("proxy = %q, want %q", tool.proxy, "http://127.0.0.1:7890")
}
- tool, err = NewWebFetchToolWithProxy(0, "http://127.0.0.1:7890", testFetchLimit)
+ tool, err = NewWebFetchToolWithProxy(0, "http://127.0.0.1:7890", format, testFetchLimit, nil)
if err != nil {
logger.ErrorCF("agent", "Failed to create web fetch tool", map[string]any{"error": err.Error()})
}
@@ -660,6 +926,16 @@ func TestNewWebFetchToolWithProxy(t *testing.T) {
}
}
+func TestNewWebFetchToolWithConfig_InvalidPrivateHostWhitelist(t *testing.T) {
+ _, err := NewWebFetchToolWithConfig(1024, "", format, testFetchLimit, []string{"not-an-ip-or-cidr"})
+ if err == nil {
+ t.Fatal("expected invalid whitelist entry to fail")
+ }
+ if !strings.Contains(err.Error(), "invalid entry") {
+ t.Fatalf("unexpected error: %v", err)
+ }
+}
+
func TestNewWebSearchTool_PropagatesProxy(t *testing.T) {
t.Run("perplexity", func(t *testing.T) {
tool, err := NewWebSearchTool(WebSearchToolOptions{
@@ -793,6 +1069,119 @@ func TestWebTool_TavilySearch_Success(t *testing.T) {
}
}
+// TestWebFetchTool_CloudflareChallenge_RetryWithHonestUA verifies that a 403 response
+// with cf-mitigated: challenge triggers a retry using the honest picoclaw User-Agent,
+// and that the retry response is returned when it succeeds.
+func TestWebFetchTool_CloudflareChallenge_RetryWithHonestUA(t *testing.T) {
+ withPrivateWebFetchHostsAllowed(t)
+
+ requestCount := 0
+ var receivedUAs []string
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ requestCount++
+ receivedUAs = append(receivedUAs, r.Header.Get("User-Agent"))
+
+ if requestCount == 1 {
+ // First request: simulate Cloudflare challenge
+ w.Header().Set("Cf-Mitigated", "challenge")
+ w.Header().Set("Content-Type", "text/html")
+ w.WriteHeader(http.StatusForbidden)
+ w.Write([]byte("Cloudflare challenge"))
+ return
+ }
+ // Second request (honest UA retry): success
+ w.Header().Set("Content-Type", "text/plain")
+ w.WriteHeader(http.StatusOK)
+ w.Write([]byte("real content"))
+ }))
+ defer server.Close()
+
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
+ if err != nil {
+ t.Fatalf("NewWebFetchTool() error: %v", err)
+ }
+
+ result := tool.Execute(context.Background(), map[string]any{"url": server.URL})
+
+ if result.IsError {
+ t.Fatalf("expected success after retry, got error: %s", result.ForLLM)
+ }
+ if !strings.Contains(result.ForLLM, "real content") {
+ t.Errorf("expected retry response content, got: %s", result.ForLLM)
+ }
+ if requestCount != 2 {
+ t.Errorf("expected exactly 2 requests, got %d", requestCount)
+ }
+
+ // First request must use the generic user agent
+ if receivedUAs[0] != userAgent {
+ t.Errorf("first request UA = %q, want %q", receivedUAs[0], userAgent)
+ }
+ // Second request must use the honest picoclaw user agent
+ if !strings.Contains(receivedUAs[1], "picoclaw") {
+ t.Errorf("retry request UA = %q, want it to contain 'picoclaw'", receivedUAs[1])
+ }
+}
+
+// TestWebFetchTool_CloudflareChallenge_NoRetryOnOtherErrors verifies that a plain 403
+// (without cf-mitigated: challenge) does NOT trigger a retry.
+func TestWebFetchTool_CloudflareChallenge_NoRetryOnOtherErrors(t *testing.T) {
+ withPrivateWebFetchHostsAllowed(t)
+
+ requestCount := 0
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ requestCount++
+ w.Header().Set("Content-Type", "text/plain")
+ w.WriteHeader(http.StatusForbidden)
+ w.Write([]byte("plain forbidden"))
+ }))
+ defer server.Close()
+
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
+ if err != nil {
+ t.Fatalf("NewWebFetchTool() error: %v", err)
+ }
+
+ tool.Execute(context.Background(), map[string]any{"url": server.URL})
+
+ if requestCount != 1 {
+ t.Errorf("expected exactly 1 request for plain 403, got %d", requestCount)
+ }
+}
+
+// TestWebFetchTool_CloudflareChallenge_RetryFailsToo verifies that if the honest-UA
+// retry also fails (e.g. still blocked), the error from the retry is returned.
+func TestWebFetchTool_CloudflareChallenge_RetryFailsToo(t *testing.T) {
+ withPrivateWebFetchHostsAllowed(t)
+
+ server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ // Always return CF challenge regardless of UA
+ w.Header().Set("Cf-Mitigated", "challenge")
+ w.Header().Set("Content-Type", "text/html")
+ w.WriteHeader(http.StatusForbidden)
+ w.Write([]byte("still blocked"))
+ }))
+ defer server.Close()
+
+ tool, err := NewWebFetchTool(50000, format, testFetchLimit)
+ if err != nil {
+ t.Fatalf("NewWebFetchTool() error: %v", err)
+ }
+
+ result := tool.Execute(context.Background(), map[string]any{"url": server.URL})
+
+ // Should not be an error — the retry response is used as-is (403 is a valid HTTP response)
+ if result.IsError {
+ t.Fatalf("expected non-error result even when retry is also blocked, got: %s", result.ForLLM)
+ }
+ // Status in the JSON result should reflect the 403
+ if !strings.Contains(result.ForLLM, "403") {
+ t.Errorf("expected status 403 in result, got: %s", result.ForLLM)
+ }
+}
+
func TestAPIKeyPool(t *testing.T) {
pool := NewAPIKeyPool([]string{"key1", "key2", "key3"})
if len(pool.keys) != 3 {
diff --git a/pkg/utils/context.go b/pkg/utils/context.go
new file mode 100644
index 000000000..2007de9a3
--- /dev/null
+++ b/pkg/utils/context.go
@@ -0,0 +1,173 @@
+// PicoClaw - Ultra-lightweight personal AI agent
+// Inspired by and based on nanobot: https://github.com/HKUDS/nanobot
+// License: MIT
+//
+// Copyright (c) 2026 PicoClaw contributors
+
+package utils
+
+import (
+ "encoding/json"
+ "fmt"
+ "unicode/utf8"
+
+ "github.com/sipeed/picoclaw/pkg/providers"
+)
+
+// CalculateDefaultMaxContextRunes computes a default context limit based on the model's context window.
+// Strategy: Use 75% of the context window and convert to rune estimate.
+//
+// Token-to-rune conversion ratios (conservative estimates):
+// - English: ~4 chars per token
+// - Chinese: ~1.5-2 chars per token
+// - Mixed: ~3 chars per token (used here for safety)
+func CalculateDefaultMaxContextRunes(contextWindow int) int {
+ if contextWindow <= 0 {
+ // Conservative fallback when context window is unknown
+ return 8000 // ~2000 tokens
+ }
+
+ // Use 75% of context window to leave headroom
+ targetTokens := int(float64(contextWindow) * 0.75)
+
+ // Convert tokens to runes using conservative ratio
+ const avgCharsPerToken = 3
+ return targetTokens * avgCharsPerToken
+}
+
+// ResolveMaxContextRunes determines the final MaxContextRunes value to use.
+// Priority: explicit config > auto-calculate > conservative default
+func ResolveMaxContextRunes(configValue, contextWindow int) int {
+ switch {
+ case configValue > 0:
+ // Explicitly configured, use as-is
+ return configValue
+ case configValue == -1:
+ // Explicitly disabled
+ return -1
+ default:
+ // 0 or unset: auto-calculate
+ return CalculateDefaultMaxContextRunes(contextWindow)
+ }
+}
+
+// MeasureContextRunes calculates the total rune count of a message list.
+// Includes content, reasoning content, and estimates for tool calls.
+func MeasureContextRunes(messages []providers.Message) int {
+ totalRunes := 0
+ for _, msg := range messages {
+ totalRunes += utf8.RuneCountInString(msg.Content)
+ totalRunes += utf8.RuneCountInString(msg.ReasoningContent)
+
+ // Tool calls: serialize to JSON and count
+ if len(msg.ToolCalls) > 0 {
+ for _, tc := range msg.ToolCalls {
+ totalRunes += utf8.RuneCountInString(tc.Name)
+ // Arguments: serialize and count
+ if argsJSON, err := json.Marshal(tc.Arguments); err == nil {
+ totalRunes += utf8.RuneCount(argsJSON)
+ } else {
+ // Fallback estimate if serialization fails
+ totalRunes += 100
+ }
+ }
+ }
+
+ // ToolCallID
+ totalRunes += utf8.RuneCountInString(msg.ToolCallID)
+ }
+ return totalRunes
+}
+
+// TruncateContextSmart intelligently truncates message history to fit within maxRunes.
+//
+// Strategy:
+// 1. Always preserve system messages (they define the agent's behavior)
+// 2. Keep the most recent messages (they contain current context)
+// 3. Drop older middle messages when necessary
+// 4. Insert a truncation notice to inform the LLM
+//
+// Returns the truncated message list.
+func TruncateContextSmart(messages []providers.Message, maxRunes int) []providers.Message {
+ if len(messages) == 0 {
+ return messages
+ }
+
+ // Separate system messages from others
+ var systemMsgs []providers.Message
+ var otherMsgs []providers.Message
+
+ for _, msg := range messages {
+ if msg.Role == "system" {
+ systemMsgs = append(systemMsgs, msg)
+ } else {
+ otherMsgs = append(otherMsgs, msg)
+ }
+ }
+
+ // Calculate system message size
+ systemRunes := 0
+ for _, msg := range systemMsgs {
+ systemRunes += utf8.RuneCountInString(msg.Content)
+ systemRunes += utf8.RuneCountInString(msg.ReasoningContent)
+ }
+
+ // Reserve space for truncation notice (estimate ~80 runes)
+ const truncationNoticeEstimate = 80
+
+ // Allocate remaining space for other messages
+ remainingRunes := maxRunes - systemRunes - truncationNoticeEstimate
+ if remainingRunes <= 0 {
+ // System messages already exceed limit - return only system messages
+ return systemMsgs
+ }
+
+ // Collect recent messages in reverse order until we hit the limit
+ var keptMsgs []providers.Message
+ currentRunes := 0
+
+ for i := len(otherMsgs) - 1; i >= 0; i-- {
+ msg := otherMsgs[i]
+ msgRunes := utf8.RuneCountInString(msg.Content) +
+ utf8.RuneCountInString(msg.ReasoningContent)
+
+ // Estimate tool call size
+ if len(msg.ToolCalls) > 0 {
+ for _, tc := range msg.ToolCalls {
+ msgRunes += utf8.RuneCountInString(tc.Name)
+ if argsJSON, err := json.Marshal(tc.Arguments); err == nil {
+ msgRunes += utf8.RuneCount(argsJSON)
+ } else {
+ msgRunes += 100
+ }
+ }
+ }
+ msgRunes += utf8.RuneCountInString(msg.ToolCallID)
+
+ if currentRunes+msgRunes > remainingRunes {
+ // Would exceed limit, stop collecting
+ break
+ }
+
+ // Prepend to maintain chronological order
+ keptMsgs = append([]providers.Message{msg}, keptMsgs...)
+ currentRunes += msgRunes
+ }
+
+ // If we dropped messages, add a truncation notice
+ result := systemMsgs
+ if len(keptMsgs) < len(otherMsgs) {
+ droppedCount := len(otherMsgs) - len(keptMsgs)
+ truncationNotice := providers.Message{
+ Role: "system",
+ Content: fmt.Sprintf(
+ "[Context truncated: %d earlier messages omitted to stay within context limits]",
+ droppedCount,
+ ),
+ }
+ result = append(result, truncationNotice)
+ }
+
+ result = append(result, keptMsgs...)
+ return result
+}
diff --git a/pkg/utils/context_test.go b/pkg/utils/context_test.go
new file mode 100644
index 000000000..450a29249
--- /dev/null
+++ b/pkg/utils/context_test.go
@@ -0,0 +1,450 @@
+// PicoClaw - Ultra-lightweight personal AI agent
+// License: MIT
+//
+// Copyright (c) 2026 PicoClaw contributors
+
+package utils
+
+import (
+ "testing"
+
+ "github.com/sipeed/picoclaw/pkg/providers"
+)
+
+func TestCalculateDefaultMaxContextRunes(t *testing.T) {
+ tests := []struct {
+ name string
+ contextWindow int
+ want int
+ }{
+ {
+ name: "zero context window uses fallback",
+ contextWindow: 0,
+ want: 8000,
+ },
+ {
+ name: "negative context window uses fallback",
+ contextWindow: -1,
+ want: 8000,
+ },
+ {
+ name: "small context window (4k tokens)",
+ contextWindow: 4000,
+ want: 9000, // 4000 * 0.75 * 3 = 9000
+ },
+ {
+ name: "medium context window (128k tokens)",
+ contextWindow: 128000,
+ want: 288000, // 128000 * 0.75 * 3 = 288000
+ },
+ {
+ name: "large context window (1M tokens)",
+ contextWindow: 1000000,
+ want: 2250000, // 1000000 * 0.75 * 3 = 2250000
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ got := CalculateDefaultMaxContextRunes(tt.contextWindow)
+ if got != tt.want {
+ t.Errorf("CalculateDefaultMaxContextRunes(%d) = %d, want %d",
+ tt.contextWindow, got, tt.want)
+ }
+ })
+ }
+}
+
+func TestResolveMaxContextRunes(t *testing.T) {
+ tests := []struct {
+ name string
+ configValue int
+ contextWindow int
+ want int
+ }{
+ {
+ name: "explicit positive value",
+ configValue: 12000,
+ contextWindow: 4000,
+ want: 12000,
+ },
+ {
+ name: "explicit disable (-1)",
+ configValue: -1,
+ contextWindow: 4000,
+ want: -1,
+ },
+ {
+ name: "zero uses auto-calculate",
+ configValue: 0,
+ contextWindow: 4000,
+ want: 9000, // 4000 * 0.75 * 3
+ },
+ {
+ name: "unset (0) with unknown context window",
+ configValue: 0,
+ contextWindow: 0,
+ want: 8000, // fallback
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ got := ResolveMaxContextRunes(tt.configValue, tt.contextWindow)
+ if got != tt.want {
+ t.Errorf("ResolveMaxContextRunes(%d, %d) = %d, want %d",
+ tt.configValue, tt.contextWindow, got, tt.want)
+ }
+ })
+ }
+}
+
+func TestMeasureContextRunes(t *testing.T) {
+ tests := []struct {
+ name string
+ messages []providers.Message
+ want int
+ }{
+ {
+ name: "empty messages",
+ messages: []providers.Message{},
+ want: 0,
+ },
+ {
+ name: "single simple message",
+ messages: []providers.Message{
+ {Role: "user", Content: "Hello"},
+ },
+ want: 5, // "Hello" = 5 runes
+ },
+ {
+ name: "message with reasoning",
+ messages: []providers.Message{
+ {
+ Role: "assistant",
+ Content: "Answer",
+ ReasoningContent: "Thinking",
+ },
+ },
+ want: 14, // "Answer" (6) + "Thinking" (8) = 14
+ },
+ {
+ name: "message with tool call",
+ messages: []providers.Message{
+ {
+ Role: "assistant",
+ Content: "Using tool",
+ ToolCalls: []providers.ToolCall{
+ {
+ Name: "test_tool",
+ Arguments: map[string]any{"key": "value"},
+ },
+ },
+ },
+ },
+ want: 10 + 9 + 15, // "Using tool" + "test_tool" + {"key":"value"}
+ },
+ {
+ name: "multiple messages",
+ messages: []providers.Message{
+ {Role: "system", Content: "You are helpful"},
+ {Role: "user", Content: "Hi"},
+ {Role: "assistant", Content: "Hello!"},
+ },
+ want: 15 + 2 + 6, // 15 + 2 + 6 = 23
+ },
+ {
+ name: "unicode characters",
+ messages: []providers.Message{
+ {Role: "user", Content: "\u4f60\u597d\u4e16\u754c"}, // 4 Chinese characters
+ },
+ want: 4,
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ got := MeasureContextRunes(tt.messages)
+ if got != tt.want {
+ t.Errorf("MeasureContextRunes() = %d, want %d", got, tt.want)
+ }
+ })
+ }
+}
+
+func TestTruncateContextSmart(t *testing.T) {
+ tests := []struct {
+ name string
+ messages []providers.Message
+ maxRunes int
+ wantLen int
+ wantHas []string // Content strings that should be present
+ wantNot []string // Content strings that should be absent
+ }{
+ {
+ name: "empty messages",
+ messages: []providers.Message{},
+ maxRunes: 100,
+ wantLen: 0,
+ },
+ {
+ name: "no truncation needed",
+ messages: []providers.Message{
+ {Role: "system", Content: "System"},
+ {Role: "user", Content: "Hello"},
+ },
+ maxRunes: 100,
+ wantLen: 2,
+ wantHas: []string{"System", "Hello"},
+ },
+ {
+ name: "truncate when limit is tight",
+ messages: []providers.Message{
+ {Role: "system", Content: "System"},
+ {Role: "user", Content: "Message 1 with some content here"},
+ {Role: "assistant", Content: "Response 1 with some content here"},
+ {Role: "user", Content: "Message 2 with some content here"},
+ {Role: "assistant", Content: "Response 2 with some content here"},
+ {Role: "user", Content: "Latest"},
+ },
+ maxRunes: 120, // Tight limit to force truncation
+ wantLen: -1, // Don't check exact length, just verify truncation occurred
+ wantHas: []string{"System", "Latest"},
+ wantNot: []string{"Message 1", "Response 1"},
+ },
+ {
+ name: "system messages exceed limit",
+ messages: []providers.Message{
+ {Role: "system", Content: "Very long system message"},
+ {Role: "user", Content: "User message"},
+ },
+ maxRunes: 10, // Less than system message
+ wantLen: 1, // Only system message
+ wantHas: []string{"Very long system message"},
+ wantNot: []string{"User message"},
+ },
+ {
+ name: "preserve multiple system messages",
+ messages: []providers.Message{
+ {Role: "system", Content: "Sys1"},
+ {Role: "system", Content: "Sys2"},
+ {Role: "user", Content: "Old"},
+ {Role: "user", Content: "New"},
+ },
+ maxRunes: 200, // Generous limit
+ wantLen: 4, // Both system + truncation notice + new
+ wantHas: []string{"Sys1", "Sys2", "New"},
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ got := TruncateContextSmart(tt.messages, tt.maxRunes)
+
+ if tt.wantLen >= 0 && len(got) != tt.wantLen {
+ t.Errorf("TruncateContextSmart() returned %d messages, want %d",
+ len(got), tt.wantLen)
+ }
+
+ // Check for expected content
+ allContent := ""
+ for _, msg := range got {
+ allContent += msg.Content + " "
+ }
+
+ for _, want := range tt.wantHas {
+ found := false
+ for _, msg := range got {
+ if msg.Content == want || containsSubstring(msg.Content, want) {
+ found = true
+ break
+ }
+ }
+ if !found {
+ t.Errorf("Expected content %q not found in truncated messages", want)
+ }
+ }
+
+ for _, notWant := range tt.wantNot {
+ for _, msg := range got {
+ if containsSubstring(msg.Content, notWant) {
+ t.Errorf("Unexpected content %q found in truncated messages", notWant)
+ }
+ }
+ }
+ })
+ }
+}
+
+func containsSubstring(s, substr string) bool {
+ return len(s) >= len(substr) && findSubstring(s, substr)
+}
+
+func findSubstring(s, substr string) bool {
+ for i := 0; i <= len(s)-len(substr); i++ {
+ if s[i:i+len(substr)] == substr {
+ return true
+ }
+ }
+ return false
+}
+
+// TestSubTurnConfigMaxContextRunes verifies that MaxContextRunes configuration
+// is properly integrated into the SubTurn execution flow.
+func TestSubTurnConfigMaxContextRunes(t *testing.T) {
+ tests := []struct {
+ name string
+ maxContextRunes int
+ contextWindow int
+ wantResolved int
+ }{
+ {
+ name: "default (0) auto-calculates from context window",
+ maxContextRunes: 0,
+ contextWindow: 4000,
+ wantResolved: 9000, // 4000 * 0.75 * 3
+ },
+ {
+ name: "explicit value is used",
+ maxContextRunes: 12000,
+ contextWindow: 4000,
+ wantResolved: 12000,
+ },
+ {
+ name: "disabled (-1) returns -1",
+ maxContextRunes: -1,
+ contextWindow: 4000,
+ wantResolved: -1,
+ },
+ {
+ name: "fallback when context window unknown",
+ maxContextRunes: 0,
+ contextWindow: 0,
+ wantResolved: 8000, // conservative fallback
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ got := ResolveMaxContextRunes(tt.maxContextRunes, tt.contextWindow)
+ if got != tt.wantResolved {
+ t.Errorf("utils.ResolveMaxContextRunes(%d, %d) = %d, want %d",
+ tt.maxContextRunes, tt.contextWindow, got, tt.wantResolved)
+ }
+ })
+ }
+}
+
+// TestContextTruncationFlow verifies the complete context truncation flow:
+// 1. Messages accumulate beyond soft limit
+// 2. Truncation is triggered
+// 3. System messages are preserved
+// 4. Recent messages are kept
+func TestContextTruncationFlow(t *testing.T) {
+ // Build a message history that exceeds the limit
+ messages := []providers.Message{
+ {Role: "system", Content: "You are a helpful assistant"}, // ~27 runes
+ {Role: "user", Content: "First question"}, // ~14 runes
+ {Role: "assistant", Content: "First answer"}, // ~12 runes
+ {Role: "user", Content: "Second question"}, // ~15 runes
+ {Role: "assistant", Content: "Second answer"}, // ~13 runes
+ {Role: "user", Content: "Third question"}, // ~14 runes
+ {Role: "assistant", Content: "Third answer"}, // ~12 runes
+ {Role: "user", Content: "Latest question"}, // ~15 runes
+ }
+
+ // Total: ~122 runes
+ totalRunes := MeasureContextRunes(messages)
+ if totalRunes < 100 {
+ t.Errorf("Expected total runes > 100, got %d", totalRunes)
+ }
+
+ // Set limit to 150 runes - should force truncation of old messages
+ // but preserve system + truncation notice + recent messages
+ maxRunes := 150
+ truncated := TruncateContextSmart(messages, maxRunes)
+
+ // Verify truncation occurred
+ if len(truncated) >= len(messages) {
+ t.Errorf("Expected truncation, but got %d messages (original: %d)",
+ len(truncated), len(messages))
+ }
+
+ // Verify system message is preserved
+ foundSystem := false
+ for _, msg := range truncated {
+ if msg.Role == "system" && msg.Content == "You are a helpful assistant" {
+ foundSystem = true
+ break
+ }
+ }
+ if !foundSystem {
+ t.Error("System message was not preserved after truncation")
+ }
+
+ // Verify latest message is preserved
+ foundLatest := false
+ for _, msg := range truncated {
+ if msg.Content == "Latest question" {
+ foundLatest = true
+ break
+ }
+ }
+ if !foundLatest {
+ t.Error("Latest message was not preserved after truncation")
+ }
+
+ // Verify truncation notice is present
+ foundNotice := false
+ for _, msg := range truncated {
+ if msg.Role == "system" && containsSubstring(msg.Content, "truncated") {
+ foundNotice = true
+ break
+ }
+ }
+ if !foundNotice {
+ t.Error("Truncation notice was not added")
+ }
+
+ // Verify result is within limit (with some tolerance for estimation)
+ resultRunes := MeasureContextRunes(truncated)
+ if resultRunes > maxRunes+20 { // Allow 20 rune tolerance
+ t.Errorf("Truncated context (%d runes) significantly exceeds limit (%d runes)",
+ resultRunes, maxRunes)
+ }
+}
+
+// TestContextTruncationPreservesToolCalls verifies that tool calls are
+// properly handled during context truncation.
+func TestContextTruncationPreservesToolCalls(t *testing.T) {
+ messages := []providers.Message{
+ {Role: "system", Content: "System"},
+ {Role: "user", Content: "Old message that should be dropped"},
+ {
+ Role: "assistant",
+ Content: "Recent tool use",
+ ToolCalls: []providers.ToolCall{
+ {
+ Name: "important_tool",
+ Arguments: map[string]any{"key": "value"},
+ },
+ },
+ },
+ }
+
+ // Set a generous limit that should keep the tool call message
+ maxRunes := 200
+ truncated := TruncateContextSmart(messages, maxRunes)
+
+ // Verify tool call message is preserved
+ foundToolCall := false
+ for _, msg := range truncated {
+ if len(msg.ToolCalls) > 0 && msg.ToolCalls[0].Name == "important_tool" {
+ foundToolCall = true
+ break
+ }
+ }
+ if !foundToolCall {
+ t.Error("Tool call message was not preserved during truncation")
+ }
+}
diff --git a/pkg/utils/markdown.go b/pkg/utils/markdown.go
new file mode 100644
index 000000000..c7873252a
--- /dev/null
+++ b/pkg/utils/markdown.go
@@ -0,0 +1,411 @@
+package utils
+
+import (
+ "bytes"
+ "net/url"
+ "regexp"
+ "strconv"
+ "strings"
+
+ "golang.org/x/net/html"
+)
+
+var (
+ reSpaces = regexp.MustCompile(`[ \t]+`)
+ reNewlines = regexp.MustCompile(`\n{3,}`)
+ reEmptyListItem = regexp.MustCompile(`(?m)^[-*]\s*$`)
+ reImageOnlyLink = regexp.MustCompile(`\[!\[\]\(<[^>]*>\)\]\(<[^>]*>\)`)
+ reEmptyHeader = regexp.MustCompile(`(?m)^#{1,6}\s*$`)
+ reLeadingLineSpace = regexp.MustCompile(`(?m)^([ \t])([^ \t\n])`)
+)
+
+var skipTags = map[string]bool{
+ "script": true, "style": true, "head": true,
+ "noscript": true, "template": true,
+ "nav": true, "footer": true, "aside": true, "header": true, "form": true, "dialog": true,
+}
+
+func isSafeHref(href string) bool {
+ lower := strings.ToLower(strings.TrimSpace(href))
+ if strings.HasPrefix(lower, "javascript:") || strings.HasPrefix(lower, "vbscript:") ||
+ strings.HasPrefix(lower, "data:") {
+ return false
+ }
+ u, err := url.Parse(strings.TrimSpace(href))
+ if err != nil {
+ return false
+ }
+ scheme := strings.ToLower(u.Scheme)
+ return scheme == "" || scheme == "http" || scheme == "https" || scheme == "mailto"
+}
+
+func isSafeImageSrc(src string) bool {
+ lower := strings.ToLower(strings.TrimSpace(src))
+ if strings.HasPrefix(lower, "data:image/") {
+ return true
+ }
+ return isSafeHref(src)
+}
+
+func escapeMdAlt(s string) string {
+ s = strings.ReplaceAll(s, `\`, `\\`)
+ s = strings.ReplaceAll(s, `[`, `\[`)
+ s = strings.ReplaceAll(s, `]`, `\]`)
+ return s
+}
+
+func getAttr(n *html.Node, key string) string {
+ for _, a := range n.Attr {
+ if a.Key == key {
+ return a.Val
+ }
+ }
+ return ""
+}
+
+func normalizeAttr(val string) string {
+ val = strings.ReplaceAll(val, "\n", "")
+ val = strings.ReplaceAll(val, "\r", "")
+ val = strings.ReplaceAll(val, "\t", "")
+ return strings.TrimSpace(val)
+}
+
+func isUnlikelyNode(n *html.Node) bool {
+ if n.Type != html.ElementNode {
+ return false
+ }
+ classId := strings.ToLower(getAttr(n, "class") + " " + getAttr(n, "id"))
+ if classId == " " {
+ return false
+ }
+ if strings.Contains(classId, "article") || strings.Contains(classId, "main") ||
+ strings.Contains(classId, "content") {
+ return false
+ }
+ unlikelyKeywords := []string{
+ "menu",
+ "nav",
+ "footer",
+ "sidebar",
+ "cookie",
+ "banner",
+ "sponsor",
+ "advert",
+ "popup",
+ "modal",
+ "newsletter",
+ "share",
+ "social",
+ }
+ for _, keyword := range unlikelyKeywords {
+ if strings.Contains(classId, keyword) {
+ return true
+ }
+ }
+ return false
+}
+
+type converter struct {
+ stack []*bytes.Buffer
+ linkHrefs []string
+ linkStates []bool
+ emphStack []string // Tracks "**", "*", "~~" for buffered emphasis
+ olCounters []int
+ inPre bool
+ listDepth int
+}
+
+func newConverter() *converter {
+ return &converter{
+ stack: []*bytes.Buffer{{}},
+ }
+}
+
+func (c *converter) write(s string) {
+ c.stack[len(c.stack)-1].WriteString(s)
+}
+
+func (c *converter) pushBuf() {
+ c.stack = append(c.stack, &bytes.Buffer{})
+}
+
+func (c *converter) popBuf() string {
+ top := c.stack[len(c.stack)-1]
+ c.stack = c.stack[:len(c.stack)-1]
+ return top.String()
+}
+
+func (c *converter) walk(n *html.Node) {
+ if n.Type == html.ElementNode {
+ if skipTags[n.Data] {
+ return
+ }
+ if isUnlikelyNode(n) {
+ return
+ }
+ }
+
+ if n.Type == html.TextNode {
+ text := n.Data
+ if !c.inPre {
+ text = strings.ReplaceAll(text, "\n", " ")
+ text = reSpaces.ReplaceAllString(text, " ")
+ }
+ if text != "" {
+ c.write(text)
+ }
+ return
+ }
+
+ if n.Type != html.ElementNode {
+ for ch := n.FirstChild; ch != nil; ch = ch.NextSibling {
+ c.walk(ch)
+ }
+ return
+ }
+
+ // Opening Tags
+ switch n.Data {
+ // Buffer emphasis content so we can TrimSpace the inner text,
+ // avoiding the regex-across-boundaries bug.
+ case "b", "strong":
+ c.emphStack = append(c.emphStack, "**")
+ c.pushBuf()
+ case "i", "em":
+ c.emphStack = append(c.emphStack, "*")
+ c.pushBuf()
+ case "del", "s":
+ c.emphStack = append(c.emphStack, "~~")
+ c.pushBuf()
+
+ case "a":
+ href := normalizeAttr(getAttr(n, "href"))
+ if href != "" && !isSafeHref(href) {
+ href = "#"
+ }
+ hasHref := href != ""
+ c.linkStates = append(c.linkStates, hasHref)
+ if hasHref {
+ c.linkHrefs = append(c.linkHrefs, href)
+ c.pushBuf()
+ }
+
+ case "h1":
+ c.write("\n\n# ")
+ case "h2":
+ c.write("\n\n## ")
+ case "h3":
+ c.write("\n\n### ")
+ case "h4":
+ c.write("\n\n#### ")
+ case "h5":
+ c.write("\n\n##### ")
+ case "h6":
+ c.write("\n\n###### ")
+
+ case "p":
+ c.write("\n\n")
+ case "br":
+ c.write("\n")
+ case "hr":
+ c.write("\n\n---\n\n")
+
+ case "ol":
+ c.olCounters = append(c.olCounters, 1)
+ // Only write leading newline for top-level list.
+ if c.listDepth == 0 {
+ c.write("\n")
+ }
+ c.listDepth++
+ case "ul":
+ if c.listDepth == 0 {
+ c.write("\n")
+ }
+ c.listDepth++
+ case "li":
+ c.write("\n")
+ if c.listDepth > 1 {
+ c.write(strings.Repeat(" ", c.listDepth-1))
+ }
+ if n.Parent != nil && n.Parent.Data == "ol" && len(c.olCounters) > 0 {
+ idx := c.olCounters[len(c.olCounters)-1]
+ c.write(strconv.Itoa(idx) + ". ")
+ c.olCounters[len(c.olCounters)-1]++
+ } else {
+ c.write("- ")
+ }
+
+ case "pre":
+ c.inPre = true
+ c.write("\n\n```\n")
+ case "code":
+ if !c.inPre {
+ c.write("`")
+ }
+
+ case "blockquote":
+ c.pushBuf()
+ for ch := n.FirstChild; ch != nil; ch = ch.NextSibling {
+ c.walk(ch)
+ }
+ inner := strings.TrimSpace(c.popBuf())
+ lines := strings.Split(inner, "\n")
+ var quoted []string
+ for _, l := range lines {
+ if strings.TrimSpace(l) == "" {
+ quoted = append(quoted, ">")
+ } else {
+ quoted = append(quoted, "> "+l)
+ }
+ }
+ var deduped []string
+ for i, line := range quoted {
+ if line == ">" && i > 0 && deduped[len(deduped)-1] == ">" {
+ continue
+ }
+ deduped = append(deduped, line)
+ }
+ c.write("\n\n" + strings.Join(deduped, "\n") + "\n\n")
+ return
+
+ case "img":
+ src := normalizeAttr(getAttr(n, "src"))
+ if src == "" {
+ src = normalizeAttr(getAttr(n, "data-src"))
+ }
+ if src == "" {
+ return
+ }
+ alt := escapeMdAlt(normalizeAttr(getAttr(n, "alt")))
+ if isSafeImageSrc(src) {
+ c.write("")
+ }
+ return
+ }
+
+ // Traverse Children
+ for ch := n.FirstChild; ch != nil; ch = ch.NextSibling {
+ c.walk(ch)
+ }
+
+ // Closing Tags
+ switch n.Data {
+ // Pop buffer, trim, wrap with the correct marker.
+ case "b", "strong", "i", "em", "del", "s":
+ if len(c.emphStack) == 0 {
+ break
+ }
+ marker := c.emphStack[len(c.emphStack)-1]
+ c.emphStack = c.emphStack[:len(c.emphStack)-1]
+ inner := strings.TrimSpace(c.popBuf())
+ if inner != "" {
+ c.write(marker + inner + marker)
+ }
+
+ case "a":
+ if len(c.linkStates) == 0 {
+ break
+ }
+ hasHref := c.linkStates[len(c.linkStates)-1]
+ c.linkStates = c.linkStates[:len(c.linkStates)-1]
+ if !hasHref {
+ break
+ }
+ href := c.linkHrefs[len(c.linkHrefs)-1]
+ c.linkHrefs = c.linkHrefs[:len(c.linkHrefs)-1]
+ inner := strings.TrimSpace(c.popBuf())
+ if strings.Contains(inner, "\n") {
+ lines := strings.Split(inner, "\n")
+ linked := false
+ for i, l := range lines {
+ cleanLine := strings.TrimSpace(l)
+ if cleanLine != "" && !strings.HasPrefix(cleanLine, "![") && !linked {
+ lines[i] = "[" + cleanLine + "](" + href + ")"
+ linked = true
+ }
+ }
+ c.write(strings.Join(lines, "\n"))
+ } else {
+ c.write("[" + inner + "](" + href + ")")
+ }
+
+ case "h1",
+ "h2",
+ "h3",
+ "h4",
+ "h5",
+ "h6",
+ "p",
+ "div",
+ "section",
+ "article",
+ "header",
+ "footer",
+ "aside",
+ "nav",
+ "figure":
+ c.write("\n")
+
+ case "ol":
+ c.listDepth--
+ if len(c.olCounters) > 0 {
+ c.olCounters = c.olCounters[:len(c.olCounters)-1]
+ }
+ if c.listDepth == 0 {
+ c.write("\n")
+ }
+ case "ul":
+ c.listDepth--
+ if c.listDepth == 0 {
+ c.write("\n")
+ }
+
+ case "pre":
+ c.inPre = false
+ c.write("\n```\n\n")
+ case "code":
+ if !c.inPre {
+ c.write("`")
+ }
+ }
+}
+
+func HtmlToMarkdown(htmlStr string) (string, error) {
+ doc, err := html.Parse(strings.NewReader(htmlStr))
+ if err != nil {
+ return "", err
+ }
+
+ c := newConverter()
+ c.walk(doc)
+
+ res := c.stack[0].String()
+
+ // Post-processing
+ res = reImageOnlyLink.ReplaceAllString(res, "")
+ res = reEmptyListItem.ReplaceAllString(res, "")
+ res = reEmptyHeader.ReplaceAllString(res, "")
+
+ lines := strings.Split(res, "\n")
+ var cleanLines []string
+ for _, line := range lines {
+ line = strings.TrimRight(line, " \t")
+ cleanTest := strings.TrimSpace(line)
+ if cleanTest == "[](>)" || cleanTest == "[](#)" || cleanTest == "-" {
+ cleanLines = append(cleanLines, "")
+ continue
+ }
+ cleanLines = append(cleanLines, line)
+ }
+ res = strings.Join(cleanLines, "\n")
+
+ res = strings.TrimSpace(res)
+ res = reNewlines.ReplaceAllString(res, "\n\n")
+
+ // Strip a single leading space from lines that are NOT list indentation.
+ // "(?m)^([ \t])([^ \t\n])" matches exactly one space/tab at line start followed
+ // by a non-whitespace char, so " - nested" (4 spaces) is left untouched.
+ res = reLeadingLineSpace.ReplaceAllString(res, "$2")
+
+ return res, nil
+}
diff --git a/pkg/utils/markdown_test.go b/pkg/utils/markdown_test.go
new file mode 100644
index 000000000..72277fb91
--- /dev/null
+++ b/pkg/utils/markdown_test.go
@@ -0,0 +1,245 @@
+package utils
+
+import (
+ "testing"
+
+ "github.com/sipeed/picoclaw/pkg/logger"
+)
+
+func TestHtmlToMarkdown(t *testing.T) {
+ // Define our test cases
+ tests := []struct {
+ name string
+ input string
+ expected string
+ }{
+ {
+ name: "Removes scripts and styles",
+ input: `
`,
+ expected: "# Main Title\n\n## Subtitle\n\n### Section",
+ },
+ {
+ name: "Handles bold and italics",
+ input: `Text bold and strong, then italic and em.`,
+ expected: "Text **bold** and **strong**, then *italic* and *em*.",
+ },
+ {
+ name: "Converts lists",
+ input: `
First element
Second element
`,
+ expected: "- First element\n- Second element",
+ },
+ {
+ name: "Handles paragraphs and line breaks ( )",
+ input: `
First paragraph
Second paragraph with a line break.
`,
+ expected: "First paragraph\n\nSecond paragraph with\na line break.",
+ },
+ {
+ name: "Decodes HTML entities",
+ input: `Math: 5 > 3 & 2 < 4. A "quote".`,
+ expected: "Math: 5 > 3 & 2 < 4. A \"quote\".",
+ },
+ {
+ name: "Cleans up residual HTML tags",
+ input: `
Text inside div and span
`,
+ expected: "Text inside div and span",
+ },
+ {
+ name: "Removes multiple spaces and excessive empty lines",
+ input: `This text has too many spaces.
And too many newlines.`,
+ expected: "This text has too many spaces.\n\nAnd too many newlines.",
+ },
+ {
+ name: "Nested lists with indentation",
+ input: "
One
Two
",
+ // Expect the sub-element to have 4 spaces of indentation
+ expected: "- One\n - Two",
+ },
+ {
+ name: "Image support",
+ input: ``,
+ // Correct Markdown syntax for images
+ expected: "",
+ },
+ {
+ name: "Image support without alt-text",
+ input: ``,
+ // If alt is missing, square brackets remain empty
+ expected: "",
+ },
+ {
+ name: "XSS Bypass on Links (Obfuscated HTML entities)",
+ // The Go HTML parser resolves entities, so this becomes "javascript:alert(1)"
+ input: `Click here`,
+ // Our isSafeHref (if updated with net/url) should neutralize it to "#"
+ expected: "[Click here](#)",
+ },
+ {
+ name: "Empty link or used as anchor",
+ input: ``,
+ // With no text or href, it shouldn't print anything (not even empty brackets)
+ expected: "",
+ },
+ {
+ name: "Link without href but with text (Textual anchor)",
+ input: `Back to top`,
+ // Should extract only plain text, without generating a broken Markdown link like [Back to top](#) or [Back to top]()
+ expected: "Back to top",
+ },
+ {
+ name: "Badly spaced bold and italics (Edge Case)",
+ input: ` Text `,
+ // In Markdown `** Text **` is often not formatted correctly. The ideal is `**Text**`
+ expected: "**Text**",
+ },
+ {
+ name: "Complex Test - Real Article",
+ input: `
+
+
+ `,
+ // Note: The indentation of the real HTML test will generate spaces that
+ // regex will clean up.
+ expected: "# Article Title\n\nThis is an **introductory text** with a [link](http://link.com).\n\n## Subtitle\n\n- Point one\n- Point two",
+ },
+ {
+ name: "Ordered list (OL)",
+ input: `
`,
+ expected: "Use the command `go test ./...` to run the tests.",
+ },
+ {
+ name: "Simple blockquote",
+ input: `
An important quote.
`,
+ expected: "> An important quote.",
+ },
+ {
+ name: "Multiline blockquote",
+ input: `
First line of the quote.
Second line of the quote.
`,
+ expected: "> First line of the quote.\n>\n> Second line of the quote.",
+ },
+ {
+ name: "Strikethrough text (del/s)",
+ input: `This text is deleted and this is crossed out.`,
+ expected: "This text is ~~deleted~~ and this is ~~crossed out~~.",
+ },
+ {
+ name: "Horizontal separator (HR)",
+ input: `
Above the line
Below the line
`,
+ expected: "Above the line\n\n---\n\nBelow the line",
+ },
+ {
+ name: "Bold nested in link",
+ input: `Linked bold text`,
+ expected: "[**Linked bold text**](https://example.com)",
+ },
+ {
+ name: "data-src Image (lazy loading)",
+ input: ``,
+ expected: "",
+ },
+ {
+ name: "Image with javascript: src blocked",
+ input: ``,
+ // src is not safe, so the image is not emitted
+ expected: "",
+ },
+ {
+ name: "Link with data: href blocked",
+ input: `Click`,
+ expected: "[Click](#)",
+ },
+ {
+ name: "Deeply nested divs",
+ input: `