検索コード

Makerfabs MaTouch ESP32-S3 2.8インチカメラ ESP32-S3でAI音声アシスタントを構築する(Azure + DeepSeek)

Makerfabs MaTouch ESP32-S3 2.8インチカメラ ESP32-S3でAI音声アシスタントを構築する(Azure + DeepSeek)

ボタンを押したまま質問すると、ボードが音声で答えます

1つの小さなボードに完全な音声アシスタントを搭載。SPEAKボタンを押したまま質問してください。ボードは2つのマイクで音声を録音し、その音声をMicrosoft Azureに送信してテキストに変換し、そのテキストをDeepSeekに送信して思考させ、回答をAzureに戻して音声に変換し、内蔵スピーカーから再生します。会話全体がチャットバブルの形で画面に表示されます。

MaTouch AI ESP32-S3ボードで動作する音声AIアシスタント MaTouch AI ESP32-S3ボードで動作する音声AIアシスタント

ボタンを押した瞬間から何が起こるか

ここでは、1つの質問の全プロセスを段階的に説明します。一度読む価値があります。画面やLEDに表示されるすべてのものが、これらの段階のいずれかに対応しているからです。

  1. SPEAKボタンを押したままにします。 タップして話すのではなく、押しながら話す方式です。指を押している間だけ録音が続き、最大6秒間です。ステータスLEDがに変わり、ボタンにLISTENINGと表示されます。

  2. 2つのマイクがあなたの声を録音します。 ボードはステレオペアから毎秒16,000回サンプリングし、2つのチャンネルを1つに平均化し、少しゲインを適用して、結果をPSRAMに保存します。話している間、ボタン上をプログレスバーがゆっくり進みます。2秒の音声は約64 KBです。

  3. ボタンを離します。 録音が停止します。ボードはオーディオの先頭に44バイトのWAVヘッダーを書き込みます。この小さなラベルが、生のサンプルをAzureが受け入れるファイルに変えるのです。

  4. 音声がAzure Speech-to-Textに送信されます。 4 KB単位で安全な接続を通じてアップロードされ、1行のテキストとして返ってきます。64 KBの音声が約25バイトの文字に変わります。LEDが琥珀色に変わります。

  5. あなたの質問が画面に表示されます 青いチャットバブルとして表示されるので、何を聞き取ったのかを正確に確認できます。これは便利です。聞き間違いがほとんどの奇妙な回答の原因だからです。

  6. テキストがDeepSeekに送信されます。 ボードはあなたの質問と、回答を2つの短い文に保つという常時指示を送信します。モデルは思考します。実際に思考する推論モデルです。そして回答を返します。

  7. 回答が画面に表示されます 灰色のバブルとして表示されます。音声で聞く前に読むことができます。

  8. 回答がAzureに戻されて音声化されます。 ボードは生の16 kHz PCMオーディオを要求します。これはアンプが正確に必要とする形式なので、このプロジェクトにはMP3デコーダーはどこにもありません。ボタンはGETTING VOICEと表示され、LEDは琥珀色のままです。まだ何も聞こえないからです。

  9. クリップ全体が、1サンプルも再生される前にPSRAMにダウンロードされます。 これは重要です。下の注記を参照してください。

  10. 再生。 オーディオがスピーカーに渡された瞬間、LEDがに変わり、ボタンにSPEAKINGと表示されます。回答が聞こえます。

音声が到着次第再生されるのではなく、先にダウンロードされる理由。 ネットワークからスピーカーへ直接ストリーミングすると、ノック音のように聞こえます。スピーカーのバッファは約0.1秒分しか保持できず、WiFi転送の一時停止がそれより長くなるとバッファが空になり、聞こえるノック音が発生します。回答全体を先にPSRAMにダウンロードすると、約1秒の追加待ち時間がかかりますが、すべてのギャップがなくなります。これが、ディスプレイがSPEAKINGと表示する前にGETTING VOICEと表示する理由でもあります。この2つは実際に異なる段階なのです。

各段階にかかる時間

実際のハードウェアで測定した、簡単な質問の場合:

段階

一般的な時間

録音

ボタンを押している間

音声認識 (Azure)

約1.8秒

思考 (DeepSeek)

簡単な質問で約1.8秒、実際の計算が必要な質問ではさらに長く

音声の取得 (Azure)

約7秒 - 最大の割合

合計、ボタン離してから最初の音まで

約11秒

各やり取りは独自のタイミングをシリアルモニターに出力するので、これらの数値を信頼するのではなく、自分で測定できます。より高速にしたい場合、最も効果的な変更はSYSTEM_PROMPTでより短い回答を要求することです。話すテキストが少なければ、合成してダウンロードする音声も少なくなります。

ボード自体は何も理解していません。優れた耳と優れた声を持つメッセンジャーであり、知性は秒単位でレンタルされています。

なぜこの3つのサービスなのか

  • Azure は音声の入力と出力を処理します。そのテキスト読み上げは生の16 kHz PCMを返すことができ、これはスピーカーチップが正確に必要とするものなので、このプロジェクトにはMP3デコーダーはどこにもありません。その音声認識はプレーンなWAVをプレーンなPOSTで受け取ります。

  • DeepSeek は会話の頭脳です。高速で、1回の返信あたり1セントの数分の1のコストです。

  • OpenAI はここでは使用されていません。プロジェクト05を参照してください。そこでは視覚処理を行います。

DeepSeekのモデル名が変更されました。 以前のdeepseek-chatdeepseek-reasonerという名前は2026年7月に廃止されました。オンラインのほとんどのチュートリアルはまだそれらを使用しており、エラーが返されます。現在の名前はdeepseek-v4-flashdeepseek-v4-proです。このプロジェクトではv4-flashを使用しています。

推論モデルの罠

DeepSeek v4-flashは回答する前に思考し、その思考はトークン制限にカウントされます。LLM_MAX_TOKENSを低く設定しすぎると、予算全体が推論に費やされ、回答は空で返され、ボードは何も表示しません。そのため、ここでは400に設定されています。難しい質問は時間もかかります。単純な事実は約2秒で回答しますが、実際の計算を必要とする質問ははるかに長くかかることがあります。

ステータスライトの読み方

意味

あなたの声を聞いています

琥珀色

クラウドが思考中、または音声を取得中

発話中 - 音声が始まった瞬間に緑に変わります

何かが失敗しました - シリアルモニタを確認してください

画面上のコントロール

チャットは電話の会話のようにスクロールし、古いメッセージは上に移動して消えていきます。CLEARで消去できます。実際のdBm表示付きのWiFi信号メーターが隅にあり、応答が遅いときにネットワークの問題かサービスの問題かを判断するのに役立ちます。

MaTouch AI ESP32-S3 2.8インチボードについて

このページのすべてのプロジェクトは、MakerfabsのMaTouch AI ESP32-S3 2.8インチTFT ST7789Vで動作します。これはオールインワンボードです:カラータッチスクリーン、3メガピクセルカメラ、2つのマイク、本格的なスピーカーアンプを備え、すべてESP32-S3と8MBのPSRAMで駆動されます。この組み合わせにより、追加の部品なしで単一のボード上でこれらのAIプロジェクトが可能になります。

8MBのPSRAMは、ここにある他のどの数値よりも重要です。これにより、ボードはカメラフレーム、数秒の録音音声、またはbase64エンコードされた写真を同時にメモリに保持できます。これらはどれもESP32の通常のRAMには収まりません。

メーカー資料:Makerfabs wikiページ

主な仕様

  • プロセッサ:ESP32-S3、デュアルコア240MHz、WiFi 2.4GHz + Bluetooth 5.0

  • メモリ:16MBフラッシュ、8MB PSRAM(ここのほぼすべてのプロジェクトで必要)

  • ディスプレイ:2.8インチIPS、320×240、ST7789Vドライバ、SPI

  • タッチ:GT911静電容量式、5本指同時追跡

  • カメラ:OV3660、3メガピクセル、最大2048×1536

  • マイク:INMP441 I2Sデジタルマイク2個(本格的なステレオペア)

  • スピーカー:MAX98357AクラスDアンプ、4Ωで3.2W

  • ストレージ:microSDカードスロット(SPIモード)

  • 電源:USB-C、JSTバッテリーコネクタ、TP4056充電器、電源スイッチ

  • その他搭載:WS2812B RGB LED、PCF8563Tバッテリーバックアップリアルタイムクロック、および公式仕様には記載されていないMAX17048バッテリー燃料ゲージ

2つのUSB-Cポートは同じではありません。ボードのスピーカーは、IO19とIO20の信号ピンをネイティブUSBポートと共有しています。これらのピンはESP32-S3のハードワイヤードUSBデータラインだからです。常にCH340K USB-Cポート(RESETボタンの横にある方)からアップロードおよび電源供給を行い、USB CDC On BootDisabledに設定してください。間違ったポートを使用すると、オーディオが誤動作したり、アップロードが失敗したりします。

Arduino IDE設定

これらの設定は重要です。このボードで報告される問題のほとんどは、これらのいずれかが間違っていることが原因であり、コアバージョンを変更するとリセットされるため、変更後は再度確認してください。

設定

ボード

ESP32S3 Dev Module

ESP32コアバージョン

2.0.17

PSRAM

OPI PSRAM

フラッシュサイズ

16MB(128Mb)

パーティションスキーム

16M Flash(3MB APP/9.9MB FATFS)

USB CDC On Boot

Disabled

アップロード速度

921600

アップロード前に全フラッシュを消去

Disabled

ポート

CH340K USB-Cポート

ESP32コア2.0.17を使用してください。3.xではありません。Espressifはコア3でデバイス上の顔検出モデルを削除したため、顔プロジェクトはそこでコンパイルできません。2.0.17に固定することで、このページのすべてのプロジェクトが1つの構成で動作します。ボードマネージャのバージョンドロップダウンで、いつでも自由に切り替えることができます。

GFX Library for Arduinoバージョン1.5.6を使用してください。1.6.xではありません。1.6リリースはESP32コア3用に構築されており、コア2.0.17では起動時にハングする可能性があります。アップロード後に画面が黒いままの場合は、これを最初に確認してください。

必要なライブラリ

Arduino IDEのツール → ライブラリを管理からこれらをインストールしてください。バージョン番号は重要です。記載されているものを使用してください。

ライブラリ

バージョン

作者

GFX Library for Arduino

1.5.6

moononournation

bb_captouch

1.3.1

Larry Bank

ArduinoJson

7.x

Benoit Blanchon

Adafruit NeoPixel

最新版ならどれでも

Adafruit

secrets.hの設定

WiFiの詳細とAPIキーは、ダウンロードにプレースホルダー値が入った状態で含まれているsecrets.hに入れてください。Arduino IDEでそのタブを開き、自分の情報に置き換えてください。

WiFiは2.4 GHzでなければなりません。 ESP32-S3は5 GHzのネットワークをまったく認識できません。ルーターが両方の帯域を1つの名前でまとめている場合(AsusではこれをSmart Connectと呼びます)、その機能をオフにするか、2.4 GHz帯に専用の名前を付けて、それをsecrets.hで使用してください。

APIキーの取得方法

このプロジェクトはクラウドAIサービスと通信するため、自分専用のキーが必要です。これまでにやったことがなくても心配しないでください。サービスにアカウントを識別させるパスワードと同じようなものです。一度だけ、数分で済みます。

キーはウェブサイトのサブスクリプションではありません。 例えばChatGPT Plusに支払ってもAPIキーは付与されません。両者は別々の製品で、請求も別々です。以下の説明にある開発者プラットフォームのアカウントが必要です。

Microsoft Azure Speech - 聞き取りと発話用

Azureは音声をテキストに変換し、回答を音声に戻します。無料枠はこのページのすべての機能に対して十分に余裕があります。

  1. portal.azure.comにアクセスし、Microsoftアカウント(無料のものでOK)でサインインします。

  2. Azureを初めて使う場合は、3つの選択肢があるAzureへようこそ画面が表示されます。Azure無料試用版で開始するを選択してください。Azureで何かを作成するにはサブスクリプションが必要です。(学生は代わりにAzure for Studentsを選択してください。同じ結果で、カードは不要です。)Microsoft Entra IDの管理は無視してください。これはまったく別のものです。

  3. リソースの作成をクリックし、Speechを検索して、Microsoft発行のSpeech serviceを選択します。

  4. フォームに記入します。任意のリソースグループ、任意の名前を付け、近くのリージョンを選択します。そのリージョンを表示されたとおり正確に書き留めてください。例:eastus

  5. 価格レベルF0(無料)を選択します。これで毎月約5時間の音声テキスト変換と50万文字のテキスト音声変換が可能です。

  6. 確認と作成をクリックし、次に作成をクリックします。約1分待ってから、リソースに移動をクリックします。

  7. 左側のメニューでキーとエンドポイントを開きます。KEY 1場所/リージョンをコピーします。

これらをsecrets.hAZURE_SPEECH_KEYAZURE_REGIONに入れてください。AZURE_STT_HOSTには、<region>.stt.speech.microsoft.comを使用します。つまり、リージョンがeastusの場合、eastus.stt.speech.microsoft.comとなります。

クレジットカードについて。 Azure無料試用版では本人確認のためにカードが必要です。請求は発生しません。30日間$200のクレジットが付与され、その後アカウントは従量課金制に移行しますが、F0 Speechレベルは無料のままで、毎月継続され、これらのプロジェクトのすべてがその範囲内に十分収まります。カードを一切提供したくない場合で学生であれば、Azure for Studentsオプションでカードなしでクレジットが得られます。

「Speech service」リソースでなければなりません。 Translator、Language、または一般的なCognitive Servicesリソースのキーは見た目が同じで完全に有効に見えますが、すべての音声リクエストでエラー401が返ります。テスト中にこれに遭遇し、1時間を費やしました。キーが正しく見えるのに音声が401で失敗する場合は、作成したリソースの種類を確認してください。

DeepSeek - 思考部分

DeepSeekは実際に質問に答える言語モデルです。低コストで、数ドルのクレジットで何千もの応答が可能です。

  1. platform.deepseek.comにアクセスしてアカウントを作成します。

  2. メニューでAPI keysを開き、Create new API keyをクリックします。

  3. すぐにコピーしてください。 キーは一度だけ表示され、二度と表示されません。失くした場合は、そのキーを削除して新しいものを作成してください。

  4. Top upで少額のクレジットを追加します。無料枠はありませんが、この使用量では最小のチャージで非常に長く持ちます。

キーをsecrets.hDEEPSEEK_KEYに入れてください。キーはsk-で始まります。

モデル名は2026年7月に変更されました。 以前のdeepseek-chatdeepseek-reasonerは廃止されたため、オンラインで見つかるほとんどのチュートリアルはエラー400で失敗します。deepseek-v4-flashを使用してください。これらのプロジェクトではすでにこれが設定されています。

実行にかかるコスト

非常に少額ですが、無料ではありません。プロジェクトを実行したままにする前に、おおよその支出を把握しておくべきです。

サービス

おおよそのコスト

Azure Speech

無料枠で毎月約5時間の聞き取りと0.5 M文字の発話をカバー

DeepSeek

回答1件あたり1セント未満 - 数ドルで何千もの応答

OpenAI vision

モデルにもよりますが、画像1枚あたり約1〜2セント

価格は変動するため、これらは見積もりではなく参考として扱ってください。これらのサービスにはすべて使用量ページがあり、支出を確認できます。また、すべてで支出制限を設定できます。これは初日に行う価値があります。

キーは秘密に保管してください。キーを入手した人は誰でもあなたの資金を使うことができます。ビデオ、スクリーンショット、フォーラムへの投稿、公開コードリポジトリにキーを入れないでください。キーが一度でも露出した場合は、プロバイダーのウェブサイトで削除し、新しいものを作成してください。数秒で完了し、これが唯一の本当の修正方法です。

トラブルシューティング

症状

原因と修正

画面が黒いまま

GFXライブラリのバージョンが間違っている(1.5.6を使用)か、ボード設定が間違っています。

PSRAM alloc failed またはカメラエラー 0xffffffff

Tools → PSRAMOPI PSRAM に設定されていません。

アップロードできない / COMポートがない

USB-Cポートが間違っているか、CH340ドライバがインストールされていません。

カメラが失敗し、回復しない

カメラのリセットラインがボードのRESETボタンに接続されているため、ソフトウェアで再起動できません。RESETを押してください。それでも失敗する場合は、カメラのリボンケーブルを挿し直してください。

コードをダウンロード

このプロジェクトの完全なArduinoスケッチは、pins.h やその他必要なものすべてとともに、無料でダウンロードできます。

04_Voice_Assistant をダウンロード

解凍し、Arduino IDEで .ino ファイルを開き、上記の設定を確認して、CH340K USB-Cポートからアップロードしてください。

884-Arduin code for MaTouch AI ESP32S3 2.8in AI Camera: AI voice assistant
言語: C++
/*
 * ===========================================================================
 *  04_Voice_Assistant  —  MaTouch AI ESP32-S3 2.8" TFT ST7789V
 * ===========================================================================
 *
 ----------
 *  ROBOJAX.COM  -  MaTouch AI ESP32-S3 2.8" project series
 *
 *    WATCH THE VIDEO
 *        https://youtu.be/6AL3g3tC_Hk
 *
 *    WRITTEN TUTORIALS - every project, with photos and full explanation
 *        Camera and touchscreen.... https://robojax.com/RTJ849
 *        Offline face recognition.. https://robojax.com/RTJ850
 *        AI voice assistant........ https://robojax.com/RTJ851
 *        AI vision................. https://robojax.com/RTJ852
 *
 *    GET THE BOARD - SAVE $5 with coupon code:  Robojax_Makerfab
 *        https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
 *        (enter the code at checkout)
 *
 *  All of this code is free. If it helped you, a subscribe on YouTube is
 *  the best way to support more of it.
 *  
 *  ---------------------------------------------------------------------------
 *
 *  A complete voice assistant on a $40 board:
 *
 *      hold SPEAK  ->  both INMP441 microphones record you
 *                  ->  Azure Speech turns the audio into text
 *                  ->  DeepSeek v4-flash thinks of an answer
 *                  ->  Azure Speech turns the answer into audio
 *                  ->  the MAX98357 speaker says it out loud
 *
 *  and the whole conversation is drawn as chat bubbles on the touchscreen.
 *
 *  WHY THIS COMBINATION OF SERVICES (each is used where it is best):
 *   - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
 *     transcription endpoint wants multipart/form-data - miserable on an MCU.
 *   - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
 *     which streams straight into the I2S speaker with NO MP3 decoder at all.
 *   - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
 *     model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
 *     Most tutorials online still use them and are broken. See secrets.h.
 *
 *  A detail the vendor examples get wrong: this board has TWO microphones on
 *  one I2S bus (left + right), but every Makerfabs demo records left-only and
 *  throws one away. This sketch records both and averages them.
 *
 *  ---------------------------------------------------------------------------
 *  *** WHICH USB PORT - THIS MATTERS ***
 *  The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
 *  through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
 *  If you use the wrong port the audio will be garbage or uploads will fail.
 *  ---------------------------------------------------------------------------
 *
 *  FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
 *
 *  BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
 *  or compile errors. These reset when you switch cores - recheck them!)
 *
 *      Board            : ESP32S3 Dev Module
 *      ESP32 core       : 2.0.17
 *      PSRAM            : OPI PSRAM        <-- required, audio buffer lives there
 *      Flash Size       : 16MB (128Mb)
 *      Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
 *      USB CDC On Boot  : Disabled         <-- required, see USB note above
 *      Upload Speed     : 921600
 *      Port             : the CH340K USB-C port (the one near RESET)
 *
 *  LIBRARIES
 *      GFX Library for Arduino   v1.5.6   (NOT 1.6.x - that pairs with core 3)
 *      bb_captouch               v1.3.1
 *      ArduinoJson               v7.x
 *      Adafruit NeoPixel         any recent
 *
 *  ---------------------------------------------------------------------------
 *  FUNCTIONS IN THIS SKETCH
 *      led(r,g,b) + LED_* macros   RGB status colours (blue/amber/green/red)
 *      getTouch(&x,&y)      read the touch panel, mapped to screen coordinates
 *      speakButtonHeld()    true while a finger is on the SPEAK button
 *      bubbleLines(t)       how many lines a message wraps to
 *      drawOneBubble(m,y)   draw a single chat bubble
 *      redrawChat()         rebuild the chat area from history, newest at bottom
 *      clearChat()          wipe the chat history (CLEAR button)
 *      chatBubble(t,user)   add a message to history and redraw
 *      drawWifi()           WiFi signal bars + dBm readout
 *      drawBar(label,col)   bottom bar: SPEAK button + CLEAR + WiFi meter
 *      micInit()            I2S input - BOTH INMP441 mics, stereo
 *      spkInit()            I2S output - MAX98357 speaker
 *      recordWhileHeld()    record while SPEAK held, downmix stereo->mono
 *      writeWavHeader(...)  prepend the 44-byte RIFF/WAVE header
 *      dumpWavToSD(...)     save the exact upload to SD (/stt_debug.wav)
 *      readHttpResponse()   read an HTTPS reply, de-chunking it properly
 *      azureSTT(...)        chunked upload of the WAV -> recognised text
 *      deepseekChat(...)    question -> deepseek-v4-flash -> answer text
 *      azureTTSSpeak(text)  answer -> Azure voice -> PSRAM -> speaker
 *      setup() / loop()     boot + WiFi / one conversation turn per press
 *
 *  Robojax.com
 * ===========================================================================
 */

#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"

/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE     16000
#define RECORD_MAX_S    6                       // hard cap on one question
#define WAV_HEADER_LEN  44
#define REC_BUF_BYTES   (SAMPLE_RATE * RECORD_MAX_S * 2)   // 16-bit mono

/* Software gain applied to the recording. The INMP441 capture is quiet at
 * 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
 * speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
 * lower it. */
#define MIC_GAIN        6

/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
 * 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
 *     back silent or garbled in stereo mode. */
#define USE_BOTH_MICS   1

/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
 * uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS  25

/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
 * shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
    TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);

BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);

/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr;      // WAV_HEADER_LEN + up to REC_BUF_BYTES

/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;

/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool  ok_sd = false;
float g_mic_peak_pct = 0;          // last recording's raw peak, % of full scale
char  g_stt_err[64] = "";          // last STT failure cause, verbatim
char  g_llm_err[64] = "";          // last DeepSeek failure cause, verbatim

/* --- layout --------------------------------------------------------------- */
#define CHAT_H     200              // chat area: y 0..199
#define BAR_Y      202              // button bar below it
#define BTN_SPEAK_X   4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W  58
#define WIFI_X      236             // signal indicator, right end of the bar
#define BTN_H        36

/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
 * This is what prevents new text printing over old - the whole area is
 * rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int     chat_count = 0;

/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
  rgb.setPixelColor(0, rgb.Color(r, g, b));
  rgb.show();
}
#define LED_IDLE()    led(0, 0, 0)
#define LED_LISTEN()  led(0, 60, 255)     // blue   - recording
#define LED_THINK()   led(255, 120, 0)    // amber  - waiting on the cloud
#define LED_SPEAK()   led(0, 255, 40)     // green  - talking
#define LED_ERROR()   led(255, 0, 0)      // red


/* ===========================================================================
 *  Touch
 * =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
  TOUCHINFO ti;
  if (!bbct.getSamples(&ti)) return false;
  if (ti.count < 1) return false;
  *x = ti.y[0];
  *y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
  return true;
}

bool speakButtonHeld() {
  uint16_t x, y;
  if (!getTouch(&x, &y)) return false;
  return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}


/* ===========================================================================
 *  Chat UI  —  word-wrapped bubbles, user right/blue, assistant left/grey
 * =========================================================================== */
#define CHAT_CHARS 42                          // chars per line at textsize 1

static int bubbleLines(const char *t) {
  int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
  return l < 1 ? 1 : l;
}

void drawOneBubble(const ChatMsg &m, int y) {
  int len = strlen(m.text);
  int lines = bubbleLines(m.text);
  int h = lines * 10 + 8;
  uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
  int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
  if (w < 30) w = 30;
  int x = m.from_user ? (316 - w) : 4;

  gfx->fillRoundRect(x, y, w, h, 5, bg);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  for (int i = 0; i < lines; i++) {
    char line[CHAT_CHARS + 1] = {0};
    strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
    gfx->setCursor(x + 5, y + 5 + i * 10);
    gfx->print(line);
  }
}

/* Rebuild the whole chat area from history: newest message anchored at the
 * bottom, older ones stacked upward until the area is full. */
void redrawChat() {
  gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
  int shown = min(chat_count, CHAT_HISTORY);
  int y = CHAT_H - 2;
  for (int i = 0; i < shown; i++) {
    ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
    int h = bubbleLines(m.text) * 10 + 8;
    y -= h;
    if (y < 0) break;                          // area full - older ones drop off
    drawOneBubble(m, y);
    y -= 4;
  }
}

void clearChat() {
  chat_count = 0;
  redrawChat();
}

void chatBubble(const char *text, bool from_user) {
  ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
  strncpy(m.text, text, sizeof(m.text) - 1);
  m.text[sizeof(m.text) - 1] = 0;
  m.from_user = from_user;
  chat_count++;
  redrawChat();
}

/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
  gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
  bool up = (WiFi.status() == WL_CONNECTED);
  long rssi = up ? WiFi.RSSI() : -100;
  // -55 dBm or better = full bars; each 10 dB drops one
  int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;

  for (int b = 0; b < 4; b++) {
    int bh = 6 + b * 6;                       // heights 6,12,18,24
    uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
    gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
  }
  gfx->setTextSize(1);
  gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
  if (up) {
    gfx->setTextColor(CYAN);
    gfx->printf("%lddBm", rssi);
  } else {
    gfx->setTextColor(RED);
    gfx->print("DOWN");
  }
  gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
  gfx->setTextColor(gfx->color565(120, 120, 120));
  gfx->print("WiFi");
}

void drawBar(const char *label, uint16_t colour) {
  gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
  gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
  gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
  gfx->setTextSize(2);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
  gfx->print(label);

  // CLEAR wipes the chat history
  gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
  gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
  gfx->print("CLEAR");

  drawWifi();
}


/* ===========================================================================
 *  I2S  —  microphones on port 0, speaker on port 1. Separate hardware
 *  ports, so recording and playback can never fight over a bus.
 * =========================================================================== */
void micInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    /* BOTH channels - this is the two-microphone fix. The vendor examples
     * use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
     * falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
    .channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = false,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_MIC_SCK,
    .ws_io_num = I2S_MIC_WS,
    .data_out_num = I2S_PIN_NO_CHANGE,
    .data_in_num = I2S_MIC_SD
  };
  i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_MIC_PORT, &pins);
}

void spkInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,   // mono - the amp downmixes anyway
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = true,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_SPK_BCLK,
    .ws_io_num = I2S_SPK_LRC,
    .data_out_num = I2S_SPK_DOUT,
    .data_in_num = I2S_PIN_NO_CHANGE
  };
  i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_SPK_PORT, &pins);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
}


/* ===========================================================================
 *  Recording  —  runs while the SPEAK button is held (up to RECORD_MAX_S).
 *  Reads stereo pairs, averages L+R into one mono stream, applies a little
 *  software gain, and fills wav_buf after the 44-byte header slot.
 *  Returns the number of audio bytes recorded.
 * =========================================================================== */
size_t recordWhileHeld() {
  int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
  size_t   mono_samples = 0;
  const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;

  int16_t chunk[512];
  uint32_t last_touch_ok = millis();
  int32_t  peak = 0;                        // loudest raw sample - mic health check

  i2s_zero_dma_buffer(I2S_MIC_PORT);

  while (mono_samples < max_samples) {
    size_t got = 0;
    i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);

#if USE_BOTH_MICS
    size_t n = got / 4;                     // 4 bytes = one L+R pair
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
    size_t n = got / 2;                     // 2 bytes = one mono sample
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = chunk[i];
#endif
      if (abs(raw) > peak) peak = abs(raw);
      int32_t mixed = raw * MIC_GAIN;
      if (mixed >  32767) mixed =  32767;
      if (mixed < -32768) mixed = -32768;
      mono[mono_samples++] = (int16_t)mixed;
    }

    /* The GT911 is polled between I2S reads. A 250 ms grace period stops a
     * momentary missed touch sample from cutting the recording short. */
    if (speakButtonHeld()) last_touch_ok = millis();
    else if (millis() - last_touch_ok > 250) break;

    // live progress on the button
    static uint32_t last_draw = 0;
    if (millis() - last_draw > 200) {
      last_draw = millis();
      gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
                    (int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
    }
  }

  /* Mic health line: peak as % of full scale BEFORE gain.
   *   0%       = the mic is not being read at all (config/pin problem)
   *   under 3% = too quiet - speak closer or raise MIC_GAIN
   *   3-40%    = healthy speech level
   */
  g_mic_peak_pct = peak * 100.0 / 32768.0;
  Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
                g_mic_peak_pct, MIC_GAIN,
                peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");

  return mono_samples * 2;
}

/* Dump the exact WAV we are about to POST onto the SD card, so it can be
 * played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
  if (!ok_sd) return;
  SD.remove("/stt_debug.wav");
  File f = SD.open("/stt_debug.wav", FILE_WRITE);
  if (!f) return;
  f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
  f.close();
  Serial.println("debug copy saved to SD as /stt_debug.wav");
}

/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
  uint32_t file_len = data_bytes + 36;
  uint32_t byte_rate = SAMPLE_RATE * 2;
  memcpy(h, "RIFF", 4);           memcpy(h + 4, &file_len, 4);
  memcpy(h + 8, "WAVEfmt ", 8);
  uint32_t fmt_len = 16;          memcpy(h + 16, &fmt_len, 4);
  uint16_t fmt = 1, ch = 1;       memcpy(h + 20, &fmt, 2);   memcpy(h + 22, &ch, 2);
  uint32_t rate = SAMPLE_RATE;    memcpy(h + 24, &rate, 4);  memcpy(h + 28, &byte_rate, 4);
  uint16_t align = 2, bits = 16;  memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
  memcpy(h + 36, "data", 4);      memcpy(h + 40, &data_bytes, 4);
}


/* ===========================================================================
 *  HTTP response reader  —  shared by all three cloud calls.
 *  Returns the status code and fills body_out. Handles chunked transfer
 *  encoding PROPERLY: the chunk-size markers must be stripped, or they end
 *  up embedded inside the JSON body and the parse fails on long replies.
 * =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
 * on timeout / closed-and-empty. Every read below goes through this, because
 * a reasoning model can think for many seconds between the response headers
 * and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
  uint32_t t0 = millis();
  while (!c.available()) {
    if (!c.connected()) return false;
    if (millis() - t0 > ms) return false;
    delay(10);
  }
  return true;
}

static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
  body_out = "";

  if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);

  bool chunked = false;
  while (waitData(client, idle_ms)) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;         // blank line = end of headers
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
  }

  if (chunked) {
    int blanks = 0;
    while (true) {
      /* The chunk-size line may not arrive for a long time while the model
       * reasons. Waiting here - instead of letting read() time out - is the
       * whole fix: a timed-out read looks exactly like "0" (final chunk),
       * which silently truncated the body to nothing. */
      if (!waitData(client, idle_ms)) {
        Serial.println("HTTP: timed out waiting for the next chunk");
        break;
      }
      String szline = client.readStringUntil('\n');
      szline.trim();
      if (szline.length() == 0) {                    // stray blank line
        if (++blanks > 4) break;
        continue;
      }
      blanks = 0;

      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;                            // genuine final chunk

      long got = 0;
      while (got < sz) {
        if (!waitData(client, idle_ms)) break;
        while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
      }
      if (waitData(client, 3000)) client.readStringUntil('\n');   // CRLF after chunk
      if (got < sz) { Serial.println("HTTP: short chunk"); break; }
    }
  } else {
    while (waitData(client, idle_ms))
      while (client.available()) body_out += (char)client.read();
  }
  return code;
}


/* ===========================================================================
 *  CLOUD CALL 1  —  Azure speech-to-text
 *  One POST, one header, plain WAV body. This is why Azure does the ears.
 * =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
  /* HTTPClient's one-shot POST fails on bodies this large (it attempts one
   * giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
   * function speaks HTTP directly and streams the WAV up in 4 KB chunks -
   * reliable, and if it ever stalls we know the exact byte it stopped at. */
  WiFiClientSecure client;
  client.setInsecure();                    // no cert bundle on-device; see notes
  client.setTimeout(15);                   // seconds, for reads

  writeWavHeader(wav_buf, audio_bytes);
  dumpWavToSD(audio_bytes);                // PC-playable copy of what we send
  size_t total = WAV_HEADER_LEN + audio_bytes;

  if (!client.connect(AZURE_STT_HOST, 443)) {
    snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
    Serial.println("STT: TLS connect failed");
    return false;
  }

  /* Two valid host forms use DIFFERENT URL paths - detect which one is in
   * secrets.h:  <resource>.cognitiveservices.azure.com -> /stt/speech/...
   *             <region>.stt.speech.microsoft.com      -> /speech/...     */
  bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
  String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
               "/speech/recognition/conversation/cognitiveservices/v1"
               "?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
               "Host: " AZURE_STT_HOST "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
               "Accept: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(total) + "\r\n\r\n";
  client.print(req);

  /* body, 4 KB at a time */
  size_t sent = 0;
  while (sent < total) {
    size_t n = min((size_t)4096, total - sent);
    size_t w = client.write(wav_buf + sent, n);
    if (w == 0) {
      delay(50);                           // brief stall - retry once
      w = client.write(wav_buf + sent, n);
      if (w == 0) {
        snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
                 (unsigned)(sent / 1024));
        Serial.printf("STT: upload stalled at %u/%u bytes\n",
                      (unsigned)sent, (unsigned)total);
        client.stop();
        return false;
      }
    }
    sent += w;
    yield();
  }
  Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);

  /* read the reply with proper de-chunking */
  String resp;
  int code = readHttpResponse(client, resp, 10000);
  client.stop();

  if (code != 200) {
    snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
    Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
    Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
    return false;
  }

  const char *status = doc["RecognitionStatus"];
  if (!status || strcmp(status, "Success") != 0) {
    /* The status names the exact failure:
     *   InitialSilenceTimeout = Azure heard silence (mic level too low)
     *   NoMatch               = heard sound but no recognisable words
     *   BabbleTimeout         = heard only noise                       */
    snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
    Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
    return false;
  }
  text_out = doc["DisplayText"].as<String>();
  if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
  return text_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 2  —  DeepSeek chat completion
 *  OpenAI-compatible format. Model name is deepseek-v4-flash - the old
 *  deepseek-chat name is dead, see secrets.h.
 * =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
  /* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
   * bodies (long answers + the model's hidden reasoning), which shows up as
   * "bad JSON reply". Reading until the server closes the connection is
   * reliable regardless of reply length. */
  JsonDocument req;
  req["model"] = DEEPSEEK_MODEL;
  req["max_tokens"] = LLM_MAX_TOKENS;
  JsonArray msgs = req["messages"].to<JsonArray>();
  JsonObject sys = msgs.add<JsonObject>();
  sys["role"] = "system";  sys["content"] = SYSTEM_PROMPT;
  JsonObject usr = msgs.add<JsonObject>();
  usr["role"] = "user";    usr["content"] = question;

  String body;
  serializeJson(req, body);

  WiFiClientSecure client;
  client.setInsecure();
  /* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
   * ~2 s, but anything that needs actual working-out (an Ohm's law problem,
   * say) can think for 10-30 s before sending a single byte. */
  client.setTimeout(60);

  if (!client.connect(DEEPSEEK_HOST, 443)) {
    snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
    Serial.println("LLM: TLS connect failed");
    return false;
  }

  client.print(String("POST /chat/completions HTTP/1.1\r\n"
               "Host: " DEEPSEEK_HOST "\r\n"
               "Authorization: Bearer " DEEPSEEK_KEY "\r\n"
               "Content-Type: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: ") + String(body.length()) + "\r\n\r\n");
  client.print(body);

  /* read the reply with proper de-chunking; generous window - long
   * questions make the model think for a while before it responds */
  String resp;
  int code = readHttpResponse(client, resp, 60000);
  client.stop();

  if (code != 200) {
    snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
    Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
    Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
    return false;
  }

  const char *content = doc["choices"][0]["message"]["content"];
  const char *finish  = doc["choices"][0]["finish_reason"];
  if (!content || !content[0]) {
    /* v4-flash is a reasoning model: if finish_reason is "length", the whole
     * token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
    if (finish && strcmp(finish, "length") == 0)
      snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
    else
      snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
    Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
    return false;
  }
  answer_out = String(content);
  answer_out.trim();
  if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
  return answer_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 3  —  Azure text-to-speech, streamed straight to the speaker
 *  We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
 *  the I2S peripheral eats. Skip the 44-byte header, forward the rest.
 *  No MP3 decoder, no audio library, no buffering the whole reply.
 * =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
  size_t got = 0;
  uint32_t t0 = millis();
  while (got < n && millis() - t0 < 10000) {
    int r = c.read(dst + got, n - got);
    if (r > 0) { got += r; t0 = millis(); }
    else if (!c.connected() && !c.available()) break;
    else delay(2);
  }
  return got;
}

bool azureTTSSpeak(const String &text) {
  // Escape the XML special characters for the SSML body
  String safe = text;
  safe.replace("&", "&amp;");
  safe.replace("<", "&lt;");
  safe.replace(">", "&gt;");

  String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
                "<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";

  /* Manual HTTP like the other two cloud calls - and for a hard reason:
   * Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
   * raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
   * into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
   * Here we parse the framing properly and keep only clean audio bytes. */
  WiFiClientSecure client;
  client.setInsecure();
  client.setTimeout(20);

  const char *host = AZURE_REGION ".tts.speech.microsoft.com";
  if (!client.connect(host, 443)) {
    Serial.println("TTS: TLS connect failed");
    return false;
  }

  client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
               "Host: ") + host + "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: application/ssml+xml\r\n"
               "X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
               "User-Agent: MaTouchRobojax\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(ssml.length()) + "\r\n\r\n");
  client.print(ssml);

  /* status + headers; note whether the body is chunked */
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
  bool chunked = false;
  long content_len = -1;
  while (client.connected() || client.available()) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
    if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
  }
  if (code != 200) {
    Serial.printf("TTS HTTP %d\n", code);
    client.stop();
    return false;
  }

  const size_t AUDIO_CAP = 1200 * 1024;    // ~37 s of speech
  uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
  if (!audio) { client.stop(); return false; }
  size_t alen = 0;

  if (chunked) {
    /* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
    while (true) {
      String szline = client.readStringUntil('\n');
      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;
      if (alen + sz > AUDIO_CAP) break;
      size_t got = readExact(client, audio + alen, sz);
      alen += got;
      client.readStringUntil('\n');        // trailing CRLF after each chunk
      if (got < (size_t)sz) break;
    }
  } else if (content_len > 0) {
    alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
  } else {
    /* no framing info: read until the server closes */
    uint32_t idle = millis();
    while ((client.connected() || client.available()) && millis() - idle < 5000) {
      int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
      if (r > 0) { alen += r; idle = millis(); }
      else delay(5);
    }
  }
  client.stop();
  Serial.printf("TTS: %u KB clean audio (%s), playing\n",
                (unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");

  bool ok = (alen > WAV_HEADER_LEN);
  if (ok) {
    /* NOW the audio actually starts - this is the honest moment to go green */
    LED_SPEAK();
    drawBar("SPEAKING...", gfx->color565(0, 130, 40));

    /* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
    static const uint8_t lead_in[640] = {0};             // 20 ms of silence
    size_t w = 0;
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
  }
  free(audio);

  // let the DMA buffers drain so the last word is not cut off
  delay(150);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
  return ok;
}


/* ===========================================================================
 *  SETUP
 * =========================================================================== */
void setup() {
  Serial.begin(115200);
  delay(400);
  Serial.println("\n=== 04 Voice Assistant  |  Robojax.com ===");
  Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");

  pinMode(TFT_BLK, OUTPUT);
  digitalWrite(TFT_BLK, LOW);
  pinMode(SD_CS, OUTPUT);
  digitalWrite(SD_CS, HIGH);

  // one shared SPI bus for TFT + SD (started before either device)
  SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);

  gfx->begin();
  gfx->fillScreen(BLACK);
  digitalWrite(TFT_BLK, HIGH);

  // SD is optional here - it only stores the /stt_debug.wav diagnostic copy
  ok_sd = SD.begin(SD_CS, SPI, 20000000);
  digitalWrite(SD_CS, HIGH);
  Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
                       : "SD not found - debug WAV dump disabled (not fatal)");

  bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
  delay(50);

  rgb.begin();
  rgb.setBrightness(LED_BRIGHTNESS);
  LED_IDLE();

  /* One recording buffer for the whole session, in PSRAM. This is the 8 MB
   * that makes the board worth buying. */
  wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
  if (!wav_buf) {
    gfx->setTextColor(RED);
    gfx->setTextSize(2);
    gfx->setCursor(10, 100);
    gfx->print("PSRAM alloc failed!");
    gfx->setTextSize(1);
    gfx->setCursor(10, 130);
    gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
    while (1) delay(1000);
  }

  micInit();
  spkInit();

  gfx->setTextSize(1);
  gfx->setTextColor(YELLOW);
  gfx->setCursor(4, 4);
  gfx->printf("Connecting to %s ...", WIFI_SSID);
  Serial.printf("Connecting to %s ", WIFI_SSID);

  WiFi.mode(WIFI_STA);
  WiFi.begin(WIFI_SSID, WIFI_PASS);
  uint32_t t0 = millis();
  while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
    delay(300);
    Serial.print(".");
  }
  Serial.println();

  clearChat();
  if (WiFi.status() == WL_CONNECTED) {
    Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
    chatBubble("Hold SPEAK and ask me anything.", false);
  } else {
    chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
    LED_ERROR();
  }
  drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}


/* ===========================================================================
 *  LOOP  —  one full conversation turn per button press
 * =========================================================================== */
void loop() {
  /* CLEAR button: edge-detected so one tap wipes once. Reading the panel
   * twice per loop (here and in speakButtonHeld) is fine - the GT911 just
   * reports its current state. */
  static bool tap_latch = false;
  static uint8_t tap_release = 0;
  if (state == ST_IDLE) {
    uint16_t cx, cy;
    if (getTouch(&cx, &cy)) {
      tap_release = 0;
      if (!tap_latch) {
        tap_latch = true;
        if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
          clearChat();
          chatBubble("Hold SPEAK and ask me anything.", false);
        }
      }
    } else if (tap_latch && ++tap_release >= 4) {
      tap_latch = false;
      tap_release = 0;
    }

    /* live WiFi signal indicator, refreshed every 2 s while idle */
    static uint32_t last_wifi = 0;
    if (millis() - last_wifi > 2000) {
      last_wifi = millis();
      drawWifi();
    }
  }

  if (state == ST_IDLE && speakButtonHeld()) {

    /* ---- record ---- */
    state = ST_RECORDING;
    LED_LISTEN();
    drawBar("LISTENING...", gfx->color565(0, 60, 200));
    uint32_t t_rec = millis();
    size_t audio_bytes = recordWhileHeld();
    t_rec = millis() - t_rec;
    Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);

    if (audio_bytes < SAMPLE_RATE / 2) {   // under a quarter second - a tap
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }

    /* ---- speech to text ---- */
    state = ST_STT;
    LED_THINK();
    drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
    uint32_t t_stt = millis();
    String question;
    if (!azureSTT(audio_bytes, question)) {
      /* Show the REAL cause on screen - no serial monitor needed. */
      char diag[96];
      snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
               g_stt_err, g_mic_peak_pct,
               ok_sd ? " | saved /stt_debug.wav" : "");
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_stt = millis() - t_stt;
    chatBubble(question.c_str(), true);
    Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());

    /* ---- think ---- */
    state = ST_LLM;
    drawBar("THINKING...", gfx->color565(150, 90, 0));
    uint32_t t_llm = millis();
    String answer;
    if (!deepseekChat(question, answer)) {
      char diag[96];
      snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_llm = millis() - t_llm;
    chatBubble(answer.c_str(), false);
    Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());

    /* ---- speak ----
     * Still amber here: the voice has to be synthesised and downloaded first
     * (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
     * at the exact moment audio starts coming out of the speaker. */
    state = ST_TTS;
    LED_THINK();
    drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
    uint32_t t_tts = millis();
    bool spoke = azureTTSSpeak(answer);
    t_tts = millis() - t_tts;

    /* Timing summary on serial - this feeds the "honest numbers" segment. */
    Serial.printf("TIMINGS  rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
                  audio_bytes / 32000.0, (unsigned long)t_stt,
                  (unsigned long)t_llm, (unsigned long)t_tts,
                  spoke ? "" : "  (TTS FAILED)");

    drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
    LED_IDLE();
    state = ST_IDLE;
  }

  delay(20);
}

リソースと参考文献

ファイル📁

必要なファイル (.h)

  • secrets.h
    Makerfabs MaTouch AI ESP32S3 2.8インチTFTカメラモジュール用ファイル
    secrets.h 0.01 MB

その他のファイル

  • pins.h
    MaTouch AI ESP32S3 2.8インチカメラLCDタッチスクリーン用のピンファイル。
    pins.h 0.01 MB

回路図

  • MaTouch_AI 2.8インチ MaTouch AI ESP32S3 2.8インチTFT ST7789V回路図
    最新のMaTouch AIボードは、I2S音声入力/I2Sスピーカー/300万画素カメラOV3660/320*240解像度ディスプレイを統合し、ESP32S3の強力なプロセッサとWi-Fi機能により、ESP32でのAI開発に適したツール/プラットフォームとなっています。
    MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF 0.15 MB