This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
The latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
Makerfabs MaTouch ESP32-S3 2.8吋相機 喺ESP32-S3上建立AI語音助手(Azure + DeepSeek)
按住一個掣,問一條問題,塊板會大聲答你
一塊細板上面嘅完整語音助手。按住 SPEAK 掣,問一條問題。塊板用兩個咪高峰錄低你,將音訊送去 Microsoft Azure 轉做文字,再將文字送去 DeepSeek 諗嘢,跟住將答案送回 Azure 轉做語音,再用自己嘅喇叭播返出嚟。成段對話會以聊天氣泡形式顯示喺螢幕上。
喺 MaTouch AI ESP32-S3 板上運行嘅對話式 AI 語音助手
由你撳掣嗰一刻開始發生咩事
以下係一條問題嘅完整旅程,逐步講解。值得睇一次,因為你喺螢幕同 LED 上見到嘅一切,都對應住其中一個階段。
你按住 SPEAK 掣。 係按住嚟講,唔係撳一下:錄音會隨住你手指按住嘅時間而進行,最長六秒。狀態 LED 會變藍色,掣上面會顯示 LISTENING。
兩個咪高峰錄低你。 塊板每秒由立體聲對採樣 16,000 次,將兩個聲道平均成一個,加少少增益,然後將結果存入 PSRAM。你講嘢嘅時候,掣上面會有一條進度條慢慢行。兩秒嘅語音大約係 64 KB。
你放開個掣。 錄音停止。塊板會喺音訊前面寫入一個 44 字節嘅 WAV 檔頭——呢個細細嘅標籤就係將原始採樣變成 Azure 會接受嘅檔案嘅關鍵。
音訊會送去 Azure 語音轉文字。 佢會以 4 KB 一塊嘅方式,透過安全連線上傳,然後以一行文字返嚟。你嘅 64 KB 聲音變成大約 25 字節嘅文字。LED 會變琥珀色。
你嘅問題會顯示喺螢幕上,以藍色聊天氣泡形式出現,等你可以清楚睇到佢聽到咩——呢個好有用,因為聽錯嘅字可以解釋好多奇怪嘅答案。
文字會送去 DeepSeek。 塊板會將你嘅問題連同一條固定指示一齊送出,要求回覆限於兩句短句。個模型會諗——真係會諗,佢係一個推理模型——然後返一個答案。
答案會顯示喺螢幕上,以灰色氣泡形式出現。你可以喺聽到之前先讀到。
答案會送回 Azure 轉做語音。 塊板要求原始 16 kHz PCM 音訊,呢個正正係佢個放大器想要嘅格式,所以成個項目入面都冇 MP3 解碼器。掣上面而家會顯示 GETTING VOICE,LED 保持琥珀色,因為暫時仲乜都聽唔到。
成段音訊會喺播任何一個採樣之前,先下載入 PSRAM。 呢點好重要——睇下面嘅備註。
播放。 當音訊交俾喇叭嗰一刻,LED 會變綠色,掣上面會顯示 SPEAKING。你就會聽到答案。
點解要預先下載音訊,而唔係邊到邊播。 直接由網絡串流去喇叭會聽到「咯咯」聲。喇叭嘅緩衝區只可以裝到大約十分之一秒,而 WiFi 傳輸每次停頓超過咁耐,就會清空緩衝區,產生可聽到嘅「咯咯」聲。預先將成段回覆下載入 PSRAM,大約會多等一秒,但可以消除所有間斷。呢個亦都係點解顯示屏會先顯示 GETTING VOICE,之後先顯示 SPEAKING——兩者確實係唔同嘅階段。
每個階段要幾耐
喺真實硬件上量度,以一條簡單問題為例:
階段 | 一般時間 |
|---|---|
錄音 | 視乎你按住個掣幾耐 |
語音轉文字(Azure) | 大約 1.8 秒 |
思考(DeepSeek) | 簡單問題大約 1.8 秒,需要真正運算嘅問題會耐好多 |
攞語音(Azure) | 大約 7 秒——最大嘅單一環節 |
總計,由放手到第一下聲 | 大約 11 秒 |
每次對話都會將自己嘅時間印喺序列監視器度,所以你可以自己量度,唔使信呢啲數字。如果你想快啲,最有效嘅改動係喺 SYSTEM_PROMPT 入面要求短啲嘅回覆——要講嘅文字少啲,即係要合成同下載嘅音訊少啲。
塊板本身永遠唔會理解任何嘢。佢只係一個有好耳朵同好聲音嘅傳訊員——啲智慧係按秒租返嚟嘅。
點解揀呢三個服務
Azure 負責語音輸入同輸出。佢嘅文字轉語音可以返原始 16 kHz PCM,正正係喇叭晶片想要嘅格式,所以成個項目入面都冇 MP3 解碼器。佢嘅語音轉文字接受一個普通 WAV 檔案,用一個普通 POST 請求就得。
DeepSeek 係對話嘅大腦。佢好快,每次回覆成本唔夠一分錢。
OpenAI 呢度冇用到——請睇項目 05,嗰度會用佢做視覺工作。
DeepSeek 模型名稱改咗。 舊嘅 deepseek-chat 同 deepseek-reasoner 名稱已經喺 2026 年 7 月退役。網上大部分教學仲用緊呢啲名,會出錯。而家嘅名稱係 deepseek-v4-flash 同 deepseek-v4-pro。呢個項目用 v4-flash。
推理模型嘅陷阱
DeepSeek v4-flash 喺回答之前會先諗嘢,而嗰啲諗嘢會計入你嘅 token 限額。如果將 LLM_MAX_TOKENS 設得太低,成個預算就會用晒喺推理度,答案會空手而回,塊板就乜都唔講。呢度設做 400 就係因為咁。難嘅問題亦都會耐啲——簡單事實大概兩秒就答到,需要真正運算嘅問題就可能要耐好多。
睇狀態燈
顏色 | 意思 |
|---|---|
藍色 | 聽緊你講嘢 |
琥珀色 | 雲端度諗緊嘢,或者攞緊把聲 |
綠色 | 講緊嘢——聲音一開就即刻變綠 |
紅色 | 有嘢失敗咗——檢查序列監視器 |
螢幕上嘅控制
對話好似電話傾偈咁捲動,最舊嘅訊息會移上去然後消失。CLEAR 會清走晒。角落有個 WiFi 訊號錶,顯示實際 dBm 讀數,當你懷疑慢回覆係網絡定服務問題嗰陣就好有用。
關於 MaTouch AI ESP32-S3 2.8" 板
呢頁嘅每個項目都係喺 Makerfabs 嘅 MaTouch AI ESP32-S3 2.8" TFT ST7789V 上面運行。佢係一塊全合一板:彩色觸控螢幕、3 百萬像素相機、兩個咪高峰同一個真正嘅喇叭放大器,全部由一個有 8 MB PSRAM 嘅 ESP32-S3 驅動。就係呢個組合令到呢啲 AI 項目可以喺一塊板上面做到,唔使駁其他嘢。
8 MB PSRAM 比呢度任何其他數字都重要。就係佢令到塊板可以同時喺記憶體度裝住一格相機畫面、幾秒錄音,或者一張 base64 編碼嘅相——呢啲嘢冇一樣放得入 ESP32 嘅普通 RAM。
廠商文件:Makerfabs wiki 頁面。
主要規格
處理器:ESP32-S3,雙核心 240 MHz,WiFi 2.4 GHz + 藍牙 5.0
記憶體:16 MB flash,8 MB PSRAM(呢度幾乎每個項目都需要)
顯示:2.8" IPS,320×240,ST7789V 驅動器,SPI
觸控:GT911 電容式,同時追蹤 5 隻手指
相機:OV3660,3 百萬像素,最高 2048×1536
咪高峰:兩個 INMP441 I2S 數碼咪(真正嘅立體聲組合)
喇叭:MAX98357A D 類放大器,4 Ω 輸出 3.2 W
儲存:microSD 卡槽(SPI 模式)
電源:USB-C、JST 電池連接器、TP4056 充電器、電源開關
板上仲有:WS2812B RGB LED、PCF8563T 電池備用實時時鐘,同一個 MAX17048 電池電量計,後者冇列喺官方規格入面
兩個 USB-C 埠係唔同嘅。塊板嘅喇叭同原生 USB 埠共用訊號腳(IO19 同 IO20),因為嗰啲腳係 ESP32-S3 硬件接死嘅 USB 數據線。永遠用 CH340K USB-C 埠(RESET 按鈕旁邊嗰個)嚟上傳同供電,並且將 USB CDC On Boot 設做 Disabled。用錯埠嘅話音訊會出問題或者上傳會失敗。
Arduino IDE 設定
呢啲設定好重要。大部分人報告呢塊板嘅問題都係其中一樣設錯,而且佢哋喺你轉核心版本嗰陣會重置,所以每次改動之後都要再檢查一次。
設定 | 數值 |
|---|---|
板 | ESP32S3 Dev Module |
ESP32 核心版本 | 2.0.17 |
PSRAM | OPI PSRAM |
Flash 大小 | 16MB (128Mb) |
分割區方案 | 16M Flash (3MB APP/9.9MB FATFS) |
USB CDC On Boot | Disabled |
上傳速度 | 921600 |
上傳前清除所有 Flash | Disabled |
埠 | CH340K USB-C 埠 |
用 ESP32 核心 2.0.17,唔好用 3.x。Espressif 喺核心 3 度移除咗裝置上面嘅人臉偵測模型,所以人臉項目喺嗰度編譯唔到。鎖定 2.0.17 就可以用一個設定令呢頁所有項目都運作到。喺 Boards Manager 度,版本下拉選單可以隨時切換。
用 GFX Library for Arduino 版本 1.5.6,唔好用 1.6.x。1.6 版本係為 ESP32 核心 3 而整,喺核心 2.0.17 上面可能會喺啟動嗰陣卡住。如果你上傳之後螢幕保持黑色,呢個就係第一樣要檢查嘅嘢。
需要嘅函式庫
透過 Arduino IDE 嘅 Tools → Manage Libraries 安裝呢啲。版本號碼好重要——請用列咗出嚟嗰啲。
函式庫 | 版本 | 作者 |
|---|---|---|
GFX Library for Arduino | 1.5.6 | moononournation |
bb_captouch | 1.3.1 | Larry Bank |
ArduinoJson | 7.x | Benoit Blanchon |
Adafruit NeoPixel | 任何近期版本 | Adafruit |
設定 secrets.h
你嘅WiFi資料同任何API金鑰都放喺secrets.h入面,呢個檔案喺下載入面已經包含咗,入面有預設值。喺Arduino IDE打開嗰個分頁,然後將佢哋換成你自己嘅資料。
WiFi一定要係2.4 GHz。 ESP32-S3完全睇唔到5 GHz嘅網絡。如果你嘅路由器將兩個頻段合併喺同一個名稱下面(Asus叫呢個做Smart Connect),你要嘛關咗佢,要嘛將2.4 GHz頻段改一個獨立名稱,然後喺secrets.h入面用嗰個名稱。
攞你嘅API金鑰
呢個項目會連接一個雲端AI服務,所以你需要自己嘅金鑰。如果你從來未做過呢樣嘢,唔使擔心——佢同密碼嘅概念一樣,用嚟向服務識別你嘅帳戶。只需要幾分鐘,做一次就得。
金鑰唔係網站訂閱。例如,俾錢買ChatGPT Plus並唔會俾你一個API金鑰——兩者係獨立產品,有獨立收費。你需要喺開發者平台開一個帳戶,下面會詳細講解。
Microsoft Azure Speech——用嚟聽同講
Azure將你嘅語音轉成文字,再將答案轉返做語音。免費方案嘅額度對呢頁所有嘢嚟講都好充足。
去portal.azure.com,用Microsoft帳戶登入(免費帳戶都得)。
如果你從來未用過Azure,你會見到一個Welcome to Azure畫面,有三個選項。揀Start with an Azure free trial——你需要先有訂閱,Azure先俾你建立任何嘢。(學生應該揀Azure for Students:結果一樣,唔使信用卡。)忽略Manage Microsoft Entra ID,嗰個係完全另一樣嘢。
撳Create a resource,搜尋Speech,然後揀由Microsoft發佈嘅Speech service。
填表格:任何資源群組、任何名稱,然後揀一個近你嘅Region——將嗰個區域名稱原樣寫低,例如
eastus。喺Pricing tier揀F0 (Free)。呢個每個月容許大約五個鐘語音轉文字,同五十萬個字元文字轉語音。
撳Review + create,然後撳Create。等大約一分鐘,然後撳Go to resource。
喺左邊選單打開Keys and Endpoint。複製KEY 1同Location/Region。
將呢啲放入secrets.h,分別係AZURE_SPEECH_KEY同AZURE_REGION。至於AZURE_STT_HOST,用<region>.stt.speech.microsoft.com——所以如果區域係eastus,就係eastus.stt.speech.microsoft.com。
關於信用卡。Azure免費試用要求一張卡嚟驗證你嘅身份。佢唔會收你錢。你會得到$200信用額,為期30日,之後帳戶會轉去Pay-As-You-Go——但F0 Speech方案仍然免費,每個月都係,而呢啲項目入面所有嘢都完全喺呢個額度之內。如果你完全唔想俾信用卡,而你又係學生,Azure for Students選項可以喺冇卡嘅情況下俾你信用額。
一定要係「Speech service」資源。嚟自Translator、Language或一般Cognitive Services資源嘅金鑰睇落一模一樣,而且完全有效——但每個語音請求都會回傳錯誤401。我哋喺測試期間就遇到呢個問題,嘥咗一個鐘。如果語音功能喺金鑰睇落正確嘅情況下回傳401,檢查吓你建立咗邊種資源。
DeepSeek——思考嘅部分
DeepSeek係實際回答你問題嘅語言模型。佢好平——幾蚊美金嘅信用額就夠覆蓋幾千次回覆。
去platform.deepseek.com開一個帳戶。
喺選單打開API keys,然後撳Create new API key。
即刻複製佢。佢只會顯示一次,之後唔會再顯示——如果你整唔見咗,刪除嗰個金鑰再整過一個。
喺Top up下面加少量信用額。冇免費方案,但最細嘅增值額以呢個使用量嚟講可以用好耐。
將金鑰放入secrets.h,作為DEEPSEEK_KEY。佢以sk-開頭。
模型名稱喺2026年7月改咗。舊嘅deepseek-chat同deepseek-reasoner已經退役,所以網上搵到嘅大部分教學都會回傳400錯誤。用deepseek-v4-flash,呢啲項目已經設定咗呢個。
運行成本大約幾多
好少,但唔係免費,你應該喺留低一個項目運行之前大概知道你會用幾多錢。
服務 | 大約成本 |
|---|---|
Azure Speech | 免費方案每個月覆蓋大約5個鐘聆聽同0.5 M字元講嘢 |
DeepSeek | 每個答案只係一仙嘅零頭——幾蚊美金就有幾千次回覆 |
OpenAI vision | 每張圖片大約一至兩仙,視乎模型而定 |
價格會變,所以將呢啲當做參考而唔係報價。呢啲服務每一個都有使用量頁面,你可以睇吓自己用咗幾多錢,而且全部都可以設定消費上限——呢個值得喺第一日就做。
將你嘅密鑰保持私密。 任何擁有密鑰嘅人都可以用你嘅錢。唔好將密鑰放喺影片、截圖、論壇帖子或者公開嘅代碼儲存庫入面。如果密鑰一旦洩露,就要喺供應商網站上刪除佢並建立一個新嘅——只需幾秒鐘,而且係唯一真正嘅解決方法。
疑難排解
症狀 | 原因同解決方法 |
|---|---|
畫面保持黑色 | GFX 庫版本錯誤(請用 1.5.6)或者開發板設定錯誤。 |
|
|
上傳唔到 / 冇 COM 埠 | 用錯 USB-C 埠,或者 CH340 驅動程式未安裝。 |
相機故障而且永遠恢復唔到 | 相機嘅重置線連住開發板嘅 RESET 按鈕,所以軟件無法重新啟動佢。請按 RESET。如果仍然失敗,請重新插好相機嘅排線。 |
下載代碼
呢個項目嘅完整 Arduino 草圖,連同 pins.h 同其他所需檔案,都可以免費下載。
解壓縮之後,喺 Arduino IDE 入面打開 .ino 檔案,檢查上面嘅設定,然後透過 CH340K USB-C 埠上傳。
This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
/*
* ===========================================================================
* 04_Voice_Assistant — MaTouch AI ESP32-S3 2.8" TFT ST7789V
* ===========================================================================
*
----------
* ROBOJAX.COM - MaTouch AI ESP32-S3 2.8" project series
*
* WATCH THE VIDEO
* https://youtu.be/6AL3g3tC_Hk
*
* WRITTEN TUTORIALS - every project, with photos and full explanation
* Camera and touchscreen.... https://robojax.com/RTJ849
* Offline face recognition.. https://robojax.com/RTJ850
* AI voice assistant........ https://robojax.com/RTJ851
* AI vision................. https://robojax.com/RTJ852
*
* GET THE BOARD - SAVE $5 with coupon code: Robojax_Makerfab
* https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
* (enter the code at checkout)
*
* All of this code is free. If it helped you, a subscribe on YouTube is
* the best way to support more of it.
*
* ---------------------------------------------------------------------------
*
* A complete voice assistant on a $40 board:
*
* hold SPEAK -> both INMP441 microphones record you
* -> Azure Speech turns the audio into text
* -> DeepSeek v4-flash thinks of an answer
* -> Azure Speech turns the answer into audio
* -> the MAX98357 speaker says it out loud
*
* and the whole conversation is drawn as chat bubbles on the touchscreen.
*
* WHY THIS COMBINATION OF SERVICES (each is used where it is best):
* - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
* transcription endpoint wants multipart/form-data - miserable on an MCU.
* - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
* which streams straight into the I2S speaker with NO MP3 decoder at all.
* - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
* model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
* Most tutorials online still use them and are broken. See secrets.h.
*
* A detail the vendor examples get wrong: this board has TWO microphones on
* one I2S bus (left + right), but every Makerfabs demo records left-only and
* throws one away. This sketch records both and averages them.
*
* ---------------------------------------------------------------------------
* *** WHICH USB PORT - THIS MATTERS ***
* The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
* through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
* If you use the wrong port the audio will be garbage or uploads will fail.
* ---------------------------------------------------------------------------
*
* FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
*
* BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
* or compile errors. These reset when you switch cores - recheck them!)
*
* Board : ESP32S3 Dev Module
* ESP32 core : 2.0.17
* PSRAM : OPI PSRAM <-- required, audio buffer lives there
* Flash Size : 16MB (128Mb)
* Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
* USB CDC On Boot : Disabled <-- required, see USB note above
* Upload Speed : 921600
* Port : the CH340K USB-C port (the one near RESET)
*
* LIBRARIES
* GFX Library for Arduino v1.5.6 (NOT 1.6.x - that pairs with core 3)
* bb_captouch v1.3.1
* ArduinoJson v7.x
* Adafruit NeoPixel any recent
*
* ---------------------------------------------------------------------------
* FUNCTIONS IN THIS SKETCH
* led(r,g,b) + LED_* macros RGB status colours (blue/amber/green/red)
* getTouch(&x,&y) read the touch panel, mapped to screen coordinates
* speakButtonHeld() true while a finger is on the SPEAK button
* bubbleLines(t) how many lines a message wraps to
* drawOneBubble(m,y) draw a single chat bubble
* redrawChat() rebuild the chat area from history, newest at bottom
* clearChat() wipe the chat history (CLEAR button)
* chatBubble(t,user) add a message to history and redraw
* drawWifi() WiFi signal bars + dBm readout
* drawBar(label,col) bottom bar: SPEAK button + CLEAR + WiFi meter
* micInit() I2S input - BOTH INMP441 mics, stereo
* spkInit() I2S output - MAX98357 speaker
* recordWhileHeld() record while SPEAK held, downmix stereo->mono
* writeWavHeader(...) prepend the 44-byte RIFF/WAVE header
* dumpWavToSD(...) save the exact upload to SD (/stt_debug.wav)
* readHttpResponse() read an HTTPS reply, de-chunking it properly
* azureSTT(...) chunked upload of the WAV -> recognised text
* deepseekChat(...) question -> deepseek-v4-flash -> answer text
* azureTTSSpeak(text) answer -> Azure voice -> PSRAM -> speaker
* setup() / loop() boot + WiFi / one conversation turn per press
*
* Robojax.com
* ===========================================================================
*/
#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"
/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE 16000
#define RECORD_MAX_S 6 // hard cap on one question
#define WAV_HEADER_LEN 44
#define REC_BUF_BYTES (SAMPLE_RATE * RECORD_MAX_S * 2) // 16-bit mono
/* Software gain applied to the recording. The INMP441 capture is quiet at
* 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
* speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
* lower it. */
#define MIC_GAIN 6
/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
* 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
* back silent or garbled in stereo mode. */
#define USE_BOTH_MICS 1
/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
* uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS 25
/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
* shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);
BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);
/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr; // WAV_HEADER_LEN + up to REC_BUF_BYTES
/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;
/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool ok_sd = false;
float g_mic_peak_pct = 0; // last recording's raw peak, % of full scale
char g_stt_err[64] = ""; // last STT failure cause, verbatim
char g_llm_err[64] = ""; // last DeepSeek failure cause, verbatim
/* --- layout --------------------------------------------------------------- */
#define CHAT_H 200 // chat area: y 0..199
#define BAR_Y 202 // button bar below it
#define BTN_SPEAK_X 4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W 58
#define WIFI_X 236 // signal indicator, right end of the bar
#define BTN_H 36
/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
* This is what prevents new text printing over old - the whole area is
* rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int chat_count = 0;
/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
rgb.setPixelColor(0, rgb.Color(r, g, b));
rgb.show();
}
#define LED_IDLE() led(0, 0, 0)
#define LED_LISTEN() led(0, 60, 255) // blue - recording
#define LED_THINK() led(255, 120, 0) // amber - waiting on the cloud
#define LED_SPEAK() led(0, 255, 40) // green - talking
#define LED_ERROR() led(255, 0, 0) // red
/* ===========================================================================
* Touch
* =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
TOUCHINFO ti;
if (!bbct.getSamples(&ti)) return false;
if (ti.count < 1) return false;
*x = ti.y[0];
*y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
return true;
}
bool speakButtonHeld() {
uint16_t x, y;
if (!getTouch(&x, &y)) return false;
return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}
/* ===========================================================================
* Chat UI — word-wrapped bubbles, user right/blue, assistant left/grey
* =========================================================================== */
#define CHAT_CHARS 42 // chars per line at textsize 1
static int bubbleLines(const char *t) {
int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
return l < 1 ? 1 : l;
}
void drawOneBubble(const ChatMsg &m, int y) {
int len = strlen(m.text);
int lines = bubbleLines(m.text);
int h = lines * 10 + 8;
uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
if (w < 30) w = 30;
int x = m.from_user ? (316 - w) : 4;
gfx->fillRoundRect(x, y, w, h, 5, bg);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
for (int i = 0; i < lines; i++) {
char line[CHAT_CHARS + 1] = {0};
strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
gfx->setCursor(x + 5, y + 5 + i * 10);
gfx->print(line);
}
}
/* Rebuild the whole chat area from history: newest message anchored at the
* bottom, older ones stacked upward until the area is full. */
void redrawChat() {
gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
int shown = min(chat_count, CHAT_HISTORY);
int y = CHAT_H - 2;
for (int i = 0; i < shown; i++) {
ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
int h = bubbleLines(m.text) * 10 + 8;
y -= h;
if (y < 0) break; // area full - older ones drop off
drawOneBubble(m, y);
y -= 4;
}
}
void clearChat() {
chat_count = 0;
redrawChat();
}
void chatBubble(const char *text, bool from_user) {
ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
strncpy(m.text, text, sizeof(m.text) - 1);
m.text[sizeof(m.text) - 1] = 0;
m.from_user = from_user;
chat_count++;
redrawChat();
}
/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
bool up = (WiFi.status() == WL_CONNECTED);
long rssi = up ? WiFi.RSSI() : -100;
// -55 dBm or better = full bars; each 10 dB drops one
int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;
for (int b = 0; b < 4; b++) {
int bh = 6 + b * 6; // heights 6,12,18,24
uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
}
gfx->setTextSize(1);
gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
if (up) {
gfx->setTextColor(CYAN);
gfx->printf("%lddBm", rssi);
} else {
gfx->setTextColor(RED);
gfx->print("DOWN");
}
gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
gfx->setTextColor(gfx->color565(120, 120, 120));
gfx->print("WiFi");
}
void drawBar(const char *label, uint16_t colour) {
gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
gfx->setTextSize(2);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
gfx->print(label);
// CLEAR wipes the chat history
gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
gfx->print("CLEAR");
drawWifi();
}
/* ===========================================================================
* I2S — microphones on port 0, speaker on port 1. Separate hardware
* ports, so recording and playback can never fight over a bus.
* =========================================================================== */
void micInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
/* BOTH channels - this is the two-microphone fix. The vendor examples
* use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
* falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
.channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = false,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_MIC_SCK,
.ws_io_num = I2S_MIC_WS,
.data_out_num = I2S_PIN_NO_CHANGE,
.data_in_num = I2S_MIC_SD
};
i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_MIC_PORT, &pins);
}
void spkInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT, // mono - the amp downmixes anyway
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = true,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_SPK_BCLK,
.ws_io_num = I2S_SPK_LRC,
.data_out_num = I2S_SPK_DOUT,
.data_in_num = I2S_PIN_NO_CHANGE
};
i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_SPK_PORT, &pins);
i2s_zero_dma_buffer(I2S_SPK_PORT);
}
/* ===========================================================================
* Recording — runs while the SPEAK button is held (up to RECORD_MAX_S).
* Reads stereo pairs, averages L+R into one mono stream, applies a little
* software gain, and fills wav_buf after the 44-byte header slot.
* Returns the number of audio bytes recorded.
* =========================================================================== */
size_t recordWhileHeld() {
int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
size_t mono_samples = 0;
const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;
int16_t chunk[512];
uint32_t last_touch_ok = millis();
int32_t peak = 0; // loudest raw sample - mic health check
i2s_zero_dma_buffer(I2S_MIC_PORT);
while (mono_samples < max_samples) {
size_t got = 0;
i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);
#if USE_BOTH_MICS
size_t n = got / 4; // 4 bytes = one L+R pair
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
size_t n = got / 2; // 2 bytes = one mono sample
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = chunk[i];
#endif
if (abs(raw) > peak) peak = abs(raw);
int32_t mixed = raw * MIC_GAIN;
if (mixed > 32767) mixed = 32767;
if (mixed < -32768) mixed = -32768;
mono[mono_samples++] = (int16_t)mixed;
}
/* The GT911 is polled between I2S reads. A 250 ms grace period stops a
* momentary missed touch sample from cutting the recording short. */
if (speakButtonHeld()) last_touch_ok = millis();
else if (millis() - last_touch_ok > 250) break;
// live progress on the button
static uint32_t last_draw = 0;
if (millis() - last_draw > 200) {
last_draw = millis();
gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
(int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
}
}
/* Mic health line: peak as % of full scale BEFORE gain.
* 0% = the mic is not being read at all (config/pin problem)
* under 3% = too quiet - speak closer or raise MIC_GAIN
* 3-40% = healthy speech level
*/
g_mic_peak_pct = peak * 100.0 / 32768.0;
Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
g_mic_peak_pct, MIC_GAIN,
peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");
return mono_samples * 2;
}
/* Dump the exact WAV we are about to POST onto the SD card, so it can be
* played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
if (!ok_sd) return;
SD.remove("/stt_debug.wav");
File f = SD.open("/stt_debug.wav", FILE_WRITE);
if (!f) return;
f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
f.close();
Serial.println("debug copy saved to SD as /stt_debug.wav");
}
/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
uint32_t file_len = data_bytes + 36;
uint32_t byte_rate = SAMPLE_RATE * 2;
memcpy(h, "RIFF", 4); memcpy(h + 4, &file_len, 4);
memcpy(h + 8, "WAVEfmt ", 8);
uint32_t fmt_len = 16; memcpy(h + 16, &fmt_len, 4);
uint16_t fmt = 1, ch = 1; memcpy(h + 20, &fmt, 2); memcpy(h + 22, &ch, 2);
uint32_t rate = SAMPLE_RATE; memcpy(h + 24, &rate, 4); memcpy(h + 28, &byte_rate, 4);
uint16_t align = 2, bits = 16; memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
memcpy(h + 36, "data", 4); memcpy(h + 40, &data_bytes, 4);
}
/* ===========================================================================
* HTTP response reader — shared by all three cloud calls.
* Returns the status code and fills body_out. Handles chunked transfer
* encoding PROPERLY: the chunk-size markers must be stripped, or they end
* up embedded inside the JSON body and the parse fails on long replies.
* =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
* on timeout / closed-and-empty. Every read below goes through this, because
* a reasoning model can think for many seconds between the response headers
* and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
uint32_t t0 = millis();
while (!c.available()) {
if (!c.connected()) return false;
if (millis() - t0 > ms) return false;
delay(10);
}
return true;
}
static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
body_out = "";
if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
while (waitData(client, idle_ms)) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break; // blank line = end of headers
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
}
if (chunked) {
int blanks = 0;
while (true) {
/* The chunk-size line may not arrive for a long time while the model
* reasons. Waiting here - instead of letting read() time out - is the
* whole fix: a timed-out read looks exactly like "0" (final chunk),
* which silently truncated the body to nothing. */
if (!waitData(client, idle_ms)) {
Serial.println("HTTP: timed out waiting for the next chunk");
break;
}
String szline = client.readStringUntil('\n');
szline.trim();
if (szline.length() == 0) { // stray blank line
if (++blanks > 4) break;
continue;
}
blanks = 0;
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break; // genuine final chunk
long got = 0;
while (got < sz) {
if (!waitData(client, idle_ms)) break;
while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
}
if (waitData(client, 3000)) client.readStringUntil('\n'); // CRLF after chunk
if (got < sz) { Serial.println("HTTP: short chunk"); break; }
}
} else {
while (waitData(client, idle_ms))
while (client.available()) body_out += (char)client.read();
}
return code;
}
/* ===========================================================================
* CLOUD CALL 1 — Azure speech-to-text
* One POST, one header, plain WAV body. This is why Azure does the ears.
* =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
/* HTTPClient's one-shot POST fails on bodies this large (it attempts one
* giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
* function speaks HTTP directly and streams the WAV up in 4 KB chunks -
* reliable, and if it ever stalls we know the exact byte it stopped at. */
WiFiClientSecure client;
client.setInsecure(); // no cert bundle on-device; see notes
client.setTimeout(15); // seconds, for reads
writeWavHeader(wav_buf, audio_bytes);
dumpWavToSD(audio_bytes); // PC-playable copy of what we send
size_t total = WAV_HEADER_LEN + audio_bytes;
if (!client.connect(AZURE_STT_HOST, 443)) {
snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
Serial.println("STT: TLS connect failed");
return false;
}
/* Two valid host forms use DIFFERENT URL paths - detect which one is in
* secrets.h: <resource>.cognitiveservices.azure.com -> /stt/speech/...
* <region>.stt.speech.microsoft.com -> /speech/... */
bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
"/speech/recognition/conversation/cognitiveservices/v1"
"?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
"Host: " AZURE_STT_HOST "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
"Accept: application/json\r\n"
"Connection: close\r\n"
"Content-Length: " + String(total) + "\r\n\r\n";
client.print(req);
/* body, 4 KB at a time */
size_t sent = 0;
while (sent < total) {
size_t n = min((size_t)4096, total - sent);
size_t w = client.write(wav_buf + sent, n);
if (w == 0) {
delay(50); // brief stall - retry once
w = client.write(wav_buf + sent, n);
if (w == 0) {
snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
(unsigned)(sent / 1024));
Serial.printf("STT: upload stalled at %u/%u bytes\n",
(unsigned)sent, (unsigned)total);
client.stop();
return false;
}
}
sent += w;
yield();
}
Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);
/* read the reply with proper de-chunking */
String resp;
int code = readHttpResponse(client, resp, 10000);
client.stop();
if (code != 200) {
snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
return false;
}
const char *status = doc["RecognitionStatus"];
if (!status || strcmp(status, "Success") != 0) {
/* The status names the exact failure:
* InitialSilenceTimeout = Azure heard silence (mic level too low)
* NoMatch = heard sound but no recognisable words
* BabbleTimeout = heard only noise */
snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
return false;
}
text_out = doc["DisplayText"].as<String>();
if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
return text_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 2 — DeepSeek chat completion
* OpenAI-compatible format. Model name is deepseek-v4-flash - the old
* deepseek-chat name is dead, see secrets.h.
* =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
/* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
* bodies (long answers + the model's hidden reasoning), which shows up as
* "bad JSON reply". Reading until the server closes the connection is
* reliable regardless of reply length. */
JsonDocument req;
req["model"] = DEEPSEEK_MODEL;
req["max_tokens"] = LLM_MAX_TOKENS;
JsonArray msgs = req["messages"].to<JsonArray>();
JsonObject sys = msgs.add<JsonObject>();
sys["role"] = "system"; sys["content"] = SYSTEM_PROMPT;
JsonObject usr = msgs.add<JsonObject>();
usr["role"] = "user"; usr["content"] = question;
String body;
serializeJson(req, body);
WiFiClientSecure client;
client.setInsecure();
/* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
* ~2 s, but anything that needs actual working-out (an Ohm's law problem,
* say) can think for 10-30 s before sending a single byte. */
client.setTimeout(60);
if (!client.connect(DEEPSEEK_HOST, 443)) {
snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
Serial.println("LLM: TLS connect failed");
return false;
}
client.print(String("POST /chat/completions HTTP/1.1\r\n"
"Host: " DEEPSEEK_HOST "\r\n"
"Authorization: Bearer " DEEPSEEK_KEY "\r\n"
"Content-Type: application/json\r\n"
"Connection: close\r\n"
"Content-Length: ") + String(body.length()) + "\r\n\r\n");
client.print(body);
/* read the reply with proper de-chunking; generous window - long
* questions make the model think for a while before it responds */
String resp;
int code = readHttpResponse(client, resp, 60000);
client.stop();
if (code != 200) {
snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
return false;
}
const char *content = doc["choices"][0]["message"]["content"];
const char *finish = doc["choices"][0]["finish_reason"];
if (!content || !content[0]) {
/* v4-flash is a reasoning model: if finish_reason is "length", the whole
* token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
if (finish && strcmp(finish, "length") == 0)
snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
else
snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
return false;
}
answer_out = String(content);
answer_out.trim();
if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
return answer_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 3 — Azure text-to-speech, streamed straight to the speaker
* We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
* the I2S peripheral eats. Skip the 44-byte header, forward the rest.
* No MP3 decoder, no audio library, no buffering the whole reply.
* =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
size_t got = 0;
uint32_t t0 = millis();
while (got < n && millis() - t0 < 10000) {
int r = c.read(dst + got, n - got);
if (r > 0) { got += r; t0 = millis(); }
else if (!c.connected() && !c.available()) break;
else delay(2);
}
return got;
}
bool azureTTSSpeak(const String &text) {
// Escape the XML special characters for the SSML body
String safe = text;
safe.replace("&", "&");
safe.replace("<", "<");
safe.replace(">", ">");
String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
"<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";
/* Manual HTTP like the other two cloud calls - and for a hard reason:
* Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
* raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
* into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
* Here we parse the framing properly and keep only clean audio bytes. */
WiFiClientSecure client;
client.setInsecure();
client.setTimeout(20);
const char *host = AZURE_REGION ".tts.speech.microsoft.com";
if (!client.connect(host, 443)) {
Serial.println("TTS: TLS connect failed");
return false;
}
client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
"Host: ") + host + "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: application/ssml+xml\r\n"
"X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
"User-Agent: MaTouchRobojax\r\n"
"Connection: close\r\n"
"Content-Length: " + String(ssml.length()) + "\r\n\r\n");
client.print(ssml);
/* status + headers; note whether the body is chunked */
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
long content_len = -1;
while (client.connected() || client.available()) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break;
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
}
if (code != 200) {
Serial.printf("TTS HTTP %d\n", code);
client.stop();
return false;
}
const size_t AUDIO_CAP = 1200 * 1024; // ~37 s of speech
uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
if (!audio) { client.stop(); return false; }
size_t alen = 0;
if (chunked) {
/* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
while (true) {
String szline = client.readStringUntil('\n');
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break;
if (alen + sz > AUDIO_CAP) break;
size_t got = readExact(client, audio + alen, sz);
alen += got;
client.readStringUntil('\n'); // trailing CRLF after each chunk
if (got < (size_t)sz) break;
}
} else if (content_len > 0) {
alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
} else {
/* no framing info: read until the server closes */
uint32_t idle = millis();
while ((client.connected() || client.available()) && millis() - idle < 5000) {
int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
if (r > 0) { alen += r; idle = millis(); }
else delay(5);
}
}
client.stop();
Serial.printf("TTS: %u KB clean audio (%s), playing\n",
(unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");
bool ok = (alen > WAV_HEADER_LEN);
if (ok) {
/* NOW the audio actually starts - this is the honest moment to go green */
LED_SPEAK();
drawBar("SPEAKING...", gfx->color565(0, 130, 40));
/* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
static const uint8_t lead_in[640] = {0}; // 20 ms of silence
size_t w = 0;
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
}
free(audio);
// let the DMA buffers drain so the last word is not cut off
delay(150);
i2s_zero_dma_buffer(I2S_SPK_PORT);
return ok;
}
/* ===========================================================================
* SETUP
* =========================================================================== */
void setup() {
Serial.begin(115200);
delay(400);
Serial.println("\n=== 04 Voice Assistant | Robojax.com ===");
Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");
pinMode(TFT_BLK, OUTPUT);
digitalWrite(TFT_BLK, LOW);
pinMode(SD_CS, OUTPUT);
digitalWrite(SD_CS, HIGH);
// one shared SPI bus for TFT + SD (started before either device)
SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);
gfx->begin();
gfx->fillScreen(BLACK);
digitalWrite(TFT_BLK, HIGH);
// SD is optional here - it only stores the /stt_debug.wav diagnostic copy
ok_sd = SD.begin(SD_CS, SPI, 20000000);
digitalWrite(SD_CS, HIGH);
Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
: "SD not found - debug WAV dump disabled (not fatal)");
bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
delay(50);
rgb.begin();
rgb.setBrightness(LED_BRIGHTNESS);
LED_IDLE();
/* One recording buffer for the whole session, in PSRAM. This is the 8 MB
* that makes the board worth buying. */
wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
if (!wav_buf) {
gfx->setTextColor(RED);
gfx->setTextSize(2);
gfx->setCursor(10, 100);
gfx->print("PSRAM alloc failed!");
gfx->setTextSize(1);
gfx->setCursor(10, 130);
gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
while (1) delay(1000);
}
micInit();
spkInit();
gfx->setTextSize(1);
gfx->setTextColor(YELLOW);
gfx->setCursor(4, 4);
gfx->printf("Connecting to %s ...", WIFI_SSID);
Serial.printf("Connecting to %s ", WIFI_SSID);
WiFi.mode(WIFI_STA);
WiFi.begin(WIFI_SSID, WIFI_PASS);
uint32_t t0 = millis();
while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
delay(300);
Serial.print(".");
}
Serial.println();
clearChat();
if (WiFi.status() == WL_CONNECTED) {
Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
chatBubble("Hold SPEAK and ask me anything.", false);
} else {
chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
LED_ERROR();
}
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}
/* ===========================================================================
* LOOP — one full conversation turn per button press
* =========================================================================== */
void loop() {
/* CLEAR button: edge-detected so one tap wipes once. Reading the panel
* twice per loop (here and in speakButtonHeld) is fine - the GT911 just
* reports its current state. */
static bool tap_latch = false;
static uint8_t tap_release = 0;
if (state == ST_IDLE) {
uint16_t cx, cy;
if (getTouch(&cx, &cy)) {
tap_release = 0;
if (!tap_latch) {
tap_latch = true;
if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
clearChat();
chatBubble("Hold SPEAK and ask me anything.", false);
}
}
} else if (tap_latch && ++tap_release >= 4) {
tap_latch = false;
tap_release = 0;
}
/* live WiFi signal indicator, refreshed every 2 s while idle */
static uint32_t last_wifi = 0;
if (millis() - last_wifi > 2000) {
last_wifi = millis();
drawWifi();
}
}
if (state == ST_IDLE && speakButtonHeld()) {
/* ---- record ---- */
state = ST_RECORDING;
LED_LISTEN();
drawBar("LISTENING...", gfx->color565(0, 60, 200));
uint32_t t_rec = millis();
size_t audio_bytes = recordWhileHeld();
t_rec = millis() - t_rec;
Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);
if (audio_bytes < SAMPLE_RATE / 2) { // under a quarter second - a tap
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
/* ---- speech to text ---- */
state = ST_STT;
LED_THINK();
drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
uint32_t t_stt = millis();
String question;
if (!azureSTT(audio_bytes, question)) {
/* Show the REAL cause on screen - no serial monitor needed. */
char diag[96];
snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
g_stt_err, g_mic_peak_pct,
ok_sd ? " | saved /stt_debug.wav" : "");
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_stt = millis() - t_stt;
chatBubble(question.c_str(), true);
Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());
/* ---- think ---- */
state = ST_LLM;
drawBar("THINKING...", gfx->color565(150, 90, 0));
uint32_t t_llm = millis();
String answer;
if (!deepseekChat(question, answer)) {
char diag[96];
snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_llm = millis() - t_llm;
chatBubble(answer.c_str(), false);
Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());
/* ---- speak ----
* Still amber here: the voice has to be synthesised and downloaded first
* (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
* at the exact moment audio starts coming out of the speaker. */
state = ST_TTS;
LED_THINK();
drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
uint32_t t_tts = millis();
bool spoke = azureTTSSpeak(answer);
t_tts = millis() - t_tts;
/* Timing summary on serial - this feeds the "honest numbers" segment. */
Serial.printf("TIMINGS rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
audio_bytes / 32000.0, (unsigned long)t_stt,
(unsigned long)t_llm, (unsigned long)t_tts,
spoke ? "" : " (TTS FAILED)");
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
}
delay(20);
}
你可能需要嘅嘢
-
其他Product page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
資源與參考資料
-
文件记录Makerfabs MaTouch ESP32-S3 2.8" Camera and Touchscreen: user's manualwiki.makerfabs.com
-
文件记录
-
文件记录Product page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
-
下載Arduino GFX Library on Githubgithub.com
文件📁
所需文件 (.h)
其他文件
示意圖
-
MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematicThe latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF0.15 MB