This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
The latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
Makerfabs MaTouch ESP32-S3 2.8" kamara Gina AI Voice Assistant akan ESP32-S3 (Azure + DeepSeek)
Riƙe maɓalli, yi tambaya, kuma allo ya ba da amsa da murya
Cikakken mataimakin murya a kan ƙaramin allo ɗaya. Riƙe maɓallin SPEAK kuma ka yi tambaya. Allon yana rikodin ka da makirufo biyu, yana aika sautin zuwa Microsoft Azure don a mayar da shi rubutu, yana aika wannan rubutu zuwa DeepSeek don tunani, yana aika amsar zuwa Azure don a mayar da ita magana, kuma yana kunna ta ta lasifikar kansa. Dukan tattaunawar tana bayyana a kan allo kamar kumfa na taɗi.
Mataimakin AI na Murya yana aiki a kan MaTouch AI ESP32-S3 board
Abin da ke faruwa daga lokacin da ka danna maɓallin
Ga dukan tafiyar tambaya ɗaya, mataki-mataki. Ya cancanci karantawa sau ɗaya, domin duk abin da kake gani a kan allo da kuma LED yana daidaita zuwa ɗaya daga cikin waɗannan matakai.
Ka danna kuma ka RIƘE maɓallin SPEAK. Riƙe-don-magana ne, ba danna-don-magana ba: rikodin yana gudana daidai gwargwado muddin yatsanka ya rage, har zuwa daƙiƙa shida. LED ɗin matsayi yana juya shuɗi kuma maɓallin yana nuna LISTENING.
Makirufo biyu suna rikodin ka. Allon yana ɗaukar samfur sau 16,000 a cikin daƙiƙa ɗaya daga nau'in stereo, yana daidaita tashoshi biyu zuwa ɗaya, yana ƙara ɗan ƙararrawa, kuma yana adana sakamakon a cikin PSRAM. Maɓallin ci gaba yana rarrafe a kan maɓallin yayin da kake magana. Magana na daƙiƙa biyu kusan 64 KB ne.
Ka saki maɓallin. Rikodin yana tsayawa. Allon yana rubuta kan WAV na bytes 44 a gaban sautin - wannan ƙaramin lakabin shine abin da ke mayar da samfurori marasa ƙarfi zuwa fayil ɗin da Azure za ta karɓa.
Sautin yana zuwa Azure Speech-to-Text. Ana ɗora shi a cikin guda 4 KB a kan hanyar sadarwa mai aminci, kuma yana dawowa a matsayin layin rubutu ɗaya. Sautin ka na 64 KB ya zama kusan bytes 25 na rubutu. LED ɗin yana juya amber.
Tambayarka tana bayyana a kan allo a matsayin kumfa taɗi mai shuɗi, don haka kana iya ganin abin da ya ji daidai - wanda ke da amfani, domin kalmomin da ba a ji daidai ba suna bayyana mafi yawan amsoshi marasa kyau.
Rubutun yana zuwa DeepSeek. Allon yana aika tambayarka tare da umarni na tsaye don kiyaye amsoshi zuwa jimloli biyu gajeru. Samfurin yana tunani - da gaske yana tunani, samfurin tunani ne - kuma yana mayar da amsa.
Amsar tana bayyana a kan allo a matsayin kumfa mai launin toka. Kana iya karanta ta kafin ka ji ta.
Amsar tana komawa zuwa Azure don a faɗa ta. Allon yana neman sautin PCM na 16 kHz mara ƙarfi, wanda shine daidai tsarin da amplifier ɗinsa ke so, don haka babu mai buɗe MP3 a ko'ina cikin wannan aikin. Maɓallin yanzu yana karanta GETTING VOICE kuma LED ɗin yana kasancewa amber, domin babu abin da ake ji tukuna.
Dukan faifan yana saukarwa zuwa PSRAM kafin a kunna samfur ɗaya. Wannan yana da muhimmanci - duba bayanin da ke ƙasa.
Kunnawa. Lokacin da aka mika sautin ga lasifika, LED ɗin yana juya kore kuma maɓallin yana karanta SPEAKING. Kana jin amsar.
Me yasa ake saukar da sautin da farko maimakon kunna shi yayin da yake zuwa. Gudanar da kai tsaye daga hanyar sadarwa zuwa lasifika yana sauti kamar ƙwanƙwasa. Buffer ɗin lasifikar yana riƙe kusan kashi ɗaya cikin goma na daƙiƙa kawai, kuma kowane tsayawa a cikin canja wurin WiFi fiye da wannan yana ɓata shi, yana haifar da ƙwanƙwasa mai ji. Saukar da dukan amsar zuwa PSRAM da farko yana kashe kusan daƙiƙa ɗaya na ƙarin jira kuma yana cire kowane rata. Wannan ma shine dalilin da yasa nuni ke cewa GETTING VOICE kafin ya ce SPEAKING - biyun da gaske matakai daban-daban ne.
Yaya tsawon kowane mataki yake ɗauka
An auna shi a kan kayan aiki na gaske, don tambaya mai sauƙi:
Mataki | Lokaci na yau da kullun |
|---|---|
Rikodi | muddin ka riƙe maɓallin |
Magana-zuwa-rubutu (Azure) | kusan daƙiƙa 1.8 |
Tunani (DeepSeek) | kusan daƙiƙa 1.8 don tambaya mai sauƙi, ya fi tsayi sosai don wadda ke buƙatar ainihin aiki |
Ɗaukar muryar (Azure) | kusan daƙiƙa 7 - mafi girman yanki ɗaya |
Jimlar, saki zuwa sauti na farko | kusan daƙiƙa 11 |
Kowane musayar yana buga lokutan kansa zuwa serial monitor, don haka kana iya auna naka maimakon amincewa da waɗannan. Idan kana son ya yi sauri, canji mafi inganci shine neman amsoshi gajeru a cikin SYSTEM_PROMPT - ƙarancin rubutu don magana yana nufin ƙarancin sauti don ƙirƙira da saukarwa.
Allon kansa bai taɓa fahimtar komai ba. Manzo ne mai kyawawan kunnuwa da kyakkyawar murya - hankalin ana hayar shi da daƙiƙa.
Me yasa waɗannan ayyuka uku
Azure yana sarrafa magana a ciki da waje. Rubutu-zuwa-maganarsa na iya mayar da PCM na 16 kHz mara ƙarfi, wanda shine daidai abin da guntun lasifikar ke so, don haka babu mai buɗe MP3 a ko'ina cikin wannan aikin. Magana-zuwa-rubutunsa yana ɗaukar WAV mai sauƙi a cikin POST mai sauƙi.
DeepSeek shine kwakwalwar tattaunawa. Yana da sauri kuma yana kashe ƙaramin kaso na cent ɗaya kowace amsa.
OpenAI ba a amfani da shi anan - duba aikin 05, inda yake yin aikin gani.
Sunayen samfurin DeepSeek sun canza. Tsoffin sunaye deepseek-chat da deepseek-reasoner an janye su a cikin Yuli 2026. Yawancin koyarwar kan layi har yanzu suna amfani da su kuma za su dawo da kuskure. Sunayen na yanzu sune deepseek-v4-flash da deepseek-v4-pro. Wannan aikin yana amfani da v4-flash.
Tarkon samfurin tunani
DeepSeek v4-flash yana tunani kafin ya amsa, kuma wannan tunanin yana ƙidaya a kan iyakar alamar ku. Saita LLM_MAX_TOKENS ƙasa da yawa kuma duk kasafin kuɗin yana kashewa a kan tunani, amsar ta dawo ba komai, kuma allo ba ya faɗi komai. An saita shi zuwa 400 a nan saboda wannan dalili. Tambayoyi masu wuya kuma suna ɗaukar lokaci mai tsawo - gaskiya mai sauƙi tana amsawa cikin kusan daƙiƙa biyu, tambaya da ke buƙatar ainihin aiki na iya ɗaukar lokaci mai tsawo.
Karanta hasken matsayi
Launi | Ma'ana |
|---|---|
Shuɗi | sauraron ku |
Amber | girgije yana tunani, ko kuma ana ɗauko murya |
Kore | magana - wannan yana juya kore a daidai lokacin da sauti ya fara |
Ja | wani abu ya gaza - duba serial monitor |
Sarrafa kan allo
Taɗi yana gudana kamar tattaunawar waya, saƙonni mafi tsufa suna motsawa sama da kashewa. CLEAR yana goge shi. Mitar siginar WiFi tare da ainihin karatun dBm yana zaune a kusurwa, wanda ke da amfani lokacin da kuke mamakin ko jinkirin amsa shine hanyar sadarwa ko sabis.
Game da allo na MaTouch AI ESP32-S3 2.8"
Kowane aiki a wannan shafi yana gudana akan MaTouch AI ESP32-S3 2.8" TFT ST7789V daga Makerfabs. Allo ne gaba ɗaya: allon taɓawa mai launi, kyamara mai megapixel 3, microphones biyu da kuma ainihin amplifier na lasifika, duk ana sarrafa su ta ESP32-S3 tare da 8 MB na PSRAM. Wannan haɗin shine abin da ke sa waɗannan ayyukan AI su yiwu akan allo ɗaya ba tare da wani abu da aka haɗa ba.
8 MB na PSRAM yana da mahimmanci fiye da kowane lamba a nan. Shi ne abin da ke ba allo damar riƙe firam ɗin kyamara, ƴan daƙiƙan rikodin sauti, ko hoto mai ɓoye base64 a cikin ƙwaƙwalwa a lokaci ɗaya - babu ɗayansu da ya dace a cikin RAM na yau da kullun na ESP32.
Takardun masana'anta: Shafin wiki na Makerfabs.
Mahimman ƙayyadaddun bayanai
Processor: ESP32-S3, dual core 240 MHz, WiFi 2.4 GHz + Bluetooth 5.0
Ƙwaƙwalwa: 16 MB flash, 8 MB PSRAM (wanda kusan kowane aiki a nan ke buƙata)
Nuni: 2.8" IPS, 320×240, direban ST7789V, SPI
Taɓawa: GT911 capacitive, yana bin yatsu 5 a lokaci ɗaya
Kyamara: OV3660, megapixel 3, har zuwa 2048×1536
Microphones: INMP441 I2S dijital mics biyu (ainihin nau'in stereo)
Lasifika: MAX98357A class-D amplifier, 3.2 W cikin 4 Ω
Adana: ramin katin microSD (yanayin SPI)
Wutar lantarki: USB-C, mai haɗin baturi JST, caja TP4056, maɓallin wuta
Hakanan akan allo: WS2812B RGB LED, agogon lokaci na gaske PCF8563T mai baturi, da ma'aunin man baturi MAX17048 wanda ba a jera shi a cikin ƙayyadaddun bayanai na hukuma
Filin USB-C guda biyu ba iri ɗaya ba ne. Lasifikan allo yana raba siginonin sa (IO19 da IO20) tare da tashar USB ta asali, saboda waɗannan fil ɗin sune layukan bayanan USB da aka haɗa da hannu na ESP32-S3. Koyaushe ɗora da kunna wuta ta tashar USB-C CH340K (wanda ke gefen maɓallin RESET), kuma saita USB CDC On Boot zuwa Disabled. Yi amfani da tashar da ba daidai ba kuma sauti zai yi kuskure ko ɗorawa zai gaza.
Saitunan Arduino IDE
Waɗannan saitunan suna da mahimmanci. Yawancin matsalolin da mutane ke ba da rahoto game da wannan allo ɗaya ne daga cikin waɗannan ba daidai ba, kuma suna sake saitawa lokacin da kuka canza sigar core, don haka sake duba su bayan kowane canji.
Saiti | Ƙima |
|---|---|
Allo | ESP32S3 Dev Module |
Sigar core ESP32 | 2.0.17 |
PSRAM | OPI PSRAM |
Girman Flash | 16MB (128Mb) |
Tsarin Rarraba | 16M Flash (3MB APP/9.9MB FATFS) |
USB CDC On Boot | Disabled |
Saurin ɗorawa | 921600 |
Goge Duk Flash Kafin ɗorawa | Disabled |
Tashar | tashar USB-C CH340K |
Yi amfani da ESP32 core 2.0.17, ba 3.x ba. Espressif ya cire samfuran gano fuska na kan na'urar a cikin core 3, don haka ayyukan fuska ba za su iya haɗawa a can ba. Ɗaure 2.0.17 yana kiyaye kowane aiki a wannan shafi yana aiki tare da tsari ɗaya. A cikin Boards Manager, jerin zaɓin sigar yana ba ku damar canzawa baya da gaba duk lokacin da kuke so.
Yi amfani da GFX Library for Arduino sigar 1.5.6, ba 1.6.x ba. Sakin 1.6 an gina su don ESP32 core 3 kuma suna iya tsayawa a farkon farawa akan core 2.0.17. Idan allonku ya kasance baƙi bayan ɗorawa, wannan shine abu na farko da za ku duba.
Libraries da ake buƙata
Shigar da waɗannan ta Tools → Manage Libraries a cikin Arduino IDE. Lambobin siga suna da mahimmanci - da fatan za a yi amfani da waɗanda aka jera.
Library | Siga | Marubuci |
|---|---|---|
GFX Library for Arduino | 1.5.6 | moononournation |
bb_captouch | 1.3.1 | Larry Bank |
ArduinoJson | 7.x | Benoit Blanchon |
Adafruit NeoPixel | kowanne na kwanan nan | Adafruit |
Saita secrets.h
Bayanan WiFi ɗinka da duk wani API keys ɗinka suna shiga cikin secrets.h, wanda aka haɗa a cikin saukarwa tare da ƙimar wuraren da za a cika. Buɗe wannan shafin a cikin Arduino IDE kuma maye gurbin su da naka.
WiFi dole ne ya kasance 2.4 GHz. ESP32-S3 ba zai iya ganin hanyar sadarwa ta 5 GHz kwata-kwata ba. Idan router ɗinka ya haɗa nau'ikan biyu a ƙarƙashin suna ɗaya (Asus yana kiran wannan Smart Connect), ko dai kashe wannan ko kuma ba wa nau'in 2.4 GHz suna na kansa kuma yi amfani da shi a cikin secrets.h.
Samun API keys ɗinka
Wannan aikin yana magana da sabis ɗin AI na yanar gizo, don haka kana buƙatar key ɗinka na kanka. Idan ba ka taɓa yin wannan ba, kada ka damu - shi ne irin ra'ayin kalmar sirri da ke gano asusunka ga sabis ɗin. Yana ɗaukar ƴan mintuna, sau ɗaya.
Key ba biyan kuɗi ba ne ga gidan yanar gizo. Biyan kuɗi na ChatGPT Plus, alal misali, ba ya ba ka API key - waɗannan biyu samfura ne daban-daban tare da biyan kuɗi daban. Kana buƙatar asusu a dandalin masu haɓakawa, kamar yadda aka bayyana a ƙasa.
Microsoft Azure Speech - don saurare da magana
Azure yana juya maganarka zuwa rubutu kuma ya mayar da amsar zuwa murya. Matsayin kyauta ya isa ga duk abin da ke wannan shafin.
Je zuwa portal.azure.com kuma shiga da asusun Microsoft (na kyauta yana da kyau).
Idan ba ka taɓa amfani da Azure ba za ka ga allon Barka da zuwa Azure yana ba da zaɓuɓɓuka uku. Zaɓi Fara da gwajin kyauta na Azure - kana buƙatar biyan kuɗi kafin Azure ya bar ka ƙirƙiri wani abu. (Dalibai su zaɓi Azure don Dalibai maimakon: sakamako ɗaya, ba a buƙatar katin.) Yi watsi da Sarrafa Microsoft Entra ID, wanda wani abu ne daban gaba ɗaya.
Danna Ƙirƙiri albarkatu, bincika Speech, kuma zaɓi Sabis ɗin Speech wanda Microsoft ya buga.
Cika fom ɗin: kowane rukunin albarkatu, kowane suna, kuma zaɓi Yanki kusa da kai - rubuta wannan yankin daidai kamar yadda ya bayyana, misali
eastus.Don Matsayin farashi zaɓi F0 (Kyauta). Wannan yana ba da damar kusan sa'o'i biyar na magana-zuwa-rubutu da rabin miliyan haruffa na rubutu-zuwa-magana kowane wata.
Danna Bita + ƙirƙira, sannan Ƙirƙira. Jira kusan minti ɗaya, sannan danna Je zuwa albarkatu.
A cikin menu na hagu buɗe Keys da Endpoint. Kwafi KEY 1 da Wuri/Yanki.
Saka waɗannan a cikin secrets.h a matsayin AZURE_SPEECH_KEY da AZURE_REGION. Don AZURE_STT_HOST, yi amfani da <yanki>.stt.speech.microsoft.com - don haka tare da yankin eastus wannan shine eastus.stt.speech.microsoft.com.
Game da katin kuɗi. Gwajin kyauta na Azure yana neman kati don tabbatar da asalin ka. Ba ya cajin ka. Kana samun $200 na kuɗi na kwanaki 30, kuma bayan haka asusun yana motsawa zuwa Biya-Yayinda-Kake-Amfani - amma Matsayin F0 Speech yana kyauta, wata bayan wata, kuma duk abin da ke cikin waɗannan ayyukan ya dace a ciki. Idan ba ka so ka ba da kati kwata-kwata kuma kai ɗalibi ne, zaɓin Azure don Dalibai yana ba ka kuɗi ba tare da shi ba.
Dole ne ya zama albarkatun "Sabis ɗin Speech". Key daga albarkatun Translator, Language, ko Cognitive Services na gaba ɗaya yana kama da shi kuma yana da inganci - amma kowane buƙatun magana yana dawo da kuskure 401. Wannan ya kama mu yayin gwaji kuma ya ɗauki awa ɗaya. Idan magana ta gaza da 401 yayin da key ɗin ya yi daidai, duba wane irin albarkatu ka ƙirƙira.
DeepSeek - ɓangaren tunani
DeepSeek shine samfurin harshe wanda ke amsa tambayarka a zahiri. Ba shi da tsada - ƴan daloli na kuɗi sun rufe dubban amsoshi.
Je zuwa platform.deepseek.com kuma ƙirƙiri asusu.
Buɗe API keys a cikin menu kuma danna Ƙirƙiri sabon API key.
Kwafi shi nan da nan. Ana nuna shi sau ɗaya kawai kuma ba za a sake ba - idan ka rasa shi, share wannan key ɗin kuma ƙirƙiri wani.
Ƙara ƙaramin adadin kuɗi a ƙarƙashin Top up. Babu matakin kyauta, amma mafi ƙarancin ƙari yana daɗe sosai a wannan amfani.
Saka key ɗin a cikin secrets.h a matsayin DEEPSEEK_KEY. Yana farawa da sk-.
Sunayen samfura sun canza a Yuli 2026. Tsoffin deepseek-chat da deepseek-reasoner an janye su, don haka yawancin koyarwar da za ka samu akan layi za su gaza da kuskure 400. Yi amfani da deepseek-v4-flash, wanda waɗannan ayyukan suka riga suka saita.
Abin da yake kashewa don gudanarwa
Kaɗan ƙwarai, amma ba kyauta ba ne, kuma ya kamata ka san kusan abin da kake kashewa kafin ka bar aiki yana gudana.
Sabis | Kimanin farashi |
|---|---|
Azure Speech | matsayin kyauta yana rufe kusan sa'o'i 5 na saurare da 0.5 M haruffa na magana kowane wata |
DeepSeek | ɗan ƙaramin kaso na cent ɗaya kowace amsa - dubban amsoshi don ƴan daloli |
OpenAI vision | kusan cent ɗaya ko biyu kowace hoto, dangane da samfurin |
Farashin yana canzawa, don haka ɗauki waɗannan a matsayin jagora maimakon ƙididdiga. Kowane ɗayan waɗannan sabis ɗin yana da shafin amfani inda za ka iya kallon abin da ka kashe, kuma dukansu suna ba ka damar saita iyakar kashewa - wanda ya cancanci yi a rana ta farko.
Kawoya maɓallan ku a asirce. Duk wanda ke da su zai iya kashe kuɗin ku. Kada ku saka su a cikin bidiyo, hoton allo, rubutu a dandali ko wurin ajiyar lambar jama'a. Idan maɓalli ya taɓa bayyana, share shi a gidan yanar gizon mai bayarwa kuma ƙirƙiri sabo - yana ɗaukar daƙiƙa kaɗan, kuma shine kawai maganin gaske.
Magance Matsala
Alama | Dalili da magani |
|---|---|
Allon ya kasance baƙi | Ba daidai sigar GFX library ba (yi amfani da 1.5.6) ko saitunan allo ba daidai ba. |
|
|
Babu abin da ke lodawa / babu tashar COM | Ba daidai tashar USB-C ba, ko direban CH340 bai shigar ba. |
Kamara ta kasa kuma ba ta sake dawowa ba | Layin sake saitin kamara yana da alaƙa da maɓallin RESET na allo, don haka software ba zai iya sake kunna ta ba. Danna RESET. Idan har yanzu ta kasa, sake shigar da kebul ɗin kamara. |
Sauke lambar
Cikakken sketch na Arduino don wannan aikin, tare da pins.h da duk abin da yake buƙata, kyauta ne don saukewa.
Buɗe shi, buɗe fayil ɗin .ino a cikin Arduino IDE, duba saitunan da ke sama, kuma loda ta tashar CH340K USB-C.
This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
/*
* ===========================================================================
* 04_Voice_Assistant — MaTouch AI ESP32-S3 2.8" TFT ST7789V
* ===========================================================================
*
----------
* ROBOJAX.COM - MaTouch AI ESP32-S3 2.8" project series
*
* WATCH THE VIDEO
* https://youtu.be/6AL3g3tC_Hk
*
* WRITTEN TUTORIALS - every project, with photos and full explanation
* Camera and touchscreen.... https://robojax.com/RTJ849
* Offline face recognition.. https://robojax.com/RTJ850
* AI voice assistant........ https://robojax.com/RTJ851
* AI vision................. https://robojax.com/RTJ852
*
* GET THE BOARD - SAVE $5 with coupon code: Robojax_Makerfab
* https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
* (enter the code at checkout)
*
* All of this code is free. If it helped you, a subscribe on YouTube is
* the best way to support more of it.
*
* ---------------------------------------------------------------------------
*
* A complete voice assistant on a $40 board:
*
* hold SPEAK -> both INMP441 microphones record you
* -> Azure Speech turns the audio into text
* -> DeepSeek v4-flash thinks of an answer
* -> Azure Speech turns the answer into audio
* -> the MAX98357 speaker says it out loud
*
* and the whole conversation is drawn as chat bubbles on the touchscreen.
*
* WHY THIS COMBINATION OF SERVICES (each is used where it is best):
* - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
* transcription endpoint wants multipart/form-data - miserable on an MCU.
* - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
* which streams straight into the I2S speaker with NO MP3 decoder at all.
* - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
* model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
* Most tutorials online still use them and are broken. See secrets.h.
*
* A detail the vendor examples get wrong: this board has TWO microphones on
* one I2S bus (left + right), but every Makerfabs demo records left-only and
* throws one away. This sketch records both and averages them.
*
* ---------------------------------------------------------------------------
* *** WHICH USB PORT - THIS MATTERS ***
* The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
* through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
* If you use the wrong port the audio will be garbage or uploads will fail.
* ---------------------------------------------------------------------------
*
* FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
*
* BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
* or compile errors. These reset when you switch cores - recheck them!)
*
* Board : ESP32S3 Dev Module
* ESP32 core : 2.0.17
* PSRAM : OPI PSRAM <-- required, audio buffer lives there
* Flash Size : 16MB (128Mb)
* Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
* USB CDC On Boot : Disabled <-- required, see USB note above
* Upload Speed : 921600
* Port : the CH340K USB-C port (the one near RESET)
*
* LIBRARIES
* GFX Library for Arduino v1.5.6 (NOT 1.6.x - that pairs with core 3)
* bb_captouch v1.3.1
* ArduinoJson v7.x
* Adafruit NeoPixel any recent
*
* ---------------------------------------------------------------------------
* FUNCTIONS IN THIS SKETCH
* led(r,g,b) + LED_* macros RGB status colours (blue/amber/green/red)
* getTouch(&x,&y) read the touch panel, mapped to screen coordinates
* speakButtonHeld() true while a finger is on the SPEAK button
* bubbleLines(t) how many lines a message wraps to
* drawOneBubble(m,y) draw a single chat bubble
* redrawChat() rebuild the chat area from history, newest at bottom
* clearChat() wipe the chat history (CLEAR button)
* chatBubble(t,user) add a message to history and redraw
* drawWifi() WiFi signal bars + dBm readout
* drawBar(label,col) bottom bar: SPEAK button + CLEAR + WiFi meter
* micInit() I2S input - BOTH INMP441 mics, stereo
* spkInit() I2S output - MAX98357 speaker
* recordWhileHeld() record while SPEAK held, downmix stereo->mono
* writeWavHeader(...) prepend the 44-byte RIFF/WAVE header
* dumpWavToSD(...) save the exact upload to SD (/stt_debug.wav)
* readHttpResponse() read an HTTPS reply, de-chunking it properly
* azureSTT(...) chunked upload of the WAV -> recognised text
* deepseekChat(...) question -> deepseek-v4-flash -> answer text
* azureTTSSpeak(text) answer -> Azure voice -> PSRAM -> speaker
* setup() / loop() boot + WiFi / one conversation turn per press
*
* Robojax.com
* ===========================================================================
*/
#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"
/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE 16000
#define RECORD_MAX_S 6 // hard cap on one question
#define WAV_HEADER_LEN 44
#define REC_BUF_BYTES (SAMPLE_RATE * RECORD_MAX_S * 2) // 16-bit mono
/* Software gain applied to the recording. The INMP441 capture is quiet at
* 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
* speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
* lower it. */
#define MIC_GAIN 6
/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
* 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
* back silent or garbled in stereo mode. */
#define USE_BOTH_MICS 1
/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
* uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS 25
/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
* shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);
BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);
/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr; // WAV_HEADER_LEN + up to REC_BUF_BYTES
/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;
/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool ok_sd = false;
float g_mic_peak_pct = 0; // last recording's raw peak, % of full scale
char g_stt_err[64] = ""; // last STT failure cause, verbatim
char g_llm_err[64] = ""; // last DeepSeek failure cause, verbatim
/* --- layout --------------------------------------------------------------- */
#define CHAT_H 200 // chat area: y 0..199
#define BAR_Y 202 // button bar below it
#define BTN_SPEAK_X 4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W 58
#define WIFI_X 236 // signal indicator, right end of the bar
#define BTN_H 36
/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
* This is what prevents new text printing over old - the whole area is
* rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int chat_count = 0;
/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
rgb.setPixelColor(0, rgb.Color(r, g, b));
rgb.show();
}
#define LED_IDLE() led(0, 0, 0)
#define LED_LISTEN() led(0, 60, 255) // blue - recording
#define LED_THINK() led(255, 120, 0) // amber - waiting on the cloud
#define LED_SPEAK() led(0, 255, 40) // green - talking
#define LED_ERROR() led(255, 0, 0) // red
/* ===========================================================================
* Touch
* =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
TOUCHINFO ti;
if (!bbct.getSamples(&ti)) return false;
if (ti.count < 1) return false;
*x = ti.y[0];
*y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
return true;
}
bool speakButtonHeld() {
uint16_t x, y;
if (!getTouch(&x, &y)) return false;
return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}
/* ===========================================================================
* Chat UI — word-wrapped bubbles, user right/blue, assistant left/grey
* =========================================================================== */
#define CHAT_CHARS 42 // chars per line at textsize 1
static int bubbleLines(const char *t) {
int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
return l < 1 ? 1 : l;
}
void drawOneBubble(const ChatMsg &m, int y) {
int len = strlen(m.text);
int lines = bubbleLines(m.text);
int h = lines * 10 + 8;
uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
if (w < 30) w = 30;
int x = m.from_user ? (316 - w) : 4;
gfx->fillRoundRect(x, y, w, h, 5, bg);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
for (int i = 0; i < lines; i++) {
char line[CHAT_CHARS + 1] = {0};
strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
gfx->setCursor(x + 5, y + 5 + i * 10);
gfx->print(line);
}
}
/* Rebuild the whole chat area from history: newest message anchored at the
* bottom, older ones stacked upward until the area is full. */
void redrawChat() {
gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
int shown = min(chat_count, CHAT_HISTORY);
int y = CHAT_H - 2;
for (int i = 0; i < shown; i++) {
ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
int h = bubbleLines(m.text) * 10 + 8;
y -= h;
if (y < 0) break; // area full - older ones drop off
drawOneBubble(m, y);
y -= 4;
}
}
void clearChat() {
chat_count = 0;
redrawChat();
}
void chatBubble(const char *text, bool from_user) {
ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
strncpy(m.text, text, sizeof(m.text) - 1);
m.text[sizeof(m.text) - 1] = 0;
m.from_user = from_user;
chat_count++;
redrawChat();
}
/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
bool up = (WiFi.status() == WL_CONNECTED);
long rssi = up ? WiFi.RSSI() : -100;
// -55 dBm or better = full bars; each 10 dB drops one
int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;
for (int b = 0; b < 4; b++) {
int bh = 6 + b * 6; // heights 6,12,18,24
uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
}
gfx->setTextSize(1);
gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
if (up) {
gfx->setTextColor(CYAN);
gfx->printf("%lddBm", rssi);
} else {
gfx->setTextColor(RED);
gfx->print("DOWN");
}
gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
gfx->setTextColor(gfx->color565(120, 120, 120));
gfx->print("WiFi");
}
void drawBar(const char *label, uint16_t colour) {
gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
gfx->setTextSize(2);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
gfx->print(label);
// CLEAR wipes the chat history
gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
gfx->print("CLEAR");
drawWifi();
}
/* ===========================================================================
* I2S — microphones on port 0, speaker on port 1. Separate hardware
* ports, so recording and playback can never fight over a bus.
* =========================================================================== */
void micInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
/* BOTH channels - this is the two-microphone fix. The vendor examples
* use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
* falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
.channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = false,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_MIC_SCK,
.ws_io_num = I2S_MIC_WS,
.data_out_num = I2S_PIN_NO_CHANGE,
.data_in_num = I2S_MIC_SD
};
i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_MIC_PORT, &pins);
}
void spkInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT, // mono - the amp downmixes anyway
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = true,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_SPK_BCLK,
.ws_io_num = I2S_SPK_LRC,
.data_out_num = I2S_SPK_DOUT,
.data_in_num = I2S_PIN_NO_CHANGE
};
i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_SPK_PORT, &pins);
i2s_zero_dma_buffer(I2S_SPK_PORT);
}
/* ===========================================================================
* Recording — runs while the SPEAK button is held (up to RECORD_MAX_S).
* Reads stereo pairs, averages L+R into one mono stream, applies a little
* software gain, and fills wav_buf after the 44-byte header slot.
* Returns the number of audio bytes recorded.
* =========================================================================== */
size_t recordWhileHeld() {
int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
size_t mono_samples = 0;
const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;
int16_t chunk[512];
uint32_t last_touch_ok = millis();
int32_t peak = 0; // loudest raw sample - mic health check
i2s_zero_dma_buffer(I2S_MIC_PORT);
while (mono_samples < max_samples) {
size_t got = 0;
i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);
#if USE_BOTH_MICS
size_t n = got / 4; // 4 bytes = one L+R pair
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
size_t n = got / 2; // 2 bytes = one mono sample
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = chunk[i];
#endif
if (abs(raw) > peak) peak = abs(raw);
int32_t mixed = raw * MIC_GAIN;
if (mixed > 32767) mixed = 32767;
if (mixed < -32768) mixed = -32768;
mono[mono_samples++] = (int16_t)mixed;
}
/* The GT911 is polled between I2S reads. A 250 ms grace period stops a
* momentary missed touch sample from cutting the recording short. */
if (speakButtonHeld()) last_touch_ok = millis();
else if (millis() - last_touch_ok > 250) break;
// live progress on the button
static uint32_t last_draw = 0;
if (millis() - last_draw > 200) {
last_draw = millis();
gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
(int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
}
}
/* Mic health line: peak as % of full scale BEFORE gain.
* 0% = the mic is not being read at all (config/pin problem)
* under 3% = too quiet - speak closer or raise MIC_GAIN
* 3-40% = healthy speech level
*/
g_mic_peak_pct = peak * 100.0 / 32768.0;
Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
g_mic_peak_pct, MIC_GAIN,
peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");
return mono_samples * 2;
}
/* Dump the exact WAV we are about to POST onto the SD card, so it can be
* played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
if (!ok_sd) return;
SD.remove("/stt_debug.wav");
File f = SD.open("/stt_debug.wav", FILE_WRITE);
if (!f) return;
f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
f.close();
Serial.println("debug copy saved to SD as /stt_debug.wav");
}
/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
uint32_t file_len = data_bytes + 36;
uint32_t byte_rate = SAMPLE_RATE * 2;
memcpy(h, "RIFF", 4); memcpy(h + 4, &file_len, 4);
memcpy(h + 8, "WAVEfmt ", 8);
uint32_t fmt_len = 16; memcpy(h + 16, &fmt_len, 4);
uint16_t fmt = 1, ch = 1; memcpy(h + 20, &fmt, 2); memcpy(h + 22, &ch, 2);
uint32_t rate = SAMPLE_RATE; memcpy(h + 24, &rate, 4); memcpy(h + 28, &byte_rate, 4);
uint16_t align = 2, bits = 16; memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
memcpy(h + 36, "data", 4); memcpy(h + 40, &data_bytes, 4);
}
/* ===========================================================================
* HTTP response reader — shared by all three cloud calls.
* Returns the status code and fills body_out. Handles chunked transfer
* encoding PROPERLY: the chunk-size markers must be stripped, or they end
* up embedded inside the JSON body and the parse fails on long replies.
* =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
* on timeout / closed-and-empty. Every read below goes through this, because
* a reasoning model can think for many seconds between the response headers
* and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
uint32_t t0 = millis();
while (!c.available()) {
if (!c.connected()) return false;
if (millis() - t0 > ms) return false;
delay(10);
}
return true;
}
static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
body_out = "";
if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
while (waitData(client, idle_ms)) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break; // blank line = end of headers
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
}
if (chunked) {
int blanks = 0;
while (true) {
/* The chunk-size line may not arrive for a long time while the model
* reasons. Waiting here - instead of letting read() time out - is the
* whole fix: a timed-out read looks exactly like "0" (final chunk),
* which silently truncated the body to nothing. */
if (!waitData(client, idle_ms)) {
Serial.println("HTTP: timed out waiting for the next chunk");
break;
}
String szline = client.readStringUntil('\n');
szline.trim();
if (szline.length() == 0) { // stray blank line
if (++blanks > 4) break;
continue;
}
blanks = 0;
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break; // genuine final chunk
long got = 0;
while (got < sz) {
if (!waitData(client, idle_ms)) break;
while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
}
if (waitData(client, 3000)) client.readStringUntil('\n'); // CRLF after chunk
if (got < sz) { Serial.println("HTTP: short chunk"); break; }
}
} else {
while (waitData(client, idle_ms))
while (client.available()) body_out += (char)client.read();
}
return code;
}
/* ===========================================================================
* CLOUD CALL 1 — Azure speech-to-text
* One POST, one header, plain WAV body. This is why Azure does the ears.
* =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
/* HTTPClient's one-shot POST fails on bodies this large (it attempts one
* giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
* function speaks HTTP directly and streams the WAV up in 4 KB chunks -
* reliable, and if it ever stalls we know the exact byte it stopped at. */
WiFiClientSecure client;
client.setInsecure(); // no cert bundle on-device; see notes
client.setTimeout(15); // seconds, for reads
writeWavHeader(wav_buf, audio_bytes);
dumpWavToSD(audio_bytes); // PC-playable copy of what we send
size_t total = WAV_HEADER_LEN + audio_bytes;
if (!client.connect(AZURE_STT_HOST, 443)) {
snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
Serial.println("STT: TLS connect failed");
return false;
}
/* Two valid host forms use DIFFERENT URL paths - detect which one is in
* secrets.h: <resource>.cognitiveservices.azure.com -> /stt/speech/...
* <region>.stt.speech.microsoft.com -> /speech/... */
bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
"/speech/recognition/conversation/cognitiveservices/v1"
"?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
"Host: " AZURE_STT_HOST "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
"Accept: application/json\r\n"
"Connection: close\r\n"
"Content-Length: " + String(total) + "\r\n\r\n";
client.print(req);
/* body, 4 KB at a time */
size_t sent = 0;
while (sent < total) {
size_t n = min((size_t)4096, total - sent);
size_t w = client.write(wav_buf + sent, n);
if (w == 0) {
delay(50); // brief stall - retry once
w = client.write(wav_buf + sent, n);
if (w == 0) {
snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
(unsigned)(sent / 1024));
Serial.printf("STT: upload stalled at %u/%u bytes\n",
(unsigned)sent, (unsigned)total);
client.stop();
return false;
}
}
sent += w;
yield();
}
Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);
/* read the reply with proper de-chunking */
String resp;
int code = readHttpResponse(client, resp, 10000);
client.stop();
if (code != 200) {
snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
return false;
}
const char *status = doc["RecognitionStatus"];
if (!status || strcmp(status, "Success") != 0) {
/* The status names the exact failure:
* InitialSilenceTimeout = Azure heard silence (mic level too low)
* NoMatch = heard sound but no recognisable words
* BabbleTimeout = heard only noise */
snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
return false;
}
text_out = doc["DisplayText"].as<String>();
if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
return text_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 2 — DeepSeek chat completion
* OpenAI-compatible format. Model name is deepseek-v4-flash - the old
* deepseek-chat name is dead, see secrets.h.
* =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
/* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
* bodies (long answers + the model's hidden reasoning), which shows up as
* "bad JSON reply". Reading until the server closes the connection is
* reliable regardless of reply length. */
JsonDocument req;
req["model"] = DEEPSEEK_MODEL;
req["max_tokens"] = LLM_MAX_TOKENS;
JsonArray msgs = req["messages"].to<JsonArray>();
JsonObject sys = msgs.add<JsonObject>();
sys["role"] = "system"; sys["content"] = SYSTEM_PROMPT;
JsonObject usr = msgs.add<JsonObject>();
usr["role"] = "user"; usr["content"] = question;
String body;
serializeJson(req, body);
WiFiClientSecure client;
client.setInsecure();
/* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
* ~2 s, but anything that needs actual working-out (an Ohm's law problem,
* say) can think for 10-30 s before sending a single byte. */
client.setTimeout(60);
if (!client.connect(DEEPSEEK_HOST, 443)) {
snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
Serial.println("LLM: TLS connect failed");
return false;
}
client.print(String("POST /chat/completions HTTP/1.1\r\n"
"Host: " DEEPSEEK_HOST "\r\n"
"Authorization: Bearer " DEEPSEEK_KEY "\r\n"
"Content-Type: application/json\r\n"
"Connection: close\r\n"
"Content-Length: ") + String(body.length()) + "\r\n\r\n");
client.print(body);
/* read the reply with proper de-chunking; generous window - long
* questions make the model think for a while before it responds */
String resp;
int code = readHttpResponse(client, resp, 60000);
client.stop();
if (code != 200) {
snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
return false;
}
const char *content = doc["choices"][0]["message"]["content"];
const char *finish = doc["choices"][0]["finish_reason"];
if (!content || !content[0]) {
/* v4-flash is a reasoning model: if finish_reason is "length", the whole
* token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
if (finish && strcmp(finish, "length") == 0)
snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
else
snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
return false;
}
answer_out = String(content);
answer_out.trim();
if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
return answer_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 3 — Azure text-to-speech, streamed straight to the speaker
* We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
* the I2S peripheral eats. Skip the 44-byte header, forward the rest.
* No MP3 decoder, no audio library, no buffering the whole reply.
* =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
size_t got = 0;
uint32_t t0 = millis();
while (got < n && millis() - t0 < 10000) {
int r = c.read(dst + got, n - got);
if (r > 0) { got += r; t0 = millis(); }
else if (!c.connected() && !c.available()) break;
else delay(2);
}
return got;
}
bool azureTTSSpeak(const String &text) {
// Escape the XML special characters for the SSML body
String safe = text;
safe.replace("&", "&");
safe.replace("<", "<");
safe.replace(">", ">");
String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
"<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";
/* Manual HTTP like the other two cloud calls - and for a hard reason:
* Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
* raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
* into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
* Here we parse the framing properly and keep only clean audio bytes. */
WiFiClientSecure client;
client.setInsecure();
client.setTimeout(20);
const char *host = AZURE_REGION ".tts.speech.microsoft.com";
if (!client.connect(host, 443)) {
Serial.println("TTS: TLS connect failed");
return false;
}
client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
"Host: ") + host + "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: application/ssml+xml\r\n"
"X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
"User-Agent: MaTouchRobojax\r\n"
"Connection: close\r\n"
"Content-Length: " + String(ssml.length()) + "\r\n\r\n");
client.print(ssml);
/* status + headers; note whether the body is chunked */
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
long content_len = -1;
while (client.connected() || client.available()) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break;
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
}
if (code != 200) {
Serial.printf("TTS HTTP %d\n", code);
client.stop();
return false;
}
const size_t AUDIO_CAP = 1200 * 1024; // ~37 s of speech
uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
if (!audio) { client.stop(); return false; }
size_t alen = 0;
if (chunked) {
/* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
while (true) {
String szline = client.readStringUntil('\n');
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break;
if (alen + sz > AUDIO_CAP) break;
size_t got = readExact(client, audio + alen, sz);
alen += got;
client.readStringUntil('\n'); // trailing CRLF after each chunk
if (got < (size_t)sz) break;
}
} else if (content_len > 0) {
alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
} else {
/* no framing info: read until the server closes */
uint32_t idle = millis();
while ((client.connected() || client.available()) && millis() - idle < 5000) {
int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
if (r > 0) { alen += r; idle = millis(); }
else delay(5);
}
}
client.stop();
Serial.printf("TTS: %u KB clean audio (%s), playing\n",
(unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");
bool ok = (alen > WAV_HEADER_LEN);
if (ok) {
/* NOW the audio actually starts - this is the honest moment to go green */
LED_SPEAK();
drawBar("SPEAKING...", gfx->color565(0, 130, 40));
/* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
static const uint8_t lead_in[640] = {0}; // 20 ms of silence
size_t w = 0;
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
}
free(audio);
// let the DMA buffers drain so the last word is not cut off
delay(150);
i2s_zero_dma_buffer(I2S_SPK_PORT);
return ok;
}
/* ===========================================================================
* SETUP
* =========================================================================== */
void setup() {
Serial.begin(115200);
delay(400);
Serial.println("\n=== 04 Voice Assistant | Robojax.com ===");
Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");
pinMode(TFT_BLK, OUTPUT);
digitalWrite(TFT_BLK, LOW);
pinMode(SD_CS, OUTPUT);
digitalWrite(SD_CS, HIGH);
// one shared SPI bus for TFT + SD (started before either device)
SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);
gfx->begin();
gfx->fillScreen(BLACK);
digitalWrite(TFT_BLK, HIGH);
// SD is optional here - it only stores the /stt_debug.wav diagnostic copy
ok_sd = SD.begin(SD_CS, SPI, 20000000);
digitalWrite(SD_CS, HIGH);
Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
: "SD not found - debug WAV dump disabled (not fatal)");
bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
delay(50);
rgb.begin();
rgb.setBrightness(LED_BRIGHTNESS);
LED_IDLE();
/* One recording buffer for the whole session, in PSRAM. This is the 8 MB
* that makes the board worth buying. */
wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
if (!wav_buf) {
gfx->setTextColor(RED);
gfx->setTextSize(2);
gfx->setCursor(10, 100);
gfx->print("PSRAM alloc failed!");
gfx->setTextSize(1);
gfx->setCursor(10, 130);
gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
while (1) delay(1000);
}
micInit();
spkInit();
gfx->setTextSize(1);
gfx->setTextColor(YELLOW);
gfx->setCursor(4, 4);
gfx->printf("Connecting to %s ...", WIFI_SSID);
Serial.printf("Connecting to %s ", WIFI_SSID);
WiFi.mode(WIFI_STA);
WiFi.begin(WIFI_SSID, WIFI_PASS);
uint32_t t0 = millis();
while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
delay(300);
Serial.print(".");
}
Serial.println();
clearChat();
if (WiFi.status() == WL_CONNECTED) {
Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
chatBubble("Hold SPEAK and ask me anything.", false);
} else {
chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
LED_ERROR();
}
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}
/* ===========================================================================
* LOOP — one full conversation turn per button press
* =========================================================================== */
void loop() {
/* CLEAR button: edge-detected so one tap wipes once. Reading the panel
* twice per loop (here and in speakButtonHeld) is fine - the GT911 just
* reports its current state. */
static bool tap_latch = false;
static uint8_t tap_release = 0;
if (state == ST_IDLE) {
uint16_t cx, cy;
if (getTouch(&cx, &cy)) {
tap_release = 0;
if (!tap_latch) {
tap_latch = true;
if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
clearChat();
chatBubble("Hold SPEAK and ask me anything.", false);
}
}
} else if (tap_latch && ++tap_release >= 4) {
tap_latch = false;
tap_release = 0;
}
/* live WiFi signal indicator, refreshed every 2 s while idle */
static uint32_t last_wifi = 0;
if (millis() - last_wifi > 2000) {
last_wifi = millis();
drawWifi();
}
}
if (state == ST_IDLE && speakButtonHeld()) {
/* ---- record ---- */
state = ST_RECORDING;
LED_LISTEN();
drawBar("LISTENING...", gfx->color565(0, 60, 200));
uint32_t t_rec = millis();
size_t audio_bytes = recordWhileHeld();
t_rec = millis() - t_rec;
Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);
if (audio_bytes < SAMPLE_RATE / 2) { // under a quarter second - a tap
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
/* ---- speech to text ---- */
state = ST_STT;
LED_THINK();
drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
uint32_t t_stt = millis();
String question;
if (!azureSTT(audio_bytes, question)) {
/* Show the REAL cause on screen - no serial monitor needed. */
char diag[96];
snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
g_stt_err, g_mic_peak_pct,
ok_sd ? " | saved /stt_debug.wav" : "");
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_stt = millis() - t_stt;
chatBubble(question.c_str(), true);
Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());
/* ---- think ---- */
state = ST_LLM;
drawBar("THINKING...", gfx->color565(150, 90, 0));
uint32_t t_llm = millis();
String answer;
if (!deepseekChat(question, answer)) {
char diag[96];
snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_llm = millis() - t_llm;
chatBubble(answer.c_str(), false);
Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());
/* ---- speak ----
* Still amber here: the voice has to be synthesised and downloaded first
* (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
* at the exact moment audio starts coming out of the speaker. */
state = ST_TTS;
LED_THINK();
drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
uint32_t t_tts = millis();
bool spoke = azureTTSSpeak(answer);
t_tts = millis() - t_tts;
/* Timing summary on serial - this feeds the "honest numbers" segment. */
Serial.printf("TIMINGS rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
audio_bytes / 32000.0, (unsigned long)t_stt,
(unsigned long)t_llm, (unsigned long)t_tts,
spoke ? "" : " (TTS FAILED)");
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
}
delay(20);
}
Abin da zaka iya bukata
-
SauranProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
Albarkatun da kuma shafukan da za a duba
-
TakarduMakerfabs MaTouch ESP32-S3 2.8" Camera and Touchscreen: user's manualwiki.makerfabs.com
-
Takardu
-
TakarduProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
-
ZazzageArduino GFX Library on Githubgithub.com
Fayiloli📁
Fayil da ake buƙata (.h)
Sauran Fayiloli
Tsarin zane
-
MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematicThe latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF0.15 MB