Search Code

Makerfabs MaTouch ESP32-S3 2.8" kamarad Ku dhis AI Codka Caawiyaha ee ESP32-S3 (Azure + DeepSeek)

Makerfabs MaTouch ESP32-S3 2.8" kamarad Ku dhis AI Codka Caawiyaha ee ESP32-S3 (Azure + DeepSeek)

Riid batoon, weydi su'aal, oo looxa ayaa kuu jawaabaya cod ahaan

Kaaliye cod oo dhammaystiran oo ku yaal loox yar oo keliya. Qabo badhanka SPEAK oo weydii su'aal. Looxa ayaa kugu duubaya adiga oo isticmaalaya labada makarafoon, wuxuuna u diraa codka Microsoft Azure si loogu beddelo qoraal, wuxuuna u diraa qoraalkaas DeepSeek si ay uga fikirdo, wuxuuna u soo celiyaa jawaabta Azure si loogu beddelo hadal, wuxuuna ku ciyaaraa isagoo isticmaalaya qalabka codkiisa. Wadada hadalka oo dhan waxay ka soo muuqataa shaashadda sidii goobo sheeko ah.

Talking AI Voice Assistant running on the MaTouch AI ESP32-S3 board Kaaliye AI oo hadlaya oo ku shaqeeya looxa MaTouch AI ESP32-S3

Waxa dhaca laga bilaabo daqiiqadda aad riixdo batoonka

Waa kuwan safarka oo dhan ee su'aal keliya, tallaabo tallaabo. Waxaa mudan in hal mar la akhriyo, sababtoo ah wax kasta oo aad ku aragto shaashadda iyo LED-ga waxay u dhigmaan mid ka mid ah heerarkan.

  1. Waxaad riixdaa oo HAYSAA batoonka SPEAK. Waa riix-o-hay, ma aha taabo-o-hadal: duubista waxay socotaa inta uu farihiigu ku jiro hoos, ilaa lix ilbiriqsi. LED-ga heerka ayaa isu beddelaya buluug oo batoonku wuxuu muujinayaa LISTENING.

  2. Labada makarafoon ayaa kugu duubaya. Looxa wuxuu muunad ka qaadaa 16,000 jeer ilbiriqsi kasta laga soo bilaabo labada kanaal, wuxuuna isku celceliyaa labada kanaal hal, wuxuuna ku dabaqaa xoog yar, wuxuuna ku kaydiyaa natiijada PSRAM. Bar horumar ah ayaa ku gurguurta batoonka inta aad ku hadlaysid. Laba ilbiriqsi oo hadal ah waxay qiyaastii yihiin 64 KB.

  3. Waxaad sii daysaa batoonka. Duubistu way joogsataa. Looxa wuxuu ku qoraa madax WAV ah oo 44-byte ah hortiisa codka - calaamaddaas yar ayaa u beddesha muunadaha cayriin file ay Azure aqbali doonto.

  4. Codka wuxuu u tagayaa Azure Speech-to-Text. Waxaa la soo galiyaa qaybo 4 KB ah isku xir ammaan ah, waxaana ku soo noqdaa hal xariiq oo qoraal ah. 64 KB ee dhawaaqaaga waxay noqdeen qiyaastii 25 byte oo qoraal ah. LED-ga wuxuu isu beddelayaa amber.

  5. Su'aashaada waxay ka soo muuqataa shaashadda sidii goobo sheeko buluug ah, si aad u aragto waxa ay maqashay - taas oo faa'iido leh, sababtoo ah ereyada qaladka loo maqlay ayaa sharaxa inta badan jawaabaha aan caadiga ahayn.

  6. Qoraalku wuxuu u tagayaa DeepSeek. Looxa wuxuu u diraa su'aashaada oo ay weheliso tilmaan joogto ah oo ah in jawaabaha lagu koobo laba jumlad oo gaaban. Qaabku wuu fikiraa - si dhab ah ayuu u fikiraa, waa qaab sababayn ah - wuxuuna soo celiyaa jawaab.

  7. Jawaabtu waxay ka soo muuqataa shaashadda sidii goobo cawlan. Waad akhriyi kartaa ka hor intaadan maqal.

  8. Jawaabtu waxay ku soo noqotaa Azure si loogu hadlo. Looxa wuxuu weydiistaa cod PCM ah oo 16 kHz oo cayriin ah, kaas oo ah si sax ah qaabka uu qalabka codkiisu rabo, marka ma jiro decoder MP3 ah meelna ka mid ah mashruucan. Batoonku hadda wuxuu akhriyaa GETTING VOICE oo LED-gu wuxuu ku sii hadhaa amber, sababtoo ah wax maqal ah weli ma jiraan.

  9. Qaybta oo dhan waxaa la soo dejiyaa PSRAM ka hor inta aan hal muunad la ciyaarin. Tani waa muhiim - eeg qoraalka hoose.

  10. Ciyaarista. Daqiiqadda codka loo dhiibo qalabka codka, LED-ga wuxuu isu beddelaa cagaar oo batoonku wuxuu akhriyaa SPEAKING. Waxaad maqashaa jawaabta.

Sababta codka ugu horayn loo soo dejiyo halkii loo ciyaari lahaa marka uu yimaado. Si toos ah uga soo qulqulaya shabakadda una gala qalabka codka waxay u egtahay garaacid. Kaydka qalabka codka wuxuu qaadi karaa qiyaastii tobnaad ilbiriqsi, oo joogsi kasta oo ka yimaada wareejinta WiFi-ga oo ka dheer intaas ayaa maranaysa, taasoo soo saarta garaac maqal ah. Soo dejinta jawaabta oo dhan PSRAM ugu horayn waxay ku kacdaa qiyaastii ilbiriqsi sugitaan dheeri ah waxayna meesha ka saartaa dhammaan farqiyada. Sidaas ayaa sidoo kale sababta shaashaddu u sheegto GETTING VOICE ka hor inta aysan sheegin SPEAKING - labaduba si daacad ah waa heerar kala duwan.

Intee ayay qaadataa heer kasta

Lagu qiyaasay qalab dhab ah, su'aal fudud:

Heer

Waqti caadi ah

Duubista

intii aad hayso batoonka

Hadal-ku-qoraal (Azure)

qiyaastii 1.8 ilbiriqsi

Fikirka (DeepSeek)

qiyaastii 1.8 ilbiriqsi su'aal fudud, wax badan oo dheer mid u baahan shaqo dhab ah

Soo helista codka (Azure)

qiyaastii 7 ilbiriqsi - qaybta ugu weyn

Wadarta, sii daynta ilaa dhawaaqa ugu horreeya

qiyaastii 11 ilbiriqsi

Wada hadal kasta wuxuu daabacaa waqtiyadiisa gaarka ah kormeerka taxanaha, marka waad qiyaasi kartaa kuwaaga halkii aad aamini lahayd kuwan. Haddii aad rabto inay ka dhakhso badato, isbeddelka ugu waxtarka badan waa weydiista jawaabo gaaban gudaha SYSTEM_PROMPT - qoraal yar oo la hadlo waxay ka dhigan tahay cod yar oo la soo saaro oo la soo dejiyo.

Looxa laftiisu waligiis ma fahmo waxba. Waa farriin-wade dhego wanaagsan oo cod wanaagsan leh - caqligu waa mid kiro laga qaato ilbiriqsi kasta.

Sababta saddexdan adeeg

  • Azure waxay maamushaa hadalka soo gala iyo soo baxa. Qoraalka-ku-hadal ayaa soo celin kara PCM 16 kHz oo cayriin, kaas oo si sax ah u ah waxa qalabka codku rabo, marka ma jiro decoder MP3 ah meelna ka mid ah mashruucan. Hadal-ku-qoraalkiisu wuxuu qaataa WAV fudud oo ku jira POST fudud.

  • DeepSeek waa maskaxda wada hadalka. Waa mid degdeg ah oo ku kacda qayb senti oo jawaab kasta.

  • OpenAI halkan lama isticmaalo - eeg mashruuca 05, halkaas oo ay ku qabato shaqada aragga.

Magacyada qaabka DeepSeek ayaa isbeddelay. Magacyadii hore ee deepseek-chat iyo deepseek-reasoner waa la joojiyay Luulyo 2026. Inta badan tababarka khadka tooska ah ayaa weli isticmaalaya waxayna soo celin doonaan qalad. Magacyada hadda jira waa deepseek-v4-flash iyo deepseek-v4-pro. Mashruucani wuxuu isticmaalaa v4-flash.

Dabinka qaabka sababaynta

DeepSeek v4-flash wax ka fikiraa ka hor inta uu ka jawaabin, taasina waxay ku tirsan tahay xadkaaga token-ka. Haddii aad LLM_MAX_TOKENS u dhigto mid aad u hooseeya, dhammaan miisaaniyadda waxaa lagu bixiyaa fikirka, jawaabtu waxay soo noqotaa mid madhan, guddiguna waxba ma dhaho. Waxaa loo dhigay 400 halkan sababtaas awgeed. Su'aalaha adag waxay kaloo qaataan waqti dheer - jawaab fudud waxay ku soo noqotaa ilaa laba ilbiriqsi, su'aal u baahan shaqo dhab ah waxay qaadan kartaa waqti aad u dheer.

Akhriska nalka xaaladda

Midab

Macne

Buluug

ku dheggan yahay adiga

Jaalle

daruurta ayaa fikiraysa, ama codka ayaa la soo qaadayaa

Cagaar

hadlaya - tani waxay isu beddeshaa cagaar isla daqiiqadda codku bilaabmo

Casaan

wax baa ku guuldarreystay - hubi kormeeraha taxanaha ah

Kontaroolada shaashadda

Wada sheekaysigu wuxuu u socdaa sida wicitaan telefoon, fariimaha ugu da'da weyn ayaa kor u kacaya oo baxaya. CLEAR ayaa tirtira. Mitirka calaamadda WiFi oo leh akhrinta dBm dhabta ah ayaa ku fadhiya geeska, taas oo faa'iido u leh markaad la yaabban tahay in jawaabta gaabista ah ay tahay shabakadda ama adeegga.

Ku saabsan guddiga MaTouch AI ESP32-S3 2.8"

Mashruuc kasta oo ku yaal boggan wuxuu ku shaqeeyaa MaTouch AI ESP32-S3 2.8" TFT ST7789V oo ka socda Makerfabs. Waa guddi dhammeystiran: shaashad taabasho midab leh, kamarad 3 megapixel ah, laba makarafoon iyo cod-weyneeye dhab ah, dhammaantood waxaa wada shideeya ESP32-S3 oo leh 8 MB oo PSRAM ah. Isku darkaas ayaa ka dhigaya mashruucyadan AI-ga kuwa suurtogal ah guddi keliya oo aan wax kale ku xirnayn.

8 MB ee PSRAM waxay ka muhiimsan tahay lambar kasta oo kale halkan. Waa waxa u oggolaanaya guddiga inuu ku hayo qaab kamarad, dhawr ilbiriqsi oo cod duuban, ama sawir base64-ku-qoran xusuusta isku waqti - kuwaas oo midna aan ku habboonayn RAM-ka caadiga ah ee ESP32.

Dokumeentiyada soo saaraha: Bogga wiki-ga Makerfabs.

Sifooyinka muhiimka ah

  • Processor: ESP32-S3, laba xudun 240 MHz, WiFi 2.4 GHz + Bluetooth 5.0

  • Xusuusta: 16 MB flash, 8 MB PSRAM (loo baahan yahay ku dhawaad dhammaan mashruucyada halkan)

  • Shaashadda: 2.8" IPS, 320×240, darawalka ST7789V, SPI

  • Taabashada: GT911 capacitive, waxay raadraacdaa 5 farood isku mar

  • Kamaradda: OV3660, 3 megapixel, ilaa 2048×1536

  • Makarafoonnada: laba INMP441 I2S makarafoon dhijitaal ah (lamaan dhab ah oo stereo ah)

  • Cod-weyneeyaha: MAX98357A class-D amplifier, 3.2 W gudaha 4 Ω

  • Kaydinta: godka kaadhka microSD (qaabka SPI)

  • Korontada: USB-C, isku-xirka batteriga JST, charger-ka TP4056, furaha korontada

  • Sidoo kale guddiga ku yaal: WS2812B RGB LED, saacadda waqtiga-dhabta ah ee batteriga-ku-tiirsan PCF8563T, iyo cabbiraadda batteriga MAX17048 oo ah mid aan ku qoran sifooyinka rasmiga ah

Labada dekedood ee USB-C isku mid ma aha. Cod-weyneeyaha guddigu wuxuu la wadaagaa fiilooyinkiisa calaamadda (IO19 iyo IO20) dekedda asalka ah ee USB, sababtoo ah fiilooyinkaas waa khadadka xogta ee adag ee ESP32-S3. Had iyo jeer ku shub oo ku shide dekedda CH340K USB-C (ta ku xigta badhanka RESET), oo dhig USB CDC On Boot si Disabled. Haddii aad isticmaasho dekedda khaldan, codku wuu khaldami doonaa ama shubista way guuldarreysan doontaa.

Dejinta Arduino IDE

Dejintan ayaa muhiim ah. Inta badan dhibaatooyinka dadku ay ka sheegaan guddigan waa mid ka mid ah kuwan oo khaldan, waxayna dib isu dejiyaan markaad beddesho nooca xudunta, marka mar kale hubi ka dib isbeddel kasta.

Dejinta

Qiimaha

Guddiga

ESP32S3 Dev Module

Nooca xudunta ESP32

2.0.17

PSRAM

OPI PSRAM

Cabbirka Flash

16MB (128Mb)

Qorshaha Qaybinta

16M Flash (3MB APP/9.9MB FATFS)

USB CDC On Boot

Disabled

Xawaaraha Shubista

921600

Erase All Flash Before Upload

Disabled

Dekedda

dekedda CH340K USB-C

Isticmaal xudunta ESP32 2.0.17, ee ma aha 3.x. Espressif waxay ka saartay moodellada lagu ogaado wejiga ee qalabka ku jira xudunta 3, sidaa darteed mashruucyada wejiga kuma ururin doonaan halkaas. Ku xidhista 2.0.17 waxay ilaalisaa mashruuc kasta oo ku yaal boggan inuu shaqeeyo hal qaab. Guddiga Maareeyaha, liiska hoos-u-dhaca nooca ayaa kuu oggolaanaya inaad dib u beddesho wakhti kasta oo aad rabto.

Isticmaal GFX Library for Arduino nooca 1.5.6, ee ma aha 1.6.x. Siidaynta 1.6 waxaa loo dhisay xudunta ESP32 3 waxayna ku hakimi karaan bilowga xudunta 2.0.17. Haddii shaashaddaadu ay madow ku sii hadho ka dib shubista, tani waa waxa ugu horreeya ee la hubiyo.

Maktabadaha loo baahan yahay

Ku rakib kuwan iyada oo loo marayo Tools → Manage Libraries gudaha Arduino IDE. Nambarada noocyada ayaa muhiim ah - fadlan isticmaal kuwa la soo taxay.

Maktabadda

Nooca

Qoraaga

GFX Library for Arduino

1.5.6

moononournation

bb_captouch

1.3.1

Larry Bank

ArduinoJson

7.x

Benoit Blanchon

Adafruit NeoPixel

mid kasta oo dhow

Adafruit

Dejinta secrets.h

Faahfaahinta WiFi-gaaga iyo furay kasta oo API ah waxay galaan secrets.h, kaas oo ku jira soo dejinta oo leh qiimayaal meel lagu hayo. Fur tab-kaas Arduino IDE-ga oo ku beddel kuwaaga.

WiFi waa inuu ahaadaa 2.4 GHz. ESP32-S3 ma arki karo shabakad 5 GHz ah haba yaraatee. Haddii router-kaagu isku daro labada xariijin hal magac hoostiisa (Asus waxay tan ugu yeedhaa Smart Connect), ama dami taas ama sii xariijinta 2.4 GHz magac u gaar ah oo u isticmaal taas gudaha secrets.h.

Helitaanka furayaasha API-gaaga

Mashruucani wuxuu la hadlaa adeeg AI-ga daruuraha ah, marka waxaad u baahan tahay furahaaga gaarka ah. Haddii aadan waligaa tan samayn, ha welwelin - waa isla fikradda sirta ah ee aqoonsisa akoonkaaga adeegga. Waxay qaadataa dhowr daqiiqo, hal mar.

Furaha ma aha rukunka website-ka. Bixinta ChatGPT Plus, tusaale ahaan, ma siiyo furaha API - labaduba waa alaabooyin kala duwan oo leh bixino kala duwan. Waxaad u baahan tahay akoon ku yaal madalda horumarinta, sida hoos ku qoran.

Microsoft Azure Speech - dhegaysiga iyo hadalka

Azure waxay u beddeshaa hadalkaaga qoraal waxayna u celisaa jawaabta cod. Heerka bilaashka ah waa ku filan wax kasta oo ku yaal boggan.

  1. Tag portal.azure.com oo ku gal akoon Microsoft ah (mid bilaash ah ayaa fiican).

  2. Haddii aadan waligaa Azure isticmaalin waxaad arki doontaa shaashad Welcome to Azure oo bixisa saddex doorasho. Dooro Start with an Azure free trial - waxaad u baahan tahay rukun ka hor inta Azure kuu oggolaan inaad wax abuurto. (Ardaydu waa inay dooraan Azure for Students beddelkeeda: natiijo isku mid ah, kaadh looma baahna.) Iska indhatiir Manage Microsoft Entra ID, kaas oo ah wax kale oo gebi ahaanba ka duwan.

  3. Guji Create a resource, raadi Speech, oo dooro Speech service oo uu daabacay Microsoft.

  4. Buuxi foomka: koox kasta oo kheyraad ah, magac kasta, oo dooro Region kuu dhow - qor gobolkaas sida uu u muuqdo, tusaale ahaan eastus.

  5. Wixii Pricing tier dooro F0 (Free). Taasi waxay oggolaanaysaa qiyaastii shan saacadood oo hadal-qoraal ah iyo nus milyan oo xarfo ah oo qoraal-hadal ah bishiiba.

  6. Guji Review + create, ka dibna Create. Sug qiyaastii daqiiqad, ka dibna guji Go to resource.

  7. Menu-ga bidixda ku xiga fur Keys and Endpoint. Nuqul ka samee KEY 1 iyo Location/Region.

Ku rid kuwaas gudaha secrets.h sida AZURE_SPEECH_KEY iyo AZURE_REGION. Wixii AZURE_STT_HOST, isticmaal <region>.stt.speech.microsoft.com - marka gobolka eastus waa eastus.stt.speech.microsoft.com.

Ku saabsan kaadhka amaahda. Tijaabada bilaashka ah ee Azure waxay weydiisaa kaadh si ay u xaqiijiso aqoonsigaaga. Kuma dalacdo. Waxaad heleysaa $200 oo amaah ah 30 maalmood, ka dibna akoonku wuxuu u gudbaa Pay-As-You-Go - laakiin heerka F0 Speech wuxuu sii ahaanayaa bilaash, bil kasta, wax kasta oo ku yaal mashruucyadan ayaa si fiican ugu habboon gudihiisa. Haddii aad doorbidayso inaadan kaadh gebi ahaanba siin oo aad tahay arday, ikhtiyaarka Azure for Students wuxuu ku siinayaa amaah iyada oo aan lahayn.

Waa inay tahay kheyraad "Speech service". Furaha ka yimid kheyraadka Translator, Language, ama Cognitive Services guud wuxuu u muuqdaa isku mid oo si buuxda u shaqeeya - laakiin codsi kasta oo hadal ah wuxuu soo celinayaa qalad 401. Tani waxay na qabatay intii lagu jiray tijaabada waxayna qaadatay saacad. Haddii hadalku ku guuldarreysto 401 iyadoo furaha uu u muuqdo mid sax ah, hubi nooca kheyraadka aad abuurtay.

DeepSeek - qaybta fikirka

DeepSeek waa moodeelka luqadda ee dhab ahaan ka jawaaba su'aashaada. Waa mid jaban - dhowr dollar oo amaah ah ayaa daboolaya kumanaan jawaabood.

  1. Tag platform.deepseek.com oo abuur akoon.

  2. Fur API keys ee menu-ga oo guji Create new API key.

  3. Isla markiiba nuqul ka samee. Waxaa la muujiyaa hal mar oo keliya oo mar dambe lama soo celiyo - haddii aad lumiso, tirtir furahaas oo samee mid kale.

  4. Ku dar xaddi yar oo amaah ah hoosta Top up. Ma jiro heer bilaash ah, laakiin amaahda ugu yar waxay socotaa waqti aad u dheer isticmaalkan.

Ku rid furaha gudaha secrets.h sida DEEPSEEK_KEY. Wuxuu ku bilaabmaa sk-.

Magacyada moodeelka ayaa isbeddelay July 2026. Kuwa hore deepseek-chat iyo deepseek-reasoner waa la joojiyay, marka inta badan tababarka online-ka ah ee aad heli doonto waxay ku guuldarreysan doonaan qalad 400. Isticmaal deepseek-v4-flash, kaas oo ah kan mashruucyadan horey u dejiyeen.

Waxa ay ku kacdo in la wado

Aad u yar, laakiin ma aha bilaash, waxaadna tahay inaad si qiyaas ah u ogaato waxa aad ku bixinayso ka hor intaadan mashruuc ka tagin oo shaqeeya.

Adeeg

Qiimaha qiyaasta ah

Azure Speech

heerka bilaashka ah wuxuu daboolaa qiyaastii 5 saacadood oo dhegaysi ah iyo 0.5 M xarfo oo hadal ah bishiiba

DeepSeek

qayb senti ah jawaab kasta - kumanaan jawaabood dhowr dollar

OpenAI vision

qiyaastii hal ama laba senti sawirkiiba, iyadoo ku xiran moodeelka

Qiimuhu waa isbeddelaa, marka ula dhaqan kuwan sida hagaha halkii ay ka ahaan lahaayeen qiime rasmi ah. Mid kasta oo ka mid ah adeegyadan wuxuu leeyahay bog isticmaalka ah oo aad ku daawato waxa aad ku bixisay, dhammaantoodna waxay kuu oggolaanayaan inaad dejiso xadka kharashka - taas oo mudan in la sameeyo maalinta koowaad.

Furaha furayada gaarka ah qarsoodi. Qof kasta oo haysta waxay ku khasbi karaan inuu ku bixiyo lacagtaada. Ha gelin fiidiyow, sawir, qoraal forum ah ama keyd code ah oo dadweyne. Haddii furay lagu soo bandhigo, tirtir websaydhka bixiyaha oo ku samee mid cusub - waxay qaadataa ilbiriqsiyo, waana kaliya hagaajinta dhabta ah.

Cilad-baaris

Calaamadda

Sababta iyo hagaajinta

Shaashadu waxay ku sii jirtaa madow

Nooca maktabadda GFX khaldan (isticmaal 1.5.6) ama dejinta board-ka khaldan.

PSRAM alloc failed ama qalad kamarada 0xffffffff

Tools → PSRAM looma dejin OPI PSRAM.

Waxba ma soo gelinayaan / ma jiro COM port

USB-C port khaldan, ama darawalka CH340 lama rakibin.

Kamaradu way fashilantaa oo dib uma soo noqoto

Khadka reset-ka kamarada wuxuu ku xiran yahay badhanka RESET ee board-ka, sidaas darteed software-ku ma soo celin karo. Riix RESET. Haddii ay weli fashilanto, dib u dhig xadhkaha ribbon-ka kamarada.

Soo deji code-ka

Qoraalka Arduino ee dhammaystiran ee mashruucan, oo ay weheliyaan pins.h iyo wax kasta oo kale oo uu u baahan yahay, waa bilaash in la soo dejiyo.

Soo deji 04_Voice_Assistant

Fur, fur faylka .ino gudaha Arduino IDE, hubi dejinta kor ku xusan, oo ku soo geli port-ka CH340K USB-C.

884-Arduin code for MaTouch AI ESP32S3 2.8in AI Camera: AI voice assistant
Luqadda: C++
/*
 * ===========================================================================
 *  04_Voice_Assistant  —  MaTouch AI ESP32-S3 2.8" TFT ST7789V
 * ===========================================================================
 *
 ----------
 *  ROBOJAX.COM  -  MaTouch AI ESP32-S3 2.8" project series
 *
 *    WATCH THE VIDEO
 *        https://youtu.be/6AL3g3tC_Hk
 *
 *    WRITTEN TUTORIALS - every project, with photos and full explanation
 *        Camera and touchscreen.... https://robojax.com/RTJ849
 *        Offline face recognition.. https://robojax.com/RTJ850
 *        AI voice assistant........ https://robojax.com/RTJ851
 *        AI vision................. https://robojax.com/RTJ852
 *
 *    GET THE BOARD - SAVE $5 with coupon code:  Robojax_Makerfab
 *        https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
 *        (enter the code at checkout)
 *
 *  All of this code is free. If it helped you, a subscribe on YouTube is
 *  the best way to support more of it.
 *  
 *  ---------------------------------------------------------------------------
 *
 *  A complete voice assistant on a $40 board:
 *
 *      hold SPEAK  ->  both INMP441 microphones record you
 *                  ->  Azure Speech turns the audio into text
 *                  ->  DeepSeek v4-flash thinks of an answer
 *                  ->  Azure Speech turns the answer into audio
 *                  ->  the MAX98357 speaker says it out loud
 *
 *  and the whole conversation is drawn as chat bubbles on the touchscreen.
 *
 *  WHY THIS COMBINATION OF SERVICES (each is used where it is best):
 *   - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
 *     transcription endpoint wants multipart/form-data - miserable on an MCU.
 *   - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
 *     which streams straight into the I2S speaker with NO MP3 decoder at all.
 *   - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
 *     model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
 *     Most tutorials online still use them and are broken. See secrets.h.
 *
 *  A detail the vendor examples get wrong: this board has TWO microphones on
 *  one I2S bus (left + right), but every Makerfabs demo records left-only and
 *  throws one away. This sketch records both and averages them.
 *
 *  ---------------------------------------------------------------------------
 *  *** WHICH USB PORT - THIS MATTERS ***
 *  The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
 *  through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
 *  If you use the wrong port the audio will be garbage or uploads will fail.
 *  ---------------------------------------------------------------------------
 *
 *  FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
 *
 *  BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
 *  or compile errors. These reset when you switch cores - recheck them!)
 *
 *      Board            : ESP32S3 Dev Module
 *      ESP32 core       : 2.0.17
 *      PSRAM            : OPI PSRAM        <-- required, audio buffer lives there
 *      Flash Size       : 16MB (128Mb)
 *      Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
 *      USB CDC On Boot  : Disabled         <-- required, see USB note above
 *      Upload Speed     : 921600
 *      Port             : the CH340K USB-C port (the one near RESET)
 *
 *  LIBRARIES
 *      GFX Library for Arduino   v1.5.6   (NOT 1.6.x - that pairs with core 3)
 *      bb_captouch               v1.3.1
 *      ArduinoJson               v7.x
 *      Adafruit NeoPixel         any recent
 *
 *  ---------------------------------------------------------------------------
 *  FUNCTIONS IN THIS SKETCH
 *      led(r,g,b) + LED_* macros   RGB status colours (blue/amber/green/red)
 *      getTouch(&x,&y)      read the touch panel, mapped to screen coordinates
 *      speakButtonHeld()    true while a finger is on the SPEAK button
 *      bubbleLines(t)       how many lines a message wraps to
 *      drawOneBubble(m,y)   draw a single chat bubble
 *      redrawChat()         rebuild the chat area from history, newest at bottom
 *      clearChat()          wipe the chat history (CLEAR button)
 *      chatBubble(t,user)   add a message to history and redraw
 *      drawWifi()           WiFi signal bars + dBm readout
 *      drawBar(label,col)   bottom bar: SPEAK button + CLEAR + WiFi meter
 *      micInit()            I2S input - BOTH INMP441 mics, stereo
 *      spkInit()            I2S output - MAX98357 speaker
 *      recordWhileHeld()    record while SPEAK held, downmix stereo->mono
 *      writeWavHeader(...)  prepend the 44-byte RIFF/WAVE header
 *      dumpWavToSD(...)     save the exact upload to SD (/stt_debug.wav)
 *      readHttpResponse()   read an HTTPS reply, de-chunking it properly
 *      azureSTT(...)        chunked upload of the WAV -> recognised text
 *      deepseekChat(...)    question -> deepseek-v4-flash -> answer text
 *      azureTTSSpeak(text)  answer -> Azure voice -> PSRAM -> speaker
 *      setup() / loop()     boot + WiFi / one conversation turn per press
 *
 *  Robojax.com
 * ===========================================================================
 */

#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"

/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE     16000
#define RECORD_MAX_S    6                       // hard cap on one question
#define WAV_HEADER_LEN  44
#define REC_BUF_BYTES   (SAMPLE_RATE * RECORD_MAX_S * 2)   // 16-bit mono

/* Software gain applied to the recording. The INMP441 capture is quiet at
 * 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
 * speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
 * lower it. */
#define MIC_GAIN        6

/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
 * 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
 *     back silent or garbled in stereo mode. */
#define USE_BOTH_MICS   1

/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
 * uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS  25

/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
 * shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
    TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);

BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);

/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr;      // WAV_HEADER_LEN + up to REC_BUF_BYTES

/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;

/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool  ok_sd = false;
float g_mic_peak_pct = 0;          // last recording's raw peak, % of full scale
char  g_stt_err[64] = "";          // last STT failure cause, verbatim
char  g_llm_err[64] = "";          // last DeepSeek failure cause, verbatim

/* --- layout --------------------------------------------------------------- */
#define CHAT_H     200              // chat area: y 0..199
#define BAR_Y      202              // button bar below it
#define BTN_SPEAK_X   4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W  58
#define WIFI_X      236             // signal indicator, right end of the bar
#define BTN_H        36

/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
 * This is what prevents new text printing over old - the whole area is
 * rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int     chat_count = 0;

/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
  rgb.setPixelColor(0, rgb.Color(r, g, b));
  rgb.show();
}
#define LED_IDLE()    led(0, 0, 0)
#define LED_LISTEN()  led(0, 60, 255)     // blue   - recording
#define LED_THINK()   led(255, 120, 0)    // amber  - waiting on the cloud
#define LED_SPEAK()   led(0, 255, 40)     // green  - talking
#define LED_ERROR()   led(255, 0, 0)      // red


/* ===========================================================================
 *  Touch
 * =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
  TOUCHINFO ti;
  if (!bbct.getSamples(&ti)) return false;
  if (ti.count < 1) return false;
  *x = ti.y[0];
  *y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
  return true;
}

bool speakButtonHeld() {
  uint16_t x, y;
  if (!getTouch(&x, &y)) return false;
  return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}


/* ===========================================================================
 *  Chat UI  —  word-wrapped bubbles, user right/blue, assistant left/grey
 * =========================================================================== */
#define CHAT_CHARS 42                          // chars per line at textsize 1

static int bubbleLines(const char *t) {
  int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
  return l < 1 ? 1 : l;
}

void drawOneBubble(const ChatMsg &m, int y) {
  int len = strlen(m.text);
  int lines = bubbleLines(m.text);
  int h = lines * 10 + 8;
  uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
  int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
  if (w < 30) w = 30;
  int x = m.from_user ? (316 - w) : 4;

  gfx->fillRoundRect(x, y, w, h, 5, bg);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  for (int i = 0; i < lines; i++) {
    char line[CHAT_CHARS + 1] = {0};
    strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
    gfx->setCursor(x + 5, y + 5 + i * 10);
    gfx->print(line);
  }
}

/* Rebuild the whole chat area from history: newest message anchored at the
 * bottom, older ones stacked upward until the area is full. */
void redrawChat() {
  gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
  int shown = min(chat_count, CHAT_HISTORY);
  int y = CHAT_H - 2;
  for (int i = 0; i < shown; i++) {
    ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
    int h = bubbleLines(m.text) * 10 + 8;
    y -= h;
    if (y < 0) break;                          // area full - older ones drop off
    drawOneBubble(m, y);
    y -= 4;
  }
}

void clearChat() {
  chat_count = 0;
  redrawChat();
}

void chatBubble(const char *text, bool from_user) {
  ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
  strncpy(m.text, text, sizeof(m.text) - 1);
  m.text[sizeof(m.text) - 1] = 0;
  m.from_user = from_user;
  chat_count++;
  redrawChat();
}

/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
  gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
  bool up = (WiFi.status() == WL_CONNECTED);
  long rssi = up ? WiFi.RSSI() : -100;
  // -55 dBm or better = full bars; each 10 dB drops one
  int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;

  for (int b = 0; b < 4; b++) {
    int bh = 6 + b * 6;                       // heights 6,12,18,24
    uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
    gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
  }
  gfx->setTextSize(1);
  gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
  if (up) {
    gfx->setTextColor(CYAN);
    gfx->printf("%lddBm", rssi);
  } else {
    gfx->setTextColor(RED);
    gfx->print("DOWN");
  }
  gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
  gfx->setTextColor(gfx->color565(120, 120, 120));
  gfx->print("WiFi");
}

void drawBar(const char *label, uint16_t colour) {
  gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
  gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
  gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
  gfx->setTextSize(2);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
  gfx->print(label);

  // CLEAR wipes the chat history
  gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
  gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
  gfx->print("CLEAR");

  drawWifi();
}


/* ===========================================================================
 *  I2S  —  microphones on port 0, speaker on port 1. Separate hardware
 *  ports, so recording and playback can never fight over a bus.
 * =========================================================================== */
void micInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    /* BOTH channels - this is the two-microphone fix. The vendor examples
     * use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
     * falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
    .channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = false,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_MIC_SCK,
    .ws_io_num = I2S_MIC_WS,
    .data_out_num = I2S_PIN_NO_CHANGE,
    .data_in_num = I2S_MIC_SD
  };
  i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_MIC_PORT, &pins);
}

void spkInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,   // mono - the amp downmixes anyway
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = true,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_SPK_BCLK,
    .ws_io_num = I2S_SPK_LRC,
    .data_out_num = I2S_SPK_DOUT,
    .data_in_num = I2S_PIN_NO_CHANGE
  };
  i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_SPK_PORT, &pins);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
}


/* ===========================================================================
 *  Recording  —  runs while the SPEAK button is held (up to RECORD_MAX_S).
 *  Reads stereo pairs, averages L+R into one mono stream, applies a little
 *  software gain, and fills wav_buf after the 44-byte header slot.
 *  Returns the number of audio bytes recorded.
 * =========================================================================== */
size_t recordWhileHeld() {
  int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
  size_t   mono_samples = 0;
  const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;

  int16_t chunk[512];
  uint32_t last_touch_ok = millis();
  int32_t  peak = 0;                        // loudest raw sample - mic health check

  i2s_zero_dma_buffer(I2S_MIC_PORT);

  while (mono_samples < max_samples) {
    size_t got = 0;
    i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);

#if USE_BOTH_MICS
    size_t n = got / 4;                     // 4 bytes = one L+R pair
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
    size_t n = got / 2;                     // 2 bytes = one mono sample
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = chunk[i];
#endif
      if (abs(raw) > peak) peak = abs(raw);
      int32_t mixed = raw * MIC_GAIN;
      if (mixed >  32767) mixed =  32767;
      if (mixed < -32768) mixed = -32768;
      mono[mono_samples++] = (int16_t)mixed;
    }

    /* The GT911 is polled between I2S reads. A 250 ms grace period stops a
     * momentary missed touch sample from cutting the recording short. */
    if (speakButtonHeld()) last_touch_ok = millis();
    else if (millis() - last_touch_ok > 250) break;

    // live progress on the button
    static uint32_t last_draw = 0;
    if (millis() - last_draw > 200) {
      last_draw = millis();
      gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
                    (int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
    }
  }

  /* Mic health line: peak as % of full scale BEFORE gain.
   *   0%       = the mic is not being read at all (config/pin problem)
   *   under 3% = too quiet - speak closer or raise MIC_GAIN
   *   3-40%    = healthy speech level
   */
  g_mic_peak_pct = peak * 100.0 / 32768.0;
  Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
                g_mic_peak_pct, MIC_GAIN,
                peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");

  return mono_samples * 2;
}

/* Dump the exact WAV we are about to POST onto the SD card, so it can be
 * played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
  if (!ok_sd) return;
  SD.remove("/stt_debug.wav");
  File f = SD.open("/stt_debug.wav", FILE_WRITE);
  if (!f) return;
  f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
  f.close();
  Serial.println("debug copy saved to SD as /stt_debug.wav");
}

/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
  uint32_t file_len = data_bytes + 36;
  uint32_t byte_rate = SAMPLE_RATE * 2;
  memcpy(h, "RIFF", 4);           memcpy(h + 4, &file_len, 4);
  memcpy(h + 8, "WAVEfmt ", 8);
  uint32_t fmt_len = 16;          memcpy(h + 16, &fmt_len, 4);
  uint16_t fmt = 1, ch = 1;       memcpy(h + 20, &fmt, 2);   memcpy(h + 22, &ch, 2);
  uint32_t rate = SAMPLE_RATE;    memcpy(h + 24, &rate, 4);  memcpy(h + 28, &byte_rate, 4);
  uint16_t align = 2, bits = 16;  memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
  memcpy(h + 36, "data", 4);      memcpy(h + 40, &data_bytes, 4);
}


/* ===========================================================================
 *  HTTP response reader  —  shared by all three cloud calls.
 *  Returns the status code and fills body_out. Handles chunked transfer
 *  encoding PROPERLY: the chunk-size markers must be stripped, or they end
 *  up embedded inside the JSON body and the parse fails on long replies.
 * =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
 * on timeout / closed-and-empty. Every read below goes through this, because
 * a reasoning model can think for many seconds between the response headers
 * and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
  uint32_t t0 = millis();
  while (!c.available()) {
    if (!c.connected()) return false;
    if (millis() - t0 > ms) return false;
    delay(10);
  }
  return true;
}

static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
  body_out = "";

  if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);

  bool chunked = false;
  while (waitData(client, idle_ms)) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;         // blank line = end of headers
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
  }

  if (chunked) {
    int blanks = 0;
    while (true) {
      /* The chunk-size line may not arrive for a long time while the model
       * reasons. Waiting here - instead of letting read() time out - is the
       * whole fix: a timed-out read looks exactly like "0" (final chunk),
       * which silently truncated the body to nothing. */
      if (!waitData(client, idle_ms)) {
        Serial.println("HTTP: timed out waiting for the next chunk");
        break;
      }
      String szline = client.readStringUntil('\n');
      szline.trim();
      if (szline.length() == 0) {                    // stray blank line
        if (++blanks > 4) break;
        continue;
      }
      blanks = 0;

      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;                            // genuine final chunk

      long got = 0;
      while (got < sz) {
        if (!waitData(client, idle_ms)) break;
        while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
      }
      if (waitData(client, 3000)) client.readStringUntil('\n');   // CRLF after chunk
      if (got < sz) { Serial.println("HTTP: short chunk"); break; }
    }
  } else {
    while (waitData(client, idle_ms))
      while (client.available()) body_out += (char)client.read();
  }
  return code;
}


/* ===========================================================================
 *  CLOUD CALL 1  —  Azure speech-to-text
 *  One POST, one header, plain WAV body. This is why Azure does the ears.
 * =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
  /* HTTPClient's one-shot POST fails on bodies this large (it attempts one
   * giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
   * function speaks HTTP directly and streams the WAV up in 4 KB chunks -
   * reliable, and if it ever stalls we know the exact byte it stopped at. */
  WiFiClientSecure client;
  client.setInsecure();                    // no cert bundle on-device; see notes
  client.setTimeout(15);                   // seconds, for reads

  writeWavHeader(wav_buf, audio_bytes);
  dumpWavToSD(audio_bytes);                // PC-playable copy of what we send
  size_t total = WAV_HEADER_LEN + audio_bytes;

  if (!client.connect(AZURE_STT_HOST, 443)) {
    snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
    Serial.println("STT: TLS connect failed");
    return false;
  }

  /* Two valid host forms use DIFFERENT URL paths - detect which one is in
   * secrets.h:  <resource>.cognitiveservices.azure.com -> /stt/speech/...
   *             <region>.stt.speech.microsoft.com      -> /speech/...     */
  bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
  String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
               "/speech/recognition/conversation/cognitiveservices/v1"
               "?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
               "Host: " AZURE_STT_HOST "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
               "Accept: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(total) + "\r\n\r\n";
  client.print(req);

  /* body, 4 KB at a time */
  size_t sent = 0;
  while (sent < total) {
    size_t n = min((size_t)4096, total - sent);
    size_t w = client.write(wav_buf + sent, n);
    if (w == 0) {
      delay(50);                           // brief stall - retry once
      w = client.write(wav_buf + sent, n);
      if (w == 0) {
        snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
                 (unsigned)(sent / 1024));
        Serial.printf("STT: upload stalled at %u/%u bytes\n",
                      (unsigned)sent, (unsigned)total);
        client.stop();
        return false;
      }
    }
    sent += w;
    yield();
  }
  Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);

  /* read the reply with proper de-chunking */
  String resp;
  int code = readHttpResponse(client, resp, 10000);
  client.stop();

  if (code != 200) {
    snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
    Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
    Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
    return false;
  }

  const char *status = doc["RecognitionStatus"];
  if (!status || strcmp(status, "Success") != 0) {
    /* The status names the exact failure:
     *   InitialSilenceTimeout = Azure heard silence (mic level too low)
     *   NoMatch               = heard sound but no recognisable words
     *   BabbleTimeout         = heard only noise                       */
    snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
    Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
    return false;
  }
  text_out = doc["DisplayText"].as<String>();
  if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
  return text_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 2  —  DeepSeek chat completion
 *  OpenAI-compatible format. Model name is deepseek-v4-flash - the old
 *  deepseek-chat name is dead, see secrets.h.
 * =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
  /* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
   * bodies (long answers + the model's hidden reasoning), which shows up as
   * "bad JSON reply". Reading until the server closes the connection is
   * reliable regardless of reply length. */
  JsonDocument req;
  req["model"] = DEEPSEEK_MODEL;
  req["max_tokens"] = LLM_MAX_TOKENS;
  JsonArray msgs = req["messages"].to<JsonArray>();
  JsonObject sys = msgs.add<JsonObject>();
  sys["role"] = "system";  sys["content"] = SYSTEM_PROMPT;
  JsonObject usr = msgs.add<JsonObject>();
  usr["role"] = "user";    usr["content"] = question;

  String body;
  serializeJson(req, body);

  WiFiClientSecure client;
  client.setInsecure();
  /* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
   * ~2 s, but anything that needs actual working-out (an Ohm's law problem,
   * say) can think for 10-30 s before sending a single byte. */
  client.setTimeout(60);

  if (!client.connect(DEEPSEEK_HOST, 443)) {
    snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
    Serial.println("LLM: TLS connect failed");
    return false;
  }

  client.print(String("POST /chat/completions HTTP/1.1\r\n"
               "Host: " DEEPSEEK_HOST "\r\n"
               "Authorization: Bearer " DEEPSEEK_KEY "\r\n"
               "Content-Type: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: ") + String(body.length()) + "\r\n\r\n");
  client.print(body);

  /* read the reply with proper de-chunking; generous window - long
   * questions make the model think for a while before it responds */
  String resp;
  int code = readHttpResponse(client, resp, 60000);
  client.stop();

  if (code != 200) {
    snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
    Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
    Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
    return false;
  }

  const char *content = doc["choices"][0]["message"]["content"];
  const char *finish  = doc["choices"][0]["finish_reason"];
  if (!content || !content[0]) {
    /* v4-flash is a reasoning model: if finish_reason is "length", the whole
     * token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
    if (finish && strcmp(finish, "length") == 0)
      snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
    else
      snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
    Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
    return false;
  }
  answer_out = String(content);
  answer_out.trim();
  if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
  return answer_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 3  —  Azure text-to-speech, streamed straight to the speaker
 *  We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
 *  the I2S peripheral eats. Skip the 44-byte header, forward the rest.
 *  No MP3 decoder, no audio library, no buffering the whole reply.
 * =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
  size_t got = 0;
  uint32_t t0 = millis();
  while (got < n && millis() - t0 < 10000) {
    int r = c.read(dst + got, n - got);
    if (r > 0) { got += r; t0 = millis(); }
    else if (!c.connected() && !c.available()) break;
    else delay(2);
  }
  return got;
}

bool azureTTSSpeak(const String &text) {
  // Escape the XML special characters for the SSML body
  String safe = text;
  safe.replace("&", "&amp;");
  safe.replace("<", "&lt;");
  safe.replace(">", "&gt;");

  String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
                "<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";

  /* Manual HTTP like the other two cloud calls - and for a hard reason:
   * Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
   * raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
   * into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
   * Here we parse the framing properly and keep only clean audio bytes. */
  WiFiClientSecure client;
  client.setInsecure();
  client.setTimeout(20);

  const char *host = AZURE_REGION ".tts.speech.microsoft.com";
  if (!client.connect(host, 443)) {
    Serial.println("TTS: TLS connect failed");
    return false;
  }

  client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
               "Host: ") + host + "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: application/ssml+xml\r\n"
               "X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
               "User-Agent: MaTouchRobojax\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(ssml.length()) + "\r\n\r\n");
  client.print(ssml);

  /* status + headers; note whether the body is chunked */
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
  bool chunked = false;
  long content_len = -1;
  while (client.connected() || client.available()) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
    if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
  }
  if (code != 200) {
    Serial.printf("TTS HTTP %d\n", code);
    client.stop();
    return false;
  }

  const size_t AUDIO_CAP = 1200 * 1024;    // ~37 s of speech
  uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
  if (!audio) { client.stop(); return false; }
  size_t alen = 0;

  if (chunked) {
    /* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
    while (true) {
      String szline = client.readStringUntil('\n');
      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;
      if (alen + sz > AUDIO_CAP) break;
      size_t got = readExact(client, audio + alen, sz);
      alen += got;
      client.readStringUntil('\n');        // trailing CRLF after each chunk
      if (got < (size_t)sz) break;
    }
  } else if (content_len > 0) {
    alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
  } else {
    /* no framing info: read until the server closes */
    uint32_t idle = millis();
    while ((client.connected() || client.available()) && millis() - idle < 5000) {
      int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
      if (r > 0) { alen += r; idle = millis(); }
      else delay(5);
    }
  }
  client.stop();
  Serial.printf("TTS: %u KB clean audio (%s), playing\n",
                (unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");

  bool ok = (alen > WAV_HEADER_LEN);
  if (ok) {
    /* NOW the audio actually starts - this is the honest moment to go green */
    LED_SPEAK();
    drawBar("SPEAKING...", gfx->color565(0, 130, 40));

    /* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
    static const uint8_t lead_in[640] = {0};             // 20 ms of silence
    size_t w = 0;
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
  }
  free(audio);

  // let the DMA buffers drain so the last word is not cut off
  delay(150);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
  return ok;
}


/* ===========================================================================
 *  SETUP
 * =========================================================================== */
void setup() {
  Serial.begin(115200);
  delay(400);
  Serial.println("\n=== 04 Voice Assistant  |  Robojax.com ===");
  Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");

  pinMode(TFT_BLK, OUTPUT);
  digitalWrite(TFT_BLK, LOW);
  pinMode(SD_CS, OUTPUT);
  digitalWrite(SD_CS, HIGH);

  // one shared SPI bus for TFT + SD (started before either device)
  SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);

  gfx->begin();
  gfx->fillScreen(BLACK);
  digitalWrite(TFT_BLK, HIGH);

  // SD is optional here - it only stores the /stt_debug.wav diagnostic copy
  ok_sd = SD.begin(SD_CS, SPI, 20000000);
  digitalWrite(SD_CS, HIGH);
  Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
                       : "SD not found - debug WAV dump disabled (not fatal)");

  bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
  delay(50);

  rgb.begin();
  rgb.setBrightness(LED_BRIGHTNESS);
  LED_IDLE();

  /* One recording buffer for the whole session, in PSRAM. This is the 8 MB
   * that makes the board worth buying. */
  wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
  if (!wav_buf) {
    gfx->setTextColor(RED);
    gfx->setTextSize(2);
    gfx->setCursor(10, 100);
    gfx->print("PSRAM alloc failed!");
    gfx->setTextSize(1);
    gfx->setCursor(10, 130);
    gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
    while (1) delay(1000);
  }

  micInit();
  spkInit();

  gfx->setTextSize(1);
  gfx->setTextColor(YELLOW);
  gfx->setCursor(4, 4);
  gfx->printf("Connecting to %s ...", WIFI_SSID);
  Serial.printf("Connecting to %s ", WIFI_SSID);

  WiFi.mode(WIFI_STA);
  WiFi.begin(WIFI_SSID, WIFI_PASS);
  uint32_t t0 = millis();
  while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
    delay(300);
    Serial.print(".");
  }
  Serial.println();

  clearChat();
  if (WiFi.status() == WL_CONNECTED) {
    Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
    chatBubble("Hold SPEAK and ask me anything.", false);
  } else {
    chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
    LED_ERROR();
  }
  drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}


/* ===========================================================================
 *  LOOP  —  one full conversation turn per button press
 * =========================================================================== */
void loop() {
  /* CLEAR button: edge-detected so one tap wipes once. Reading the panel
   * twice per loop (here and in speakButtonHeld) is fine - the GT911 just
   * reports its current state. */
  static bool tap_latch = false;
  static uint8_t tap_release = 0;
  if (state == ST_IDLE) {
    uint16_t cx, cy;
    if (getTouch(&cx, &cy)) {
      tap_release = 0;
      if (!tap_latch) {
        tap_latch = true;
        if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
          clearChat();
          chatBubble("Hold SPEAK and ask me anything.", false);
        }
      }
    } else if (tap_latch && ++tap_release >= 4) {
      tap_latch = false;
      tap_release = 0;
    }

    /* live WiFi signal indicator, refreshed every 2 s while idle */
    static uint32_t last_wifi = 0;
    if (millis() - last_wifi > 2000) {
      last_wifi = millis();
      drawWifi();
    }
  }

  if (state == ST_IDLE && speakButtonHeld()) {

    /* ---- record ---- */
    state = ST_RECORDING;
    LED_LISTEN();
    drawBar("LISTENING...", gfx->color565(0, 60, 200));
    uint32_t t_rec = millis();
    size_t audio_bytes = recordWhileHeld();
    t_rec = millis() - t_rec;
    Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);

    if (audio_bytes < SAMPLE_RATE / 2) {   // under a quarter second - a tap
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }

    /* ---- speech to text ---- */
    state = ST_STT;
    LED_THINK();
    drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
    uint32_t t_stt = millis();
    String question;
    if (!azureSTT(audio_bytes, question)) {
      /* Show the REAL cause on screen - no serial monitor needed. */
      char diag[96];
      snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
               g_stt_err, g_mic_peak_pct,
               ok_sd ? " | saved /stt_debug.wav" : "");
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_stt = millis() - t_stt;
    chatBubble(question.c_str(), true);
    Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());

    /* ---- think ---- */
    state = ST_LLM;
    drawBar("THINKING...", gfx->color565(150, 90, 0));
    uint32_t t_llm = millis();
    String answer;
    if (!deepseekChat(question, answer)) {
      char diag[96];
      snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_llm = millis() - t_llm;
    chatBubble(answer.c_str(), false);
    Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());

    /* ---- speak ----
     * Still amber here: the voice has to be synthesised and downloaded first
     * (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
     * at the exact moment audio starts coming out of the speaker. */
    state = ST_TTS;
    LED_THINK();
    drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
    uint32_t t_tts = millis();
    bool spoke = azureTTSSpeak(answer);
    t_tts = millis() - t_tts;

    /* Timing summary on serial - this feeds the "honest numbers" segment. */
    Serial.printf("TIMINGS  rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
                  audio_bytes / 32000.0, (unsigned long)t_stt,
                  (unsigned long)t_llm, (unsigned long)t_tts,
                  spoke ? "" : "  (TTS FAILED)");

    drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
    LED_IDLE();
    state = ST_IDLE;
  }

  delay(20);
}

Faylasha📁

Faylka Loo Baahan Yahay (.h)

  • secrets.h
    file for Makerfabs MaTouch AI ESP32S3 2.8" TFT Camera module
    secrets.h 0.01 MB

Faylal Kale

  • pins.h
    pins file for MaTouch AI ESP32S3 2.8" camera LCD touch screen.
    pins.h 0.01 MB

Muuqaal

  • MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematic
    The latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
    MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF 0.15 MB