This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
The latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
دې ESP32-S3 (Azure + DeepSeek) باندې د AI غږیز مرستیال جوړ کړئ
یو تڼۍ ونیسئ، یوه پوښتنه وکړئ، او بورډ به په لوړ غږ ځواب ووایی
په یوه کوچني بورډ کې بشپړ غږیز معاون. د SPEAK تڼۍ ونیسئ او یوه پوښتنه وکړئ. بورډ تاسو د دواړو مایکروفونونو سره ثبتوي، آډیو Microsoft Azure ته لیږي ترڅو متن ته واړول شي، هغه متن DeepSeek ته لیږي ترڅو فکر وکړي، ځواب بیرته Azure ته لیږي ترڅو غږ ته واړول شي، او د خپل سپیکر له لارې یې غږوي. ټوله خبرې اترې په سکرین کې د چیټ بلبلونو په توګه ښکاري.
د MaTouch AI ESP32-S3 بورډ کې روان Talking AI غږیز معاون
له هغه شیبې څخه چې تاسو تڼۍ فشاروئ څه پیښیږي
دلته د یوې پوښتنې ټوله لاره ده، ګام په ګام. دا یو ځل لوستل ارزښت لري، ځکه چې هر هغه څه چې تاسو په سکرین او LED کې وینئ د دې مرحلو څخه په یوه پورې اړه لري.
تاسو د SPEAK تڼۍ فشاروئ او نیسئ. دا نیول-او-خبرې کول دي، نه ټکول-او-خبرې کول: ثبتونه دقیقا هومره دوام کوي څومره چې ستاسو ګوته ښکته پاتې کیږي، تر شپږو ثانیو پورې. د حالت LED نیلي کیږي او تڼۍ LISTENING ښیي.
دواړه مایکروفونونه تاسو ثبتوي. بورډ په هره ثانیه کې ۱۶،۰۰۰ ځله د سټیریو جوړې څخه نمونه اخلي، دوه چینلونه په یوه کې اوسط کوي، لږ ګټه پلي کوي، او پایله په PSRAM کې ذخیره کوي. لکه څنګه چې تاسو خبرې کوئ د پرمختګ بار د تڼۍ په اوږدو کې خپریږي. دوه ثانیې خبرې شاوخوا ۶۴ KB دي.
تاسو تڼۍ خوشې کوئ. ثبتونه ودریږي. بورډ د آډیو په مخ کې د ۴۴ بایټه WAV سرلیک لیکي - دا کوچنی لیبل هر هغه څه دي چې خام نمونې په داسې فایل بدلوي چې Azure به یې ومني.
آډیو Azure Speech-to-Text ته ځي. دا په ۴ KB ټوټو کې د خوندي اړیکې له لارې پورته کیږي، او د متن د یوې کرښې په توګه بیرته راځي. ستاسو ۶۴ KB غږ شاوخوا ۲۵ بایټه لیکل شوی متن شوی دی. LED کهري کیږي.
ستاسو پوښتنه په سکرین کې ښکاري د نیلي چیټ بلبل په توګه، نو تاسو کولی شئ په سمه توګه وګورئ چې دا څه اوریدلي دي - کوم چې ګټور دی، ځکه چې غلط اوریدل شوي کلمې ډیری عجیب ځوابونه تشریح کوي.
متن DeepSeek ته ځي. بورډ ستاسو پوښتنه د یوې تلپاتې لارښوونې سره لیږي چې ځوابونه په دوو لنډو جملو کې وساتي. ماډل فکر کوي - واقعیا فکر کوي، دا د استدلال ماډل دی - او یو ځواب بیرته راګرځوي.
ځواب په سکرین کې ښکاري د خړ بلبل په توګه. تاسو کولی شئ دا مخکې له دې چې واورئ ولولئ.
ځواب بیرته Azure ته ځي ترڅو وویل شي. بورډ د خام ۱۶ kHz PCM آډیو غوښتنه کوي، کوم چې دقیقا هغه بڼه ده چې د هغه امپلیفایر غواړي، نو په دې پروژه کې هیڅ MP3 ډیکوډر شتون نلري. تڼۍ اوس GETTING VOICE لوستل کیږي او LED کهري پاتې کیږي، ځکه چې تر اوسه هیڅ د اوریدو وړ نه دی.
ټوله کلیپ د لوبولو دمخه PSRAM ته ډاونلوډ کیږي. دا مهم دی - لاندې یادونه وګورئ.
بیا غږول. هغه شیبه چې آډیو سپیکر ته سپارل کیږي، LED شین کیږي او تڼۍ SPEAKING لوستل کیږي. تاسو ځواب اورئ.
ولې آډیو لومړی ډاونلوډ کیږي پرته له دې چې لکه څنګه چې رسیږي وغږول شي. مستقیم له شبکې څخه سپیکر ته جریان کول د ټکولو په څیر غږیږي. د سپیکر بفر یوازې شاوخوا لسمه ثانیه ساتي، او د WiFi لیږد کې هر هغه وقفه چې له دې څخه اوږده وي هغه خالي کوي، چې د اوریدو وړ ټک تولیدوي. لومړی ټول ځواب PSRAM ته ډاونلوډ کول شاوخوا یوه ثانیه اضافي انتظار لګوي او هر تشه لرې کوي. له همدې امله ښودنه د SPEAKING ویلو دمخه GETTING VOICE وایی - دواړه په ریښتیا جلا مرحلې دي.
هره مرحله څومره وخت نیسي
په ریښتیني هارډویر کې اندازه شوی، د یوې ساده پوښتنې لپاره:
مرحله | عادي وخت |
|---|---|
ثبتونه | تر هغه چې تاسو تڼۍ ونیسئ |
غږ-ته-متن (Azure) | شاوخوا ۱.۸ ثانیې |
فکر کول (DeepSeek) | د ساده پوښتنې لپاره شاوخوا ۱.۸ ثانیې، د هغه لپاره چې ریښتیني محاسبې ته اړتیا لري ډیر اوږد |
د غږ راوړل (Azure) | شاوخوا ۷ ثانیې - ترټولو لویه برخه |
ټولټال، له خوشې کولو څخه تر لومړي غږ پورې | نږدې ۱۱ ثانیې |
هره خبرې اترې خپل وختونه سریال مانیټر ته چاپوي، نو تاسو کولی شئ خپل ځان اندازه کړئ پرته له دې چې دې باور وکړئ. که تاسو دا ګړندی غواړئ، ترټولو اغیزمن بدلون په SYSTEM_PROMPT کې د لنډو ځوابونو غوښتنه کول دي - د ویلو لپاره لږ متن د ترکیب او ډاونلوډ لپاره لږ آډیو معنی لري.
بورډ پخپله هیڅکله هیڅ نه پوهیږي. دا یو پیغام رسونکی دی چې ښه غوږونه او ښه غږ لري - استخبارات په ثانیه کې کرایه کیږي.
ولې دا درې خدمتونه
Azure دننه او بهر غږ اداره کوي. د هغه متن-ته-غږ کولی شي خام ۱۶ kHz PCM بیرته راولي، کوم چې دقیقا هغه څه دي چې د سپیکر چپ غواړي، نو په دې پروژه کې هیڅ MP3 ډیکوډر شتون نلري. د هغه غږ-ته-متن یو ساده WAV په یو ساده POST کې اخلي.
DeepSeek د خبرو اترو دماغ دی. دا ګړندی دی او د هر ځواب لپاره یوازې د یو سینټ یوه برخه لګښت لري.
OpenAI دلته نه کارول کیږي - پروژه ۰۵ وګورئ، چیرې چې دا د لید کار کوي.
د DeepSeek ماډل نومونه بدل شوي. زاړه deepseek-chat او deepseek-reasoner نومونه د جولای ۲۰۲۶ کې تقاعد شوي. ډیری آنلاین ټیوټوریلونه لاهم یې کاروي او تېروتنه به بیرته راولي. اوسني نومونه deepseek-v4-flash او deepseek-v4-pro دي. دا پروژه v4-flash کاروي.
د استدلال-ماډل جال
DeepSeek v4-flash د ځواب ویلو دمخه فکر کوي، او هغه فکر ستاسو د ټوکن حد ته شمېرل کیږي. که LLM_MAX_TOKENS ډېر ټیټ وټاکئ، ټوله بودیجه په استدلال مصرفیږي، ځواب تش راګرځي، او بورډ هېڅ نه وايي. دلته د همدې دلیل لپاره 400 ټاکل شوی دی. سختې پوښتنې هم ډېر وخت نیسي - یوه ساده حقیقت شاوخوا دوه ثانیو کې ځواب ورکوي، یوه پوښتنه چې واقعي محاسبې ته اړتیا لري کولی شي ډېر ډیر وخت ونیسي.
د حالت د څراغ لوستل
رنګ | مانا |
|---|---|
شين | تاسو ته غوږ نیسي |
کهري | کلاوډ فکر کوي، یا غږ راوړل کیږي |
زرغون | خبرې کوي - دا په هغه دقیقه شیبه کې زرغون کیږي چې غږ پیل شي |
سور | یو څه ناکام شو - سریال مانیټر وګورئ |
د پردې کنټرولونه
چټ د تلیفون د خبرو اترو په څیر سکرول کیږي، زاړه پیغامونه پورته خوځي او ورک کیږي. CLEAR دا پاکوي. د WiFi سیګنال متر چې اصلي dBm لوستل لري په کونج کې ناست دی، کوم چې ګټور دی کله چې تاسو حیران یاست چې ایا ورو ځواب شبکه ده یا خدمت.
د MaTouch AI ESP32-S3 2.8" بورډ په اړه
په دې پاڼه کې هره پروژه په MaTouch AI ESP32-S3 2.8" TFT ST7789V له Makerfabs څخه چلیږي. دا یو ټول-په-یو بورډ دی: یو رنګین ټ�چ سکرین، یو 3 میګاپکسله کامره، دوه مایکروفونونه او یو ریښتینی سپیکر امپلیفیر، ټول د ESP32-S3 لخوا د 8 MB PSRAM سره پرمخ وړل کیږي. دا ترکیب دی چې دا AI پروژې په یوه واحده بورډ کې د بل هیڅ شی سره نه نښلول ممکن کوي.
8 MB PSRAM دلته د هر بل شمیر څخه ډیر مهم دی. دا هغه څه دي چې بورډ ته اجازه ورکوي چې د کامرې یوه چوکاټ، د ثبت شوي غږ څو ثانیې، یا یو base64-کوډ شوی عکس په حافظه کې په ورته وخت کې وساتي - چې هیڅ یو یې د ESP32 په عادي RAM کې نه ځای کیږي.
د تولیدونکي اسناد: د Makerfabs ویکي پاڼه.
کلیدي مشخصات
پروسسر: ESP32-S3، دوه کور 240 MHz، WiFi 2.4 GHz + بلوتوث 5.0
حافظه: 16 MB فلش، 8 MB PSRAM (دلته د نږدې هرې پروژې لپاره اړین)
ښکاره: 2.8" IPS، 320×240، ST7789V ډرایور، SPI
ټچ: GT911 کپیسیټیو، په یو وخت کې 5 ګوتې تعقیبوي
کامره: OV3660، 3 میګاپکسله، تر 2048×1536 پورې
مایکروفونونه: دوه INMP441 I2S ډیجیټل مایکونه (یو ریښتینی سټیریو جوړه)
سپیکر: MAX98357A کلاس-D امپلیفیر، 3.2 W په 4 Ω کې
ذخیره: د microSD کارت سلاټ (SPI حالت)
بریښنا: USB-C، د JST بیټرۍ نښلونکی، TP4056 چارجر، د بریښنا سویچ
همدارنګه په بورډ کې: WS2812B RGB LED، PCF8563T د بیټرۍ ملاتړ شوی ریښتینی وخت ساعت، او د MAX17048 بیټرۍ د سونګ ګیج چې په رسمي مشخصاتو کې نه لیست شوی
دوه USB-C بندرونه یو شان ندي. د بورډ سپیکر خپل سیګنال پنونه (IO19 او IO20) د اصلي USB بندر سره شریکوي، ځکه چې دا پنونه د ESP32-S3 هارډوایر شوي USB ډیټا کرښې دي. تل د CH340K USB-C بندر له لارې اپلوډ او بریښنا ورکړئ (هغه چې د RESET تڼۍ تر څنګ دی)، او USB CDC On Boot په Disabled وټاکئ. غلط بندر وکاروئ او غږ به خراب چلند وکړي یا اپلوډونه به ناکام شي.
د Arduino IDE ترتیبات
دا ترتیبات مهم دي. ډیری ستونزې چې خلک یې د دې بورډ سره راپور ورکوي د دې څخه یوه غلطه وي، او دوی کله چې تاسو د کور نسخه بدله کړئ بیا تنظیمیږي، نو د هر بدلون وروسته یې بیا وګورئ.
ترتیب | ارزښت |
|---|---|
بورډ | ESP32S3 Dev Module |
د ESP32 کور نسخه | 2.0.17 |
PSRAM | OPI PSRAM |
د فلش اندازه | 16MB (128Mb) |
د پارټیشن سکیم | 16M Flash (3MB APP/9.9MB FATFS) |
USB CDC On Boot | Disabled |
د اپلوډ سرعت | 921600 |
د اپلوډ دمخه ټول فلش پاک کړئ | Disabled |
بندر | د CH340K USB-C بندر |
د ESP32 کور 2.0.17 وکاروئ، نه 3.x. Espressif په کور 3 کې د وسیلې پر مخ د مخ کشف ماډلونه لرې کړل، نو د مخ پروژې به هلته کمپایل نشي. د 2.0.17 پن کول په دې پاڼه کې هره پروژه د یو ترتیب سره کار کوي. په Boards Manager کې، د نسخې ډراپ-ډاون تاسو ته اجازه درکوي چې هر وخت چې وغواړئ شا او خوا تبدیلي وکړئ.
د GFX Library for Arduino نسخه 1.5.6 وکاروئ، نه 1.6.x. د 1.6 خپرونې د ESP32 کور 3 لپاره جوړې شوې دي او کولی شي په کور 2.0.17 کې په پیل کې ځړول شي. که ستاسو سکرین د اپلوډ وروسته تور پاتې شي، دا لومړی شی دی چې وګورئ.
اړین کتابتونونه
دا د Arduino IDE کې د Tools → Manage Libraries له لارې نصب کړئ. د نسخې شمیرې مهمې دي - مهرباني وکړئ هغه وکاروئ چې لیست شوي دي.
کتابتون | نسخه | لیکوال |
|---|---|---|
GFX Library for Arduino | 1.5.6 | moononournation |
bb_captouch | 1.3.1 | Larry Bank |
ArduinoJson | 7.x | Benoit Blanchon |
Adafruit NeoPixel | هر وروستی | Adafruit |
د secrets.h تنظیمول
ستاسو د WiFi تفصیلات او هر API کیلي په secrets.h کې ځي، کوم چې د ځای نیونکو ارزښتونو سره په ډاونلوډ کې شامل دی. په Arduino IDE کې هغه ټب خلاص کړئ او دوی د خپلو سره بدل کړئ.
WiFi باید 2.4 GHz وي. ESP32-S3 کولی شي د 5 GHz شبکه په هیڅ صورت ونه ګوري. که ستاسو روټر دواړه بانډونه د یو نوم لاندې یوځای کوي (Asus دا Smart Connect بولي)، یا دا بند کړئ یا د 2.4 GHz بانډ ته خپل جلا نوم ورکړئ او په secrets.h کې هغه وکاروئ.
ستاسو د API کیلي ترلاسه کول
دا پروژه د کلاوډ AI خدمت سره خبرې کوي، نو تاسو خپله کیلي ته اړتیا لرئ. که تاسو مخکې دا هیڅکله نه وي کړي، اندیښنه مه کوئ - دا د هغه پاسورډ په څیر ورته نظر دی چې ستاسو حساب خدمت ته پیژني. دا یو څو دقیقې وخت نیسي، یو ځل.
یوه کیلي د ویب پاڼې ګډون نه دی. د مثال په توګه، د ChatGPT Plus لپاره پیسې ورکول تاسو ته API کیلي نه درکوي - دواړه جلا محصولات دي چې جلا بیل لري. تاسو د پراختیا کونکي پلیټ فارم کې حساب ته اړتیا لرئ، لکه څنګه چې لاندې تشریح شوی.
Microsoft Azure Speech - د اوریدو او خبرو کولو لپاره
Azure ستاسو خبرې متن ته اړوي او ځواب بیرته غږ ته اړوي. وړیا کچه د دې پاڼې د هرڅه لپاره کافي ده.
portal.azure.com ته لاړ شئ او د Microsoft حساب سره ننوځئ (وړیا حساب سم دی).
که تاسو مخکې Azure نه وي کارولی، تاسو به د Welcome to Azure سکرین وګورئ چې درې انتخابونه وړاندې کوي. Start with an Azure free trial غوره کړئ - تاسو ګډون ته اړتیا لرئ مخکې لدې چې Azure تاسو ته اجازه درکړي کوم څه جوړ کړئ. (زده کونکي باید پرځای Azure for Students غوره کړي: ورته پایله، کارت ته اړتیا نشته.) Manage Microsoft Entra ID له پامه غورځوئ، کوم چې په بشپړ ډول بل څه دی.
د Create a resource تڼۍ کلیک وکړئ، د Speech لپاره لټون وکړئ، او د Microsoft لخوا خپور شوی Speech service غوره کړئ.
فورمه ډکه کړئ: هر resource group، هر نوم، او ستاسو ته نږدې Region غوره کړئ - هغه سیمه په سمه توګه ولیکئ لکه څنګه چې ښکاري، د مثال په توګه
eastus.د Pricing tier لپاره F0 (Free) غوره کړئ. دا په میاشت کې شاوخوا پنځه ساعته د خبرو څخه متن ته او نیم ملیون حروف د متن څخه خبرو ته اجازه ورکوي.
د Review + create تڼۍ کلیک وکړئ، بیا Create. شاوخوا یوه دقیقه انتظار وکړئ، بیا د Go to resource تڼۍ کلیک وکړئ.
په کیڼ اړخ مینو کې Keys and Endpoint خلاص کړئ. KEY 1 او Location/Region کاپي کړئ.
هغه په secrets.h کې د AZURE_SPEECH_KEY او AZURE_REGION په توګه واچوئ. د AZURE_STT_HOST لپاره، <region>.stt.speech.microsoft.com وکاروئ - نو د eastus سیمې سره دا eastus.stt.speech.microsoft.com دی.
د کریډیټ کارت په اړه. د Azure وړیا آزموینه ستاسو د هویت تصدیق کولو لپاره کارت غوښتنه کوي. دا تاسو ته پیسې نه اخلي. تاسو د 30 ورځو لپاره $200 کریډیټ ترلاسه کوئ، او وروسته حساب Pay-As-You-Go ته ځي - مګر F0 Speech کچه وړیا پاتې کیږي، میاشت په میاشت، او په دې پروژو کې هرڅه په آرامۍ سره په کې ځای لري. که تاسو غواړئ په هیڅ صورت کارت ورنکړئ او تاسو زده کونکی یاست، د Azure for Students اختیار تاسو ته پرته له کارت څخه کریډیټ درکوي.
دا باید د "Speech service" سرچینه وي. د Translator، Language، یا عمومي Cognitive Services سرچینې کیلي یو شان ښکاري او په بشپړ ډول معتبره ده - مګر هر د خبرو غوښتنه 401 تېروتنه بیرته راګرځوي. دا موږ د ازموینې پرمهال ونیول او یو ساعت یې مصرف کړ. که خبرې د 401 سره ناکامې شي پداسې حال کې چې کیلي سمه ښکاري، وګورئ چې تاسو کوم ډول سرچینه جوړه کړې ده.
DeepSeek - د فکر کولو برخه
DeepSeek هغه ژبنی ماډل دی چې په حقیقت کې ستاسو پوښتنې ته ځواب ورکوي. دا ارزانه دی - د څو ډالرو کریډیټ زرګونه ځوابونه پوښي.
platform.deepseek.com ته لاړ شئ او حساب جوړ کړئ.
په مینو کې API keys خلاص کړئ او د Create new API key تڼۍ کلیک وکړئ.
سمدستي یې کاپي کړئ. دا یو ځل ښودل کیږي او بیا هیڅکله نه - که تاسو یې له لاسه ورکړئ، هغه کیلي ړنګه کړئ او بله جوړه کړئ.
د Top up لاندې لږ مقدار کریډیټ اضافه کړئ. وړیا کچه نشته، مګر ترټولو کوچنی اضافه په دې کارونې کې ډیر اوږد دوام کوي.
کیلي په secrets.h کې د DEEPSEEK_KEY په توګه واچوئ. دا د sk- سره پیل کیږي.
د ماډل نومونه د جولای 2026 کې بدل شوي. زاړه deepseek-chat او deepseek-reasoner تقاعد شوي، نو ډیری ټیوټوریلونه چې تاسو آنلاین ومومئ د 400 تېروتنې سره ناکام به شي. deepseek-v4-flash وکاروئ، کوم چې دا پروژې دمخه ټاکلی دی.
د چلولو لګښت څومره دی
ډیر لږ، مګر دا وړیا نه دی، او تاسو باید تقریبا پوه شئ چې تاسو څه مصرف کوئ مخکې لدې چې پروژه روانه پریږدئ.
خدمت | تقریبي لګښت |
|---|---|
Azure Speech | وړیا کچه په میاشت کې شاوخوا 5 ساعته اوریدل او 0.5 M حروف خبرې پوښي |
DeepSeek | د هر ځواب لپاره د یو سینټ یوه برخه - د څو ډالرو لپاره زرګونه ځوابونه |
OpenAI vision | د هر انځور لپاره تقریبا یو یا دوه سینټ، د ماډل پورې اړه لري |
بیې بدلیږي، نو دا د نرخ پرځای د لارښود په توګه وګورئ. د دې هر خدمت د کارونې پاڼه لري چیرې چې تاسو کولی شئ وګورئ چې څه مصرف کړي دي، او ټول تاسو ته اجازه درکوي د مصرف حد وټاکئ - کوم چې په لومړۍ ورځ کولو ارزښت لري.
خپل کليډونه شخصي وساتئ. هر څوک چې دا ولري کولای شي ستاسو پیسې مصرف کړي. دوی په ویډیو، سکرین شاټ، فورم پوسټ یا عامه کوډ ذخیره کې مه اچوئ. که کوم کليډ ښکاره شي، د چمتو کونکي په ویب پاڼه کې یې ړنګ کړئ او نوی جوړ کړئ - دا یوازې څو ثانیې وخت نیسي، او دا یوازینی ریښتینی حل دی.
ستونزې حل کول
نښه | علت او حل |
|---|---|
سکرین تور پاتې کیږي | د GFX کتابتون غلطه نسخه (1.5.6 وکاروئ) یا د بورډ غلط ترتیبات. |
|
|
هیڅ نه پورته کیږي / د COM بندر نشته | غلط USB-C بندر، یا د CH340 ډرایور نصب شوی نه دی. |
کیمره ناکامه کیږي او بیرته نه راګرځي | د کیمرې ریسیټ کرښه د بورډ د RESET تڼۍ سره تړلې ده، نو سافټویر نشي کولای دا بیا پیل کړي. RESET فشار کړئ. که بیا هم ناکامه شي، د کیمرې د ربن کیبل بیرته ځای په ځای کړئ. |
کوډ ډاونلوډ کړئ
د دې پروژې بشپړ Arduino سکیچ، د pins.h او نورو اړینو شیانو سره، په وړیا توګه ډاونلوډ لپاره شتون لري.
04_Voice_Assistant ډاونلوډ کړئ
دا انزیپ کړئ، د .ino فایل په Arduino IDE کې خلاص کړئ، پورته ترتیبات وګورئ، او د CH340K USB-C بندر له لارې پورته یې کړئ.
This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
/*
* ===========================================================================
* 04_Voice_Assistant — MaTouch AI ESP32-S3 2.8" TFT ST7789V
* ===========================================================================
*
----------
* ROBOJAX.COM - MaTouch AI ESP32-S3 2.8" project series
*
* WATCH THE VIDEO
* https://youtu.be/6AL3g3tC_Hk
*
* WRITTEN TUTORIALS - every project, with photos and full explanation
* Camera and touchscreen.... https://robojax.com/RTJ849
* Offline face recognition.. https://robojax.com/RTJ850
* AI voice assistant........ https://robojax.com/RTJ851
* AI vision................. https://robojax.com/RTJ852
*
* GET THE BOARD - SAVE $5 with coupon code: Robojax_Makerfab
* https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
* (enter the code at checkout)
*
* All of this code is free. If it helped you, a subscribe on YouTube is
* the best way to support more of it.
*
* ---------------------------------------------------------------------------
*
* A complete voice assistant on a $40 board:
*
* hold SPEAK -> both INMP441 microphones record you
* -> Azure Speech turns the audio into text
* -> DeepSeek v4-flash thinks of an answer
* -> Azure Speech turns the answer into audio
* -> the MAX98357 speaker says it out loud
*
* and the whole conversation is drawn as chat bubbles on the touchscreen.
*
* WHY THIS COMBINATION OF SERVICES (each is used where it is best):
* - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
* transcription endpoint wants multipart/form-data - miserable on an MCU.
* - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
* which streams straight into the I2S speaker with NO MP3 decoder at all.
* - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
* model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
* Most tutorials online still use them and are broken. See secrets.h.
*
* A detail the vendor examples get wrong: this board has TWO microphones on
* one I2S bus (left + right), but every Makerfabs demo records left-only and
* throws one away. This sketch records both and averages them.
*
* ---------------------------------------------------------------------------
* *** WHICH USB PORT - THIS MATTERS ***
* The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
* through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
* If you use the wrong port the audio will be garbage or uploads will fail.
* ---------------------------------------------------------------------------
*
* FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
*
* BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
* or compile errors. These reset when you switch cores - recheck them!)
*
* Board : ESP32S3 Dev Module
* ESP32 core : 2.0.17
* PSRAM : OPI PSRAM <-- required, audio buffer lives there
* Flash Size : 16MB (128Mb)
* Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
* USB CDC On Boot : Disabled <-- required, see USB note above
* Upload Speed : 921600
* Port : the CH340K USB-C port (the one near RESET)
*
* LIBRARIES
* GFX Library for Arduino v1.5.6 (NOT 1.6.x - that pairs with core 3)
* bb_captouch v1.3.1
* ArduinoJson v7.x
* Adafruit NeoPixel any recent
*
* ---------------------------------------------------------------------------
* FUNCTIONS IN THIS SKETCH
* led(r,g,b) + LED_* macros RGB status colours (blue/amber/green/red)
* getTouch(&x,&y) read the touch panel, mapped to screen coordinates
* speakButtonHeld() true while a finger is on the SPEAK button
* bubbleLines(t) how many lines a message wraps to
* drawOneBubble(m,y) draw a single chat bubble
* redrawChat() rebuild the chat area from history, newest at bottom
* clearChat() wipe the chat history (CLEAR button)
* chatBubble(t,user) add a message to history and redraw
* drawWifi() WiFi signal bars + dBm readout
* drawBar(label,col) bottom bar: SPEAK button + CLEAR + WiFi meter
* micInit() I2S input - BOTH INMP441 mics, stereo
* spkInit() I2S output - MAX98357 speaker
* recordWhileHeld() record while SPEAK held, downmix stereo->mono
* writeWavHeader(...) prepend the 44-byte RIFF/WAVE header
* dumpWavToSD(...) save the exact upload to SD (/stt_debug.wav)
* readHttpResponse() read an HTTPS reply, de-chunking it properly
* azureSTT(...) chunked upload of the WAV -> recognised text
* deepseekChat(...) question -> deepseek-v4-flash -> answer text
* azureTTSSpeak(text) answer -> Azure voice -> PSRAM -> speaker
* setup() / loop() boot + WiFi / one conversation turn per press
*
* Robojax.com
* ===========================================================================
*/
#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"
/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE 16000
#define RECORD_MAX_S 6 // hard cap on one question
#define WAV_HEADER_LEN 44
#define REC_BUF_BYTES (SAMPLE_RATE * RECORD_MAX_S * 2) // 16-bit mono
/* Software gain applied to the recording. The INMP441 capture is quiet at
* 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
* speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
* lower it. */
#define MIC_GAIN 6
/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
* 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
* back silent or garbled in stereo mode. */
#define USE_BOTH_MICS 1
/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
* uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS 25
/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
* shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);
BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);
/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr; // WAV_HEADER_LEN + up to REC_BUF_BYTES
/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;
/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool ok_sd = false;
float g_mic_peak_pct = 0; // last recording's raw peak, % of full scale
char g_stt_err[64] = ""; // last STT failure cause, verbatim
char g_llm_err[64] = ""; // last DeepSeek failure cause, verbatim
/* --- layout --------------------------------------------------------------- */
#define CHAT_H 200 // chat area: y 0..199
#define BAR_Y 202 // button bar below it
#define BTN_SPEAK_X 4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W 58
#define WIFI_X 236 // signal indicator, right end of the bar
#define BTN_H 36
/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
* This is what prevents new text printing over old - the whole area is
* rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int chat_count = 0;
/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
rgb.setPixelColor(0, rgb.Color(r, g, b));
rgb.show();
}
#define LED_IDLE() led(0, 0, 0)
#define LED_LISTEN() led(0, 60, 255) // blue - recording
#define LED_THINK() led(255, 120, 0) // amber - waiting on the cloud
#define LED_SPEAK() led(0, 255, 40) // green - talking
#define LED_ERROR() led(255, 0, 0) // red
/* ===========================================================================
* Touch
* =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
TOUCHINFO ti;
if (!bbct.getSamples(&ti)) return false;
if (ti.count < 1) return false;
*x = ti.y[0];
*y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
return true;
}
bool speakButtonHeld() {
uint16_t x, y;
if (!getTouch(&x, &y)) return false;
return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}
/* ===========================================================================
* Chat UI — word-wrapped bubbles, user right/blue, assistant left/grey
* =========================================================================== */
#define CHAT_CHARS 42 // chars per line at textsize 1
static int bubbleLines(const char *t) {
int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
return l < 1 ? 1 : l;
}
void drawOneBubble(const ChatMsg &m, int y) {
int len = strlen(m.text);
int lines = bubbleLines(m.text);
int h = lines * 10 + 8;
uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
if (w < 30) w = 30;
int x = m.from_user ? (316 - w) : 4;
gfx->fillRoundRect(x, y, w, h, 5, bg);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
for (int i = 0; i < lines; i++) {
char line[CHAT_CHARS + 1] = {0};
strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
gfx->setCursor(x + 5, y + 5 + i * 10);
gfx->print(line);
}
}
/* Rebuild the whole chat area from history: newest message anchored at the
* bottom, older ones stacked upward until the area is full. */
void redrawChat() {
gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
int shown = min(chat_count, CHAT_HISTORY);
int y = CHAT_H - 2;
for (int i = 0; i < shown; i++) {
ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
int h = bubbleLines(m.text) * 10 + 8;
y -= h;
if (y < 0) break; // area full - older ones drop off
drawOneBubble(m, y);
y -= 4;
}
}
void clearChat() {
chat_count = 0;
redrawChat();
}
void chatBubble(const char *text, bool from_user) {
ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
strncpy(m.text, text, sizeof(m.text) - 1);
m.text[sizeof(m.text) - 1] = 0;
m.from_user = from_user;
chat_count++;
redrawChat();
}
/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
bool up = (WiFi.status() == WL_CONNECTED);
long rssi = up ? WiFi.RSSI() : -100;
// -55 dBm or better = full bars; each 10 dB drops one
int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;
for (int b = 0; b < 4; b++) {
int bh = 6 + b * 6; // heights 6,12,18,24
uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
}
gfx->setTextSize(1);
gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
if (up) {
gfx->setTextColor(CYAN);
gfx->printf("%lddBm", rssi);
} else {
gfx->setTextColor(RED);
gfx->print("DOWN");
}
gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
gfx->setTextColor(gfx->color565(120, 120, 120));
gfx->print("WiFi");
}
void drawBar(const char *label, uint16_t colour) {
gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
gfx->setTextSize(2);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
gfx->print(label);
// CLEAR wipes the chat history
gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
gfx->print("CLEAR");
drawWifi();
}
/* ===========================================================================
* I2S — microphones on port 0, speaker on port 1. Separate hardware
* ports, so recording and playback can never fight over a bus.
* =========================================================================== */
void micInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
/* BOTH channels - this is the two-microphone fix. The vendor examples
* use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
* falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
.channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = false,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_MIC_SCK,
.ws_io_num = I2S_MIC_WS,
.data_out_num = I2S_PIN_NO_CHANGE,
.data_in_num = I2S_MIC_SD
};
i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_MIC_PORT, &pins);
}
void spkInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT, // mono - the amp downmixes anyway
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = true,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_SPK_BCLK,
.ws_io_num = I2S_SPK_LRC,
.data_out_num = I2S_SPK_DOUT,
.data_in_num = I2S_PIN_NO_CHANGE
};
i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_SPK_PORT, &pins);
i2s_zero_dma_buffer(I2S_SPK_PORT);
}
/* ===========================================================================
* Recording — runs while the SPEAK button is held (up to RECORD_MAX_S).
* Reads stereo pairs, averages L+R into one mono stream, applies a little
* software gain, and fills wav_buf after the 44-byte header slot.
* Returns the number of audio bytes recorded.
* =========================================================================== */
size_t recordWhileHeld() {
int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
size_t mono_samples = 0;
const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;
int16_t chunk[512];
uint32_t last_touch_ok = millis();
int32_t peak = 0; // loudest raw sample - mic health check
i2s_zero_dma_buffer(I2S_MIC_PORT);
while (mono_samples < max_samples) {
size_t got = 0;
i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);
#if USE_BOTH_MICS
size_t n = got / 4; // 4 bytes = one L+R pair
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
size_t n = got / 2; // 2 bytes = one mono sample
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = chunk[i];
#endif
if (abs(raw) > peak) peak = abs(raw);
int32_t mixed = raw * MIC_GAIN;
if (mixed > 32767) mixed = 32767;
if (mixed < -32768) mixed = -32768;
mono[mono_samples++] = (int16_t)mixed;
}
/* The GT911 is polled between I2S reads. A 250 ms grace period stops a
* momentary missed touch sample from cutting the recording short. */
if (speakButtonHeld()) last_touch_ok = millis();
else if (millis() - last_touch_ok > 250) break;
// live progress on the button
static uint32_t last_draw = 0;
if (millis() - last_draw > 200) {
last_draw = millis();
gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
(int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
}
}
/* Mic health line: peak as % of full scale BEFORE gain.
* 0% = the mic is not being read at all (config/pin problem)
* under 3% = too quiet - speak closer or raise MIC_GAIN
* 3-40% = healthy speech level
*/
g_mic_peak_pct = peak * 100.0 / 32768.0;
Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
g_mic_peak_pct, MIC_GAIN,
peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");
return mono_samples * 2;
}
/* Dump the exact WAV we are about to POST onto the SD card, so it can be
* played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
if (!ok_sd) return;
SD.remove("/stt_debug.wav");
File f = SD.open("/stt_debug.wav", FILE_WRITE);
if (!f) return;
f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
f.close();
Serial.println("debug copy saved to SD as /stt_debug.wav");
}
/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
uint32_t file_len = data_bytes + 36;
uint32_t byte_rate = SAMPLE_RATE * 2;
memcpy(h, "RIFF", 4); memcpy(h + 4, &file_len, 4);
memcpy(h + 8, "WAVEfmt ", 8);
uint32_t fmt_len = 16; memcpy(h + 16, &fmt_len, 4);
uint16_t fmt = 1, ch = 1; memcpy(h + 20, &fmt, 2); memcpy(h + 22, &ch, 2);
uint32_t rate = SAMPLE_RATE; memcpy(h + 24, &rate, 4); memcpy(h + 28, &byte_rate, 4);
uint16_t align = 2, bits = 16; memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
memcpy(h + 36, "data", 4); memcpy(h + 40, &data_bytes, 4);
}
/* ===========================================================================
* HTTP response reader — shared by all three cloud calls.
* Returns the status code and fills body_out. Handles chunked transfer
* encoding PROPERLY: the chunk-size markers must be stripped, or they end
* up embedded inside the JSON body and the parse fails on long replies.
* =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
* on timeout / closed-and-empty. Every read below goes through this, because
* a reasoning model can think for many seconds between the response headers
* and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
uint32_t t0 = millis();
while (!c.available()) {
if (!c.connected()) return false;
if (millis() - t0 > ms) return false;
delay(10);
}
return true;
}
static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
body_out = "";
if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
while (waitData(client, idle_ms)) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break; // blank line = end of headers
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
}
if (chunked) {
int blanks = 0;
while (true) {
/* The chunk-size line may not arrive for a long time while the model
* reasons. Waiting here - instead of letting read() time out - is the
* whole fix: a timed-out read looks exactly like "0" (final chunk),
* which silently truncated the body to nothing. */
if (!waitData(client, idle_ms)) {
Serial.println("HTTP: timed out waiting for the next chunk");
break;
}
String szline = client.readStringUntil('\n');
szline.trim();
if (szline.length() == 0) { // stray blank line
if (++blanks > 4) break;
continue;
}
blanks = 0;
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break; // genuine final chunk
long got = 0;
while (got < sz) {
if (!waitData(client, idle_ms)) break;
while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
}
if (waitData(client, 3000)) client.readStringUntil('\n'); // CRLF after chunk
if (got < sz) { Serial.println("HTTP: short chunk"); break; }
}
} else {
while (waitData(client, idle_ms))
while (client.available()) body_out += (char)client.read();
}
return code;
}
/* ===========================================================================
* CLOUD CALL 1 — Azure speech-to-text
* One POST, one header, plain WAV body. This is why Azure does the ears.
* =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
/* HTTPClient's one-shot POST fails on bodies this large (it attempts one
* giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
* function speaks HTTP directly and streams the WAV up in 4 KB chunks -
* reliable, and if it ever stalls we know the exact byte it stopped at. */
WiFiClientSecure client;
client.setInsecure(); // no cert bundle on-device; see notes
client.setTimeout(15); // seconds, for reads
writeWavHeader(wav_buf, audio_bytes);
dumpWavToSD(audio_bytes); // PC-playable copy of what we send
size_t total = WAV_HEADER_LEN + audio_bytes;
if (!client.connect(AZURE_STT_HOST, 443)) {
snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
Serial.println("STT: TLS connect failed");
return false;
}
/* Two valid host forms use DIFFERENT URL paths - detect which one is in
* secrets.h: <resource>.cognitiveservices.azure.com -> /stt/speech/...
* <region>.stt.speech.microsoft.com -> /speech/... */
bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
"/speech/recognition/conversation/cognitiveservices/v1"
"?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
"Host: " AZURE_STT_HOST "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
"Accept: application/json\r\n"
"Connection: close\r\n"
"Content-Length: " + String(total) + "\r\n\r\n";
client.print(req);
/* body, 4 KB at a time */
size_t sent = 0;
while (sent < total) {
size_t n = min((size_t)4096, total - sent);
size_t w = client.write(wav_buf + sent, n);
if (w == 0) {
delay(50); // brief stall - retry once
w = client.write(wav_buf + sent, n);
if (w == 0) {
snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
(unsigned)(sent / 1024));
Serial.printf("STT: upload stalled at %u/%u bytes\n",
(unsigned)sent, (unsigned)total);
client.stop();
return false;
}
}
sent += w;
yield();
}
Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);
/* read the reply with proper de-chunking */
String resp;
int code = readHttpResponse(client, resp, 10000);
client.stop();
if (code != 200) {
snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
return false;
}
const char *status = doc["RecognitionStatus"];
if (!status || strcmp(status, "Success") != 0) {
/* The status names the exact failure:
* InitialSilenceTimeout = Azure heard silence (mic level too low)
* NoMatch = heard sound but no recognisable words
* BabbleTimeout = heard only noise */
snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
return false;
}
text_out = doc["DisplayText"].as<String>();
if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
return text_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 2 — DeepSeek chat completion
* OpenAI-compatible format. Model name is deepseek-v4-flash - the old
* deepseek-chat name is dead, see secrets.h.
* =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
/* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
* bodies (long answers + the model's hidden reasoning), which shows up as
* "bad JSON reply". Reading until the server closes the connection is
* reliable regardless of reply length. */
JsonDocument req;
req["model"] = DEEPSEEK_MODEL;
req["max_tokens"] = LLM_MAX_TOKENS;
JsonArray msgs = req["messages"].to<JsonArray>();
JsonObject sys = msgs.add<JsonObject>();
sys["role"] = "system"; sys["content"] = SYSTEM_PROMPT;
JsonObject usr = msgs.add<JsonObject>();
usr["role"] = "user"; usr["content"] = question;
String body;
serializeJson(req, body);
WiFiClientSecure client;
client.setInsecure();
/* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
* ~2 s, but anything that needs actual working-out (an Ohm's law problem,
* say) can think for 10-30 s before sending a single byte. */
client.setTimeout(60);
if (!client.connect(DEEPSEEK_HOST, 443)) {
snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
Serial.println("LLM: TLS connect failed");
return false;
}
client.print(String("POST /chat/completions HTTP/1.1\r\n"
"Host: " DEEPSEEK_HOST "\r\n"
"Authorization: Bearer " DEEPSEEK_KEY "\r\n"
"Content-Type: application/json\r\n"
"Connection: close\r\n"
"Content-Length: ") + String(body.length()) + "\r\n\r\n");
client.print(body);
/* read the reply with proper de-chunking; generous window - long
* questions make the model think for a while before it responds */
String resp;
int code = readHttpResponse(client, resp, 60000);
client.stop();
if (code != 200) {
snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
return false;
}
const char *content = doc["choices"][0]["message"]["content"];
const char *finish = doc["choices"][0]["finish_reason"];
if (!content || !content[0]) {
/* v4-flash is a reasoning model: if finish_reason is "length", the whole
* token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
if (finish && strcmp(finish, "length") == 0)
snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
else
snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
return false;
}
answer_out = String(content);
answer_out.trim();
if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
return answer_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 3 — Azure text-to-speech, streamed straight to the speaker
* We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
* the I2S peripheral eats. Skip the 44-byte header, forward the rest.
* No MP3 decoder, no audio library, no buffering the whole reply.
* =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
size_t got = 0;
uint32_t t0 = millis();
while (got < n && millis() - t0 < 10000) {
int r = c.read(dst + got, n - got);
if (r > 0) { got += r; t0 = millis(); }
else if (!c.connected() && !c.available()) break;
else delay(2);
}
return got;
}
bool azureTTSSpeak(const String &text) {
// Escape the XML special characters for the SSML body
String safe = text;
safe.replace("&", "&");
safe.replace("<", "<");
safe.replace(">", ">");
String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
"<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";
/* Manual HTTP like the other two cloud calls - and for a hard reason:
* Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
* raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
* into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
* Here we parse the framing properly and keep only clean audio bytes. */
WiFiClientSecure client;
client.setInsecure();
client.setTimeout(20);
const char *host = AZURE_REGION ".tts.speech.microsoft.com";
if (!client.connect(host, 443)) {
Serial.println("TTS: TLS connect failed");
return false;
}
client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
"Host: ") + host + "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: application/ssml+xml\r\n"
"X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
"User-Agent: MaTouchRobojax\r\n"
"Connection: close\r\n"
"Content-Length: " + String(ssml.length()) + "\r\n\r\n");
client.print(ssml);
/* status + headers; note whether the body is chunked */
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
long content_len = -1;
while (client.connected() || client.available()) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break;
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
}
if (code != 200) {
Serial.printf("TTS HTTP %d\n", code);
client.stop();
return false;
}
const size_t AUDIO_CAP = 1200 * 1024; // ~37 s of speech
uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
if (!audio) { client.stop(); return false; }
size_t alen = 0;
if (chunked) {
/* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
while (true) {
String szline = client.readStringUntil('\n');
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break;
if (alen + sz > AUDIO_CAP) break;
size_t got = readExact(client, audio + alen, sz);
alen += got;
client.readStringUntil('\n'); // trailing CRLF after each chunk
if (got < (size_t)sz) break;
}
} else if (content_len > 0) {
alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
} else {
/* no framing info: read until the server closes */
uint32_t idle = millis();
while ((client.connected() || client.available()) && millis() - idle < 5000) {
int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
if (r > 0) { alen += r; idle = millis(); }
else delay(5);
}
}
client.stop();
Serial.printf("TTS: %u KB clean audio (%s), playing\n",
(unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");
bool ok = (alen > WAV_HEADER_LEN);
if (ok) {
/* NOW the audio actually starts - this is the honest moment to go green */
LED_SPEAK();
drawBar("SPEAKING...", gfx->color565(0, 130, 40));
/* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
static const uint8_t lead_in[640] = {0}; // 20 ms of silence
size_t w = 0;
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
}
free(audio);
// let the DMA buffers drain so the last word is not cut off
delay(150);
i2s_zero_dma_buffer(I2S_SPK_PORT);
return ok;
}
/* ===========================================================================
* SETUP
* =========================================================================== */
void setup() {
Serial.begin(115200);
delay(400);
Serial.println("\n=== 04 Voice Assistant | Robojax.com ===");
Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");
pinMode(TFT_BLK, OUTPUT);
digitalWrite(TFT_BLK, LOW);
pinMode(SD_CS, OUTPUT);
digitalWrite(SD_CS, HIGH);
// one shared SPI bus for TFT + SD (started before either device)
SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);
gfx->begin();
gfx->fillScreen(BLACK);
digitalWrite(TFT_BLK, HIGH);
// SD is optional here - it only stores the /stt_debug.wav diagnostic copy
ok_sd = SD.begin(SD_CS, SPI, 20000000);
digitalWrite(SD_CS, HIGH);
Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
: "SD not found - debug WAV dump disabled (not fatal)");
bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
delay(50);
rgb.begin();
rgb.setBrightness(LED_BRIGHTNESS);
LED_IDLE();
/* One recording buffer for the whole session, in PSRAM. This is the 8 MB
* that makes the board worth buying. */
wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
if (!wav_buf) {
gfx->setTextColor(RED);
gfx->setTextSize(2);
gfx->setCursor(10, 100);
gfx->print("PSRAM alloc failed!");
gfx->setTextSize(1);
gfx->setCursor(10, 130);
gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
while (1) delay(1000);
}
micInit();
spkInit();
gfx->setTextSize(1);
gfx->setTextColor(YELLOW);
gfx->setCursor(4, 4);
gfx->printf("Connecting to %s ...", WIFI_SSID);
Serial.printf("Connecting to %s ", WIFI_SSID);
WiFi.mode(WIFI_STA);
WiFi.begin(WIFI_SSID, WIFI_PASS);
uint32_t t0 = millis();
while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
delay(300);
Serial.print(".");
}
Serial.println();
clearChat();
if (WiFi.status() == WL_CONNECTED) {
Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
chatBubble("Hold SPEAK and ask me anything.", false);
} else {
chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
LED_ERROR();
}
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}
/* ===========================================================================
* LOOP — one full conversation turn per button press
* =========================================================================== */
void loop() {
/* CLEAR button: edge-detected so one tap wipes once. Reading the panel
* twice per loop (here and in speakButtonHeld) is fine - the GT911 just
* reports its current state. */
static bool tap_latch = false;
static uint8_t tap_release = 0;
if (state == ST_IDLE) {
uint16_t cx, cy;
if (getTouch(&cx, &cy)) {
tap_release = 0;
if (!tap_latch) {
tap_latch = true;
if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
clearChat();
chatBubble("Hold SPEAK and ask me anything.", false);
}
}
} else if (tap_latch && ++tap_release >= 4) {
tap_latch = false;
tap_release = 0;
}
/* live WiFi signal indicator, refreshed every 2 s while idle */
static uint32_t last_wifi = 0;
if (millis() - last_wifi > 2000) {
last_wifi = millis();
drawWifi();
}
}
if (state == ST_IDLE && speakButtonHeld()) {
/* ---- record ---- */
state = ST_RECORDING;
LED_LISTEN();
drawBar("LISTENING...", gfx->color565(0, 60, 200));
uint32_t t_rec = millis();
size_t audio_bytes = recordWhileHeld();
t_rec = millis() - t_rec;
Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);
if (audio_bytes < SAMPLE_RATE / 2) { // under a quarter second - a tap
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
/* ---- speech to text ---- */
state = ST_STT;
LED_THINK();
drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
uint32_t t_stt = millis();
String question;
if (!azureSTT(audio_bytes, question)) {
/* Show the REAL cause on screen - no serial monitor needed. */
char diag[96];
snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
g_stt_err, g_mic_peak_pct,
ok_sd ? " | saved /stt_debug.wav" : "");
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_stt = millis() - t_stt;
chatBubble(question.c_str(), true);
Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());
/* ---- think ---- */
state = ST_LLM;
drawBar("THINKING...", gfx->color565(150, 90, 0));
uint32_t t_llm = millis();
String answer;
if (!deepseekChat(question, answer)) {
char diag[96];
snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_llm = millis() - t_llm;
chatBubble(answer.c_str(), false);
Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());
/* ---- speak ----
* Still amber here: the voice has to be synthesised and downloaded first
* (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
* at the exact moment audio starts coming out of the speaker. */
state = ST_TTS;
LED_THINK();
drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
uint32_t t_tts = millis();
bool spoke = azureTTSSpeak(answer);
t_tts = millis() - t_tts;
/* Timing summary on serial - this feeds the "honest numbers" segment. */
Serial.printf("TIMINGS rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
audio_bytes / 32000.0, (unsigned long)t_stt,
(unsigned long)t_llm, (unsigned long)t_tts,
spoke ? "" : " (TTS FAILED)");
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
}
delay(20);
}
Things you might need
-
OtherProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
Resources & references
-
DocumentationMakerfabs MaTouch ESP32-S3 2.8" Camera and Touchscreen: user's manualwiki.makerfabs.com
-
Documentation
-
DocumentationProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
-
DownloadArduino GFX Library on Githubgithub.com
Files📁
Required File (.h)
Other Files
Schematic
-
MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematicThe latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF0.15 MB