شِفر (کود) جستجو

سازنده‌ی مِیکرفبز ما تاچ ESP32-S3 با دوربین ۲.۸ اینچی — ساخت دستیار صوتی هوش مصنوعی روی ESP32-S3 (آژور + دیپ‌سیک)

سازنده‌ی مِیکرفبز ما تاچ ESP32-S3 با دوربین ۲.۸ اینچی — ساخت دستیار صوتی هوش مصنوعی روی ESP32-S3 (آژور + دیپ‌سیک)

یک دکمه را نگه دارید، سوالی بپرسید و برد با صدای بلند پاسخ دهد

یک دستیار صوتی کامل روی یک برد کوچک. دکمه SPEAK را نگه دارید و سوالی بپرسید. برد با هر دو میکروفون صدای شما را ضبط می‌کند، صدا را به Microsoft Azure می‌فرستد تا به متن تبدیل شود، آن متن را به DeepSeek می‌فرستد تا فکر کند، پاسخ را دوباره به Azure می‌فرستد تا به گفتار تبدیل شود و آن را از طریق بلندگوی خود پخش می‌کند. کل مکالمه به صورت حباب‌های چت روی صفحه ظاهر می‌شود.

دستیار صوتی هوش مصنوعی سخنگو در حال اجرا روی برد MaTouch AI ESP32-S3 دستیار صوتی هوش مصنوعی سخنگو در حال اجرا روی برد MaTouch AI ESP32-S3

از لحظه‌ای که دکمه را فشار می‌دهید چه اتفاقی می‌افتد

در اینجا کل مسیر یک سوال، قدم به قدم آمده است. ارزش دارد یک بار خوانده شود، زیرا هر چیزی که روی صفحه و روی LED می‌بینید به یکی از این مراحل نگاشت می‌شود.

  1. دکمه SPEAK را فشار می‌دهید و نگه می‌دارید. این حالت صحبت با نگه‌داشتن است، نه صحبت با ضربه زدن: ضبط دقیقاً تا زمانی که انگشت شما پایین بماند ادامه دارد، حداکثر تا شش ثانیه. LED وضعیت به رنگ آبی تغییر می‌کند و دکمه LISTENING را نشان می‌دهد.

  2. هر دو میکروفون صدای شما را ضبط می‌کنند. برد ۱۶,۰۰۰ بار در ثانیه از جفت استریو نمونه‌برداری می‌کند، دو کانال را به یک کانال میانگین می‌گیرد، کمی بهره اعمال می‌کند و نتیجه را در PSRAM ذخیره می‌کند. یک نوار پیشرفت هنگام صحبت کردن روی دکمه حرکت می‌کند. دو ثانیه گفتار حدود ۶۴ KB است.

  3. دکمه را رها می‌کنید. ضبط متوقف می‌شود. برد یک هدر WAV ۴۴ بایتی به ابتدای صدا اضافه می‌کند - آن برچسب کوچک تنها چیزی است که نمونه‌های خام را به فایلی تبدیل می‌کند که Azure می‌پذیرد.

  4. صدا به Azure Speech-to-Text می‌رود. در قطعات ۴ KB از طریق یک اتصال امن آپلود می‌شود و به صورت یک خط متن برمی‌گردد. ۶۴ KB صدای شما به حدود ۲۵ بایت نوشتار تبدیل شده است. LED به رنگ کهربایی تغییر می‌کند.

  5. سوال شما روی صفحه ظاهر می‌شود به صورت یک حباب چت آبی، تا دقیقاً ببینید چه چیزی شنیده است - که مفید است، زیرا کلمات اشتباه شنیده شده بیشتر پاسخ‌های عجیب را توضیح می‌دهند.

  6. متن به DeepSeek می‌رود. برد سوال شما را به همراه یک دستورالعمل ثابت برای کوتاه نگه‌داشتن پاسخ‌ها به دو جمله کوتاه می‌فرستد. مدل فکر می‌کند - واقعاً فکر می‌کند، یک مدل استدلالی است - و پاسخی برمی‌گرداند.

  7. پاسخ روی صفحه ظاهر می‌شود به صورت یک حباب خاکستری. می‌توانید قبل از شنیدن آن را بخوانید.

  8. پاسخ به Azure برگردانده می‌شود تا گفته شود. برد صوتی PCM خام ۱۶ kHz درخواست می‌کند، که دقیقاً فرمتی است که تقویت‌کننده آن می‌خواهد، بنابراین هیچ رمزگشای MP3 در هیچ جای این پروژه وجود ندارد. دکمه اکنون GETTING VOICE را نشان می‌دهد و LED کهربایی می‌ماند، زیرا هنوز چیزی قابل شنیدن نیست.

  9. کل کلیپ قبل از پخش حتی یک نمونه، در PSRAM دانلود می‌شود. این مهم است - به یادداشت زیر مراجعه کنید.

  10. پخش. لحظه‌ای که صدا به بلندگو تحویل داده می‌شود، LED به رنگ سبز تغییر می‌کند و دکمه SPEAKING را نشان می‌دهد. پاسخ را می‌شنوید.

چرا صدا ابتدا دانلود می‌شود به جای اینکه هنگام رسیدن پخش شود. پخش مستقیم از شبکه به بلندگو مانند صدای تق‌تق به نظر می‌رسد. بافر بلندگو فقط حدود یک دهم ثانیه را نگه می‌دارد و هر مکث در انتقال WiFi طولانی‌تر از آن آن را خالی می‌کند و یک تق‌تق قابل شنیدن ایجاد می‌کند. دانلود کل پاسخ ابتدا در PSRAM حدود یک ثانیه انتظار اضافی هزینه دارد و همه شکاف‌ها را حذف می‌کند. به همین دلیل است که نمایشگر قبل از گفتن SPEAKING، GETTING VOICE را نشان می‌دهد - این دو واقعاً مراحل متفاوتی هستند.

هر مرحله چقدر طول می‌کشد

اندازه‌گیری شده روی سخت‌افزار واقعی، برای یک سوال ساده:

مرحله

زمان معمول

ضبط

تا زمانی که دکمه را نگه دارید

گفتار به متن (Azure)

حدود ۱.۸ ثانیه

فکر کردن (DeepSeek)

حدود ۱.۸ ثانیه برای یک سوال ساده، بسیار بیشتر برای سوالی که نیاز به محاسبه واقعی دارد

دریافت صدا (Azure)

حدود ۷ ثانیه - بزرگترین بخش

کل، از رها کردن تا اولین صدا

تقریباً ۱۱ ثانیه

هر تبادل زمان‌بندی‌های خود را در نمایشگر مسلسل چاپ می‌کند، بنابراین می‌توانید زمان‌های خودتان را اندازه بگیرید به جای اعتماد به این‌ها. اگر آن را سریع‌تر می‌خواهید، مؤثرترین تغییر درخواست پاسخ‌های کوتاه‌تر در SYSTEM_PROMPT است - متن کمتر برای گفتن یعنی صدای کمتر برای سنتز و دانلود.

خود برد هرگز چیزی را نمی‌فهمد. این یک پیام‌رسان با گوش‌های خوب و صدای خوب است - هوش به ازای هر ثانیه اجاره می‌شود.

چرا این سه سرویس

  • Azure گفتار ورودی و خروجی را مدیریت می‌کند. متن به گفتار آن می‌تواند PCM خام ۱۶ kHz برگرداند، که دقیقاً همان چیزی است که تراشه بلندگو می‌خواهد، بنابراین هیچ رمزگشای MP3 در هیچ جای این پروژه وجود ندارد. گفتار به متن آن یک WAV ساده را در یک POST ساده می‌پذیرد.

  • DeepSeek مغز مکالمه است. سریع است و هزینه هر پاسخ کسری از یک سنت است.

  • OpenAI در اینجا استفاده نمی‌شود - به پروژه ۰۵ مراجعه کنید، جایی که کار بینایی را انجام می‌دهد.

نام مدل‌های DeepSeek تغییر کرده است. نام‌های قدیمی deepseek-chat و deepseek-reasoner در ژوئیه ۲۰۲۶ بازنشسته شدند. بیشتر آموزش‌های آنلاین هنوز از آن‌ها استفاده می‌کنند و خطا برمی‌گردانند. نام‌های فعلی deepseek-v4-flash و deepseek-v4-pro هستند. این پروژه از v4-flash استفاده می‌کند.

تله مدل استدلالی

DeepSeek v4-flash قبل از پاسخ دادن فکر می‌کند، و این تفکر از سهمیه توکن شما کم می‌کند. اگر LLM_MAX_TOKENS را خیلی پایین تنظیم کنید، کل بودجه صرف استدلال می‌شود، پاسخ خالی برمی‌گردد، و برد هیچ چیزی نمی‌گوید. به همین دلیل اینجا روی 400 تنظیم شده است. سوالات سخت نیز زمان بیشتری می‌برند - یک پاسخ ساده حدود دو ثانیه طول می‌کشد، سوالی که نیاز به محاسبه واقعی دارد می‌تواند خیلی بیشتر طول بکشد.

خواندن چراغ وضعیت

رنگ

معنی

آبی

در حال گوش دادن به شما

کهربایی

ابر در حال فکر کردن است، یا صدا در حال دریافت است

سبز

در حال صحبت کردن - این دقیقاً در لحظه شروع صدا سبز می‌شود

قرمز

خطایی رخ داده است - نمایشگر مسلسل را بررسی کنید

کنترل‌های روی صفحه

چت مانند یک مکالمه تلفنی اسکرول می‌شود، پیام‌های قدیمی‌تر به بالا و خارج می‌روند. CLEAR آن را پاک می‌کند. یک نشانگر قدرت سیگنال WiFi با خوانش واقعی dBm در گوشه قرار دارد، که وقتی تعجب می‌کنید که آیا پاسخ کند به خاطر شبکه است یا سرویس، مفید است.

درباره برد MaTouch AI ESP32-S3 2.8"

هر پروژه در این صفحه بر روی MaTouch AI ESP32-S3 2.8" TFT ST7789V از Makerfabs اجرا می‌شود. این یک برد همه‌کاره است: یک صفحه لمسی رنگی، یک دوربین 3 مگاپیکسلی، دو میکروفون و یک تقویت‌کننده بلندگوی واقعی، همه توسط یک ESP32-S3 با 8 مگابایت PSRAM هدایت می‌شوند. این ترکیب همان چیزی است که این پروژه‌های هوش مصنوعی را روی یک برد واحد بدون هیچ چیز متصل دیگری ممکن می‌سازد.

8 مگابایت PSRAM بیشتر از هر عدد دیگری در اینجا اهمیت دارد. این همان چیزی است که به برد اجازه می‌دهد یک فریم دوربین، چند ثانیه صدای ضبط‌شده، یا یک عکس کدگذاری‌شده base64 را همزمان در حافظه نگه دارد - هیچ‌کدام از این‌ها در RAM معمولی ESP32 جا نمی‌شود.

مستندات سازنده: صفحه ویکی Makerfabs.

مشخصات کلیدی

  • پردازنده: ESP32-S3، دو هسته 240 مگاهرتز، WiFi 2.4 گیگاهرتز + بلوتوث 5.0

  • حافظه: 16 مگابایت فلش، 8 مگابایت PSRAM (تقریباً برای هر پروژه اینجا الزامی است)

  • نمایشگر: 2.8" IPS، 320×240، درایور ST7789V، SPI

  • لمس: GT911 خازنی، ردیابی 5 انگشت به طور همزمان

  • دوربین: OV3660، 3 مگاپیکسل، تا 2048×1536

  • میکروفون‌ها: دو میکروفون دیجیتال I2S INMP441 (یک جفت استریوی واقعی)

  • بلندگو: تقویت‌کننده کلاس-D MAX98357A، 3.2 وات در 4 اهم

  • ذخیره‌سازی: اسلات کارت microSD (حالت SPI)

  • برق: USB-C، کانکتور باتری JST، شارژر TP4056، کلید برق

  • همچنین روی برد: LED RGB WS2812B، ساعت واقعی با باتری PCF8563T، و یک نشانگر سطح باتری MAX17048 که در مشخصات رسمی لیست نشده است

دو پورت USB-C یکسان نیستند. بلندگوی برد پایه‌های سیگنال خود (IO19 و IO20) را با پورت USB بومی به اشتراک می‌گذارد، زیرا آن پایه‌ها خطوط داده USB سخت‌افزاری ESP32-S3 هستند. همیشه از طریق پورت USB-C CH340K (آن یکی کنار دکمه RESET) آپلود و برق دهید، و USB CDC On Boot را روی Disabled تنظیم کنید. اگر از پورت اشتباه استفاده کنید، صدا رفتار نامناسبی خواهد داشت یا آپلودها ناموفق خواهند بود.

تنظیمات Arduino IDE

این تنظیمات مهم هستند. بیشتر مشکلاتی که مردم با این برد گزارش می‌دهند یکی از این موارد است که اشتباه است، و آنها وقتی نسخه هسته را تغییر می‌دهید بازنشانی می‌شوند، بنابراین بعد از هر تغییری دوباره آن‌ها را بررسی کنید.

تنظیم

مقدار

برد

ESP32S3 Dev Module

نسخه هسته ESP32

2.0.17

PSRAM

OPI PSRAM

اندازه فلش

16MB (128Mb)

طرح پارتیشن

16M Flash (3MB APP/9.9MB FATFS)

USB CDC On Boot

Disabled

سرعت آپلود

921600

پاک کردن همه فلش قبل از آپلود

Disabled

پورت

پورت USB-C CH340K

از هسته ESP32 نسخه 2.0.17 استفاده کنید، نه 3.x. Espressif مدل‌های تشخیص چهره روی دستگاه را در هسته 3 حذف کرده است، بنابراین پروژه‌های چهره آنجا کامپایل نمی‌شوند. ثابت نگه داشتن 2.0.17 باعث می‌شود هر پروژه در این صفحه با یک پیکربندی کار کند. در Boards Manager، منوی کشویی نسخه به شما اجازه می‌دهد هر زمان که خواستید بین آن‌ها جابه‌جا شوید.

از GFX Library for Arduino نسخه 1.5.6 استفاده کنید، نه 1.6.x. نسخه‌های 1.6 برای هسته ESP32 نسخه 3 ساخته شده‌اند و می‌توانند در راه‌اندازی روی هسته 2.0.17 متوقف شوند. اگر صفحه شما بعد از آپلود سیاه ماند، این اولین چیزی است که باید بررسی کنید.

کتابخانه‌های مورد نیاز

این‌ها را از طریق Tools → Manage Libraries در Arduino IDE نصب کنید. شماره نسخه‌ها مهم هستند - لطفاً از موارد ذکر شده استفاده کنید.

کتابخانه

نسخه

نویسنده

GFX Library for Arduino

1.5.6

moononournation

bb_captouch

1.3.1

Larry Bank

ArduinoJson

7.x

Benoit Blanchon

Adafruit NeoPixel

هر نسخه اخیر

Adafruit

راه‌اندازی secrets.h

جزئیات وای‌فای شما و هر کلید API در secrets.h قرار می‌گیرد که در دانلود با مقادیر نمونه گنجانده شده است. آن تب را در Arduino IDE باز کنید و آن‌ها را با مقادیر خودتان جایگزین کنید.

وای‌فای باید 2.4 گیگاهرتز باشد. ESP32-S3 به هیچ وجه نمی‌تواند شبکه 5 گیگاهرتز را ببیند. اگر روتر شما هر دو باند را تحت یک نام ترکیب می‌کند (ایسوس این را Smart Connect می‌نامد)، یا آن را خاموش کنید یا به باند 2.4 گیگاهرتز نام جداگانه‌ای بدهید و از آن در secrets.h استفاده کنید.

دریافت کلیدهای API شما

این پروژه با یک سرویس هوش مصنوعی ابری ارتباط برقرار می‌کند، بنابراین به کلید اختصاصی خودتان نیاز دارید. اگر قبلاً این کار را انجام نداده‌اید، نگران نباشید - این همان ایده رمز عبوری است که حساب شما را به سرویس معرفی می‌کند. چند دقیقه طول می‌کشد، فقط یک بار.

یک کلید اشتراک یک وب‌سایت نیست. برای مثال، پرداخت برای ChatGPT Plus به شما کلید API نمی‌دهد - این دو محصول جداگانه با صورتحساب جداگانه هستند. شما به یک حساب در پلتفرم توسعه‌دهنده نیاز دارید، که در زیر توضیح داده شده است.

Microsoft Azure Speech - برای گوش دادن و صحبت کردن

Azure گفتار شما را به متن تبدیل می‌کند و پاسخ را دوباره به صدا تبدیل می‌کند. سطح رایگان برای همه چیز در این صفحه به اندازه کافی سخاوتمندانه است.

  1. به portal.azure.com بروید و با یک حساب مایکروسافت وارد شوید (حساب رایگان کافی است).

  2. اگر قبلاً از Azure استفاده نکرده‌اید، صفحه Welcome to Azure را با سه گزینه مشاهده خواهید کرد. Start with an Azure free trial را انتخاب کنید - قبل از اینکه Azure به شما اجازه ایجاد هر چیزی را بدهد، به یک اشتراک نیاز دارید. (دانشجویان باید به جای آن Azure for Students را انتخاب کنند: نتیجه یکسان، بدون نیاز به کارت.) گزینه Manage Microsoft Entra ID را نادیده بگیرید که کاملاً چیز دیگری است.

  3. روی Create a resource کلیک کنید، برای Speech جستجو کنید و Speech service منتشر شده توسط مایکروسافت را انتخاب کنید.

  4. فرم را پر کنید: هر گروه منبع، هر نام، و یک Region نزدیک خود انتخاب کنید - آن منطقه را دقیقاً همان‌طور که ظاهر می‌شود یادداشت کنید، برای مثال eastus.

  5. برای Pricing tier گزینه F0 (Free) را انتخاب کنید. این حدود پنج ساعت تبدیل گفتار به متن و پانصد هزار خصیصه تبدیل متن به گفتار در هر ماه را امکان‌پذیر می‌کند.

  6. روی Review + create و سپس Create کلیک کنید. حدود یک دقیقه صبر کنید، سپس روی Go to resource کلیک کنید.

  7. در منوی سمت چپ Keys and Endpoint را باز کنید. KEY 1 و Location/Region را کپی کنید.

آن‌ها را در secrets.h به عنوان AZURE_SPEECH_KEY و AZURE_REGION قرار دهید. برای AZURE_STT_HOST، از <region>.stt.speech.microsoft.com استفاده کنید - بنابراین با منطقه eastus آن eastus.stt.speech.microsoft.com می‌شود.

درباره کارت اعتباری. آزمایش رایگان Azure برای تأیید هویت شما کارت می‌خواهد. از شما هزینه‌ای دریافت نمی‌کند. شما 200 دلار اعتبار برای 30 روز دریافت می‌کنید و پس از آن حساب به Pay-As-You-Go منتقل می‌شود - اما سطح F0 Speech رایگان می‌ماند، ماه به ماه، و همه چیز در این پروژه‌ها به راحتی در آن جا می‌شود. اگر ترجیح می‌دهید اصلاً کارت ندهید و دانشجو هستید، گزینه Azure for Students بدون کارت به شما اعتبار می‌دهد.

باید منبع "Speech service" باشد. کلید از منبع Translator، Language یا Cognitive Services عمومی یکسان به نظر می‌رسد و کاملاً معتبر است - اما هر درخواست گفتار خطای 401 برمی‌گرداند. این در طول آزمایش ما را گرفتار کرد و یک ساعت هزینه داشت. اگر گفتار با 401 شکست خورد در حالی که کلید درست به نظر می‌رسد، بررسی کنید چه نوع منبعی ایجاد کرده‌اید.

DeepSeek - بخش تفکر

DeepSeek مدل زبانی است که در واقع به سؤال شما پاسخ می‌دهد. ارزان است - چند دلار اعتبار هزاران پاسخ را پوشش می‌دهد.

  1. به platform.deepseek.com بروید و یک حساب ایجاد کنید.

  2. API keys را در منو باز کنید و روی Create new API key کلیک کنید.

  3. فوراً آن را کپی کنید. فقط یک بار نمایش داده می‌شود و هرگز دوباره - اگر آن را گم کردید، آن کلید را حذف کنید و یکی دیگر بسازید.

  4. مقدار کمی اعتبار زیر Top up اضافه کنید. سطح رایگان وجود ندارد، اما کوچکترین شارژ در این میزان استفاده مدت بسیار طولانی دوام می‌آورد.

کلید را در secrets.h به عنوان DEEPSEEK_KEY قرار دهید. با sk- شروع می‌شود.

نام مدل‌ها در ژوئیه 2026 تغییر کرد. مدل‌های قدیمی deepseek-chat و deepseek-reasoner بازنشسته شدند، بنابراین بیشتر آموزش‌هایی که آنلاین پیدا می‌کنید با خطای 400 شکست می‌خورند. از deepseek-v4-flash استفاده کنید که این پروژه‌ها از قبل تنظیم کرده‌اند.

هزینه اجرا چقدر است

بسیار کم، اما رایگان نیست، و باید تقریباً بدانید قبل از اینکه پروژه‌ای را در حال اجرا بگذارید چه هزینه‌ای می‌کنید.

سرویس

هزینه تقریبی

Azure Speech

سطح رایگان حدود 5 ساعت گوش دادن و 0.5 میلیون خصیصه صحبت کردن در ماه را پوشش می‌دهد

DeepSeek

کسری از سنت برای هر پاسخ - هزاران پاسخ با چند دلار

OpenAI vision

تقریباً یک یا دو سنت برای هر تصویر، بسته به مدل

قیمت‌ها تغییر می‌کنند، بنابراین این‌ها را به عنوان راهنما در نظر بگیرید نه قیمت قطعی. هر یک از این سرویس‌ها صفحه مصرف دارد که می‌توانید هزینه‌های خود را تماشا کنید و همه آن‌ها به شما اجازه تنظیم سقف هزینه می‌دهند - که ارزش انجام دادن از روز اول را دارد.

کلیدهای خود را خصوصی نگه دارید. هر کسی که به آن‌ها دسترسی داشته باشد می‌تواند پول شما را خرج کند. آن‌ها را در ویدیو، اسکرین‌شات، پست انجمن یا مخزن شِفر (کود) عمومی قرار ندهید. اگر کلیدی فاش شد، آن را در وب‌سایت ارائه‌دهنده حذف کنید و یک کلید جدید بسازید - این کار چند ثانیه طول می‌کشد و تنها راه‌حل واقعی است.

عیب‌یابی

علامت

علت و راه‌حل

صفحه سیاه می‌ماند

نسخه کتابخانه GFX اشتباه است (از 1.5.6 استفاده کنید) یا تنظیمات برد اشتباه است.

PSRAM alloc failed یا خطای دوربین 0xffffffff

Tools → PSRAM روی OPI PSRAM تنظیم نشده است.

هیچ چیزی آپلود نمی‌شود / پورت COM وجود ندارد

پورت USB-C اشتباه است، یا درایور CH340 نصب نشده است.

دوربین از کار می‌افتد و هرگز بازیابی نمی‌شود

خط ریست دوربین به دکمه RESET برد متصل است، بنابراین نرم‌افزار نمی‌تواند آن را راه‌اندازی مجدد کند. RESET را فشار دهید. اگر همچنان از کار می‌افتد، کابل نواری دوربین را دوباره وصل کنید.

دانلود شِفر (کود)

اسکچ کامل آردوینو برای این پروژه، همراه با pins.h و هر چیز دیگری که نیاز دارد، به صورت رایگان قابل دانلود است.

دانلود 04_Voice_Assistant

آن را از حالت فشرده خارج کنید، فایل .ino را در Arduino IDE باز کنید، تنظیمات بالا را بررسی کنید و از طریق پورت CH340K USB-C آپلود کنید.

884-Arduin code for MaTouch AI ESP32S3 2.8in AI Camera: AI voice assistant
زبان: C++
/*
 * ===========================================================================
 *  04_Voice_Assistant  —  MaTouch AI ESP32-S3 2.8" TFT ST7789V
 * ===========================================================================
 *
 ----------
 *  ROBOJAX.COM  -  MaTouch AI ESP32-S3 2.8" project series
 *
 *    WATCH THE VIDEO
 *        https://youtu.be/6AL3g3tC_Hk
 *
 *    WRITTEN TUTORIALS - every project, with photos and full explanation
 *        Camera and touchscreen.... https://robojax.com/RTJ849
 *        Offline face recognition.. https://robojax.com/RTJ850
 *        AI voice assistant........ https://robojax.com/RTJ851
 *        AI vision................. https://robojax.com/RTJ852
 *
 *    GET THE BOARD - SAVE $5 with coupon code:  Robojax_Makerfab
 *        https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
 *        (enter the code at checkout)
 *
 *  All of this code is free. If it helped you, a subscribe on YouTube is
 *  the best way to support more of it.
 *  
 *  ---------------------------------------------------------------------------
 *
 *  A complete voice assistant on a $40 board:
 *
 *      hold SPEAK  ->  both INMP441 microphones record you
 *                  ->  Azure Speech turns the audio into text
 *                  ->  DeepSeek v4-flash thinks of an answer
 *                  ->  Azure Speech turns the answer into audio
 *                  ->  the MAX98357 speaker says it out loud
 *
 *  and the whole conversation is drawn as chat bubbles on the touchscreen.
 *
 *  WHY THIS COMBINATION OF SERVICES (each is used where it is best):
 *   - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
 *     transcription endpoint wants multipart/form-data - miserable on an MCU.
 *   - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
 *     which streams straight into the I2S speaker with NO MP3 decoder at all.
 *   - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
 *     model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
 *     Most tutorials online still use them and are broken. See secrets.h.
 *
 *  A detail the vendor examples get wrong: this board has TWO microphones on
 *  one I2S bus (left + right), but every Makerfabs demo records left-only and
 *  throws one away. This sketch records both and averages them.
 *
 *  ---------------------------------------------------------------------------
 *  *** WHICH USB PORT - THIS MATTERS ***
 *  The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
 *  through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
 *  If you use the wrong port the audio will be garbage or uploads will fail.
 *  ---------------------------------------------------------------------------
 *
 *  FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
 *
 *  BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
 *  or compile errors. These reset when you switch cores - recheck them!)
 *
 *      Board            : ESP32S3 Dev Module
 *      ESP32 core       : 2.0.17
 *      PSRAM            : OPI PSRAM        <-- required, audio buffer lives there
 *      Flash Size       : 16MB (128Mb)
 *      Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
 *      USB CDC On Boot  : Disabled         <-- required, see USB note above
 *      Upload Speed     : 921600
 *      Port             : the CH340K USB-C port (the one near RESET)
 *
 *  LIBRARIES
 *      GFX Library for Arduino   v1.5.6   (NOT 1.6.x - that pairs with core 3)
 *      bb_captouch               v1.3.1
 *      ArduinoJson               v7.x
 *      Adafruit NeoPixel         any recent
 *
 *  ---------------------------------------------------------------------------
 *  FUNCTIONS IN THIS SKETCH
 *      led(r,g,b) + LED_* macros   RGB status colours (blue/amber/green/red)
 *      getTouch(&x,&y)      read the touch panel, mapped to screen coordinates
 *      speakButtonHeld()    true while a finger is on the SPEAK button
 *      bubbleLines(t)       how many lines a message wraps to
 *      drawOneBubble(m,y)   draw a single chat bubble
 *      redrawChat()         rebuild the chat area from history, newest at bottom
 *      clearChat()          wipe the chat history (CLEAR button)
 *      chatBubble(t,user)   add a message to history and redraw
 *      drawWifi()           WiFi signal bars + dBm readout
 *      drawBar(label,col)   bottom bar: SPEAK button + CLEAR + WiFi meter
 *      micInit()            I2S input - BOTH INMP441 mics, stereo
 *      spkInit()            I2S output - MAX98357 speaker
 *      recordWhileHeld()    record while SPEAK held, downmix stereo->mono
 *      writeWavHeader(...)  prepend the 44-byte RIFF/WAVE header
 *      dumpWavToSD(...)     save the exact upload to SD (/stt_debug.wav)
 *      readHttpResponse()   read an HTTPS reply, de-chunking it properly
 *      azureSTT(...)        chunked upload of the WAV -> recognised text
 *      deepseekChat(...)    question -> deepseek-v4-flash -> answer text
 *      azureTTSSpeak(text)  answer -> Azure voice -> PSRAM -> speaker
 *      setup() / loop()     boot + WiFi / one conversation turn per press
 *
 *  Robojax.com
 * ===========================================================================
 */

#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"

/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE     16000
#define RECORD_MAX_S    6                       // hard cap on one question
#define WAV_HEADER_LEN  44
#define REC_BUF_BYTES   (SAMPLE_RATE * RECORD_MAX_S * 2)   // 16-bit mono

/* Software gain applied to the recording. The INMP441 capture is quiet at
 * 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
 * speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
 * lower it. */
#define MIC_GAIN        6

/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
 * 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
 *     back silent or garbled in stereo mode. */
#define USE_BOTH_MICS   1

/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
 * uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS  25

/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
 * shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
    TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);

BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);

/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr;      // WAV_HEADER_LEN + up to REC_BUF_BYTES

/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;

/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool  ok_sd = false;
float g_mic_peak_pct = 0;          // last recording's raw peak, % of full scale
char  g_stt_err[64] = "";          // last STT failure cause, verbatim
char  g_llm_err[64] = "";          // last DeepSeek failure cause, verbatim

/* --- layout --------------------------------------------------------------- */
#define CHAT_H     200              // chat area: y 0..199
#define BAR_Y      202              // button bar below it
#define BTN_SPEAK_X   4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W  58
#define WIFI_X      236             // signal indicator, right end of the bar
#define BTN_H        36

/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
 * This is what prevents new text printing over old - the whole area is
 * rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int     chat_count = 0;

/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
  rgb.setPixelColor(0, rgb.Color(r, g, b));
  rgb.show();
}
#define LED_IDLE()    led(0, 0, 0)
#define LED_LISTEN()  led(0, 60, 255)     // blue   - recording
#define LED_THINK()   led(255, 120, 0)    // amber  - waiting on the cloud
#define LED_SPEAK()   led(0, 255, 40)     // green  - talking
#define LED_ERROR()   led(255, 0, 0)      // red


/* ===========================================================================
 *  Touch
 * =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
  TOUCHINFO ti;
  if (!bbct.getSamples(&ti)) return false;
  if (ti.count < 1) return false;
  *x = ti.y[0];
  *y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
  return true;
}

bool speakButtonHeld() {
  uint16_t x, y;
  if (!getTouch(&x, &y)) return false;
  return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}


/* ===========================================================================
 *  Chat UI  —  word-wrapped bubbles, user right/blue, assistant left/grey
 * =========================================================================== */
#define CHAT_CHARS 42                          // chars per line at textsize 1

static int bubbleLines(const char *t) {
  int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
  return l < 1 ? 1 : l;
}

void drawOneBubble(const ChatMsg &m, int y) {
  int len = strlen(m.text);
  int lines = bubbleLines(m.text);
  int h = lines * 10 + 8;
  uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
  int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
  if (w < 30) w = 30;
  int x = m.from_user ? (316 - w) : 4;

  gfx->fillRoundRect(x, y, w, h, 5, bg);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  for (int i = 0; i < lines; i++) {
    char line[CHAT_CHARS + 1] = {0};
    strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
    gfx->setCursor(x + 5, y + 5 + i * 10);
    gfx->print(line);
  }
}

/* Rebuild the whole chat area from history: newest message anchored at the
 * bottom, older ones stacked upward until the area is full. */
void redrawChat() {
  gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
  int shown = min(chat_count, CHAT_HISTORY);
  int y = CHAT_H - 2;
  for (int i = 0; i < shown; i++) {
    ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
    int h = bubbleLines(m.text) * 10 + 8;
    y -= h;
    if (y < 0) break;                          // area full - older ones drop off
    drawOneBubble(m, y);
    y -= 4;
  }
}

void clearChat() {
  chat_count = 0;
  redrawChat();
}

void chatBubble(const char *text, bool from_user) {
  ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
  strncpy(m.text, text, sizeof(m.text) - 1);
  m.text[sizeof(m.text) - 1] = 0;
  m.from_user = from_user;
  chat_count++;
  redrawChat();
}

/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
  gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
  bool up = (WiFi.status() == WL_CONNECTED);
  long rssi = up ? WiFi.RSSI() : -100;
  // -55 dBm or better = full bars; each 10 dB drops one
  int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;

  for (int b = 0; b < 4; b++) {
    int bh = 6 + b * 6;                       // heights 6,12,18,24
    uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
    gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
  }
  gfx->setTextSize(1);
  gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
  if (up) {
    gfx->setTextColor(CYAN);
    gfx->printf("%lddBm", rssi);
  } else {
    gfx->setTextColor(RED);
    gfx->print("DOWN");
  }
  gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
  gfx->setTextColor(gfx->color565(120, 120, 120));
  gfx->print("WiFi");
}

void drawBar(const char *label, uint16_t colour) {
  gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
  gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
  gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
  gfx->setTextSize(2);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
  gfx->print(label);

  // CLEAR wipes the chat history
  gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
  gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
  gfx->print("CLEAR");

  drawWifi();
}


/* ===========================================================================
 *  I2S  —  microphones on port 0, speaker on port 1. Separate hardware
 *  ports, so recording and playback can never fight over a bus.
 * =========================================================================== */
void micInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    /* BOTH channels - this is the two-microphone fix. The vendor examples
     * use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
     * falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
    .channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = false,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_MIC_SCK,
    .ws_io_num = I2S_MIC_WS,
    .data_out_num = I2S_PIN_NO_CHANGE,
    .data_in_num = I2S_MIC_SD
  };
  i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_MIC_PORT, &pins);
}

void spkInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,   // mono - the amp downmixes anyway
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = true,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_SPK_BCLK,
    .ws_io_num = I2S_SPK_LRC,
    .data_out_num = I2S_SPK_DOUT,
    .data_in_num = I2S_PIN_NO_CHANGE
  };
  i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_SPK_PORT, &pins);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
}


/* ===========================================================================
 *  Recording  —  runs while the SPEAK button is held (up to RECORD_MAX_S).
 *  Reads stereo pairs, averages L+R into one mono stream, applies a little
 *  software gain, and fills wav_buf after the 44-byte header slot.
 *  Returns the number of audio bytes recorded.
 * =========================================================================== */
size_t recordWhileHeld() {
  int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
  size_t   mono_samples = 0;
  const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;

  int16_t chunk[512];
  uint32_t last_touch_ok = millis();
  int32_t  peak = 0;                        // loudest raw sample - mic health check

  i2s_zero_dma_buffer(I2S_MIC_PORT);

  while (mono_samples < max_samples) {
    size_t got = 0;
    i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);

#if USE_BOTH_MICS
    size_t n = got / 4;                     // 4 bytes = one L+R pair
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
    size_t n = got / 2;                     // 2 bytes = one mono sample
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = chunk[i];
#endif
      if (abs(raw) > peak) peak = abs(raw);
      int32_t mixed = raw * MIC_GAIN;
      if (mixed >  32767) mixed =  32767;
      if (mixed < -32768) mixed = -32768;
      mono[mono_samples++] = (int16_t)mixed;
    }

    /* The GT911 is polled between I2S reads. A 250 ms grace period stops a
     * momentary missed touch sample from cutting the recording short. */
    if (speakButtonHeld()) last_touch_ok = millis();
    else if (millis() - last_touch_ok > 250) break;

    // live progress on the button
    static uint32_t last_draw = 0;
    if (millis() - last_draw > 200) {
      last_draw = millis();
      gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
                    (int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
    }
  }

  /* Mic health line: peak as % of full scale BEFORE gain.
   *   0%       = the mic is not being read at all (config/pin problem)
   *   under 3% = too quiet - speak closer or raise MIC_GAIN
   *   3-40%    = healthy speech level
   */
  g_mic_peak_pct = peak * 100.0 / 32768.0;
  Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
                g_mic_peak_pct, MIC_GAIN,
                peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");

  return mono_samples * 2;
}

/* Dump the exact WAV we are about to POST onto the SD card, so it can be
 * played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
  if (!ok_sd) return;
  SD.remove("/stt_debug.wav");
  File f = SD.open("/stt_debug.wav", FILE_WRITE);
  if (!f) return;
  f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
  f.close();
  Serial.println("debug copy saved to SD as /stt_debug.wav");
}

/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
  uint32_t file_len = data_bytes + 36;
  uint32_t byte_rate = SAMPLE_RATE * 2;
  memcpy(h, "RIFF", 4);           memcpy(h + 4, &file_len, 4);
  memcpy(h + 8, "WAVEfmt ", 8);
  uint32_t fmt_len = 16;          memcpy(h + 16, &fmt_len, 4);
  uint16_t fmt = 1, ch = 1;       memcpy(h + 20, &fmt, 2);   memcpy(h + 22, &ch, 2);
  uint32_t rate = SAMPLE_RATE;    memcpy(h + 24, &rate, 4);  memcpy(h + 28, &byte_rate, 4);
  uint16_t align = 2, bits = 16;  memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
  memcpy(h + 36, "data", 4);      memcpy(h + 40, &data_bytes, 4);
}


/* ===========================================================================
 *  HTTP response reader  —  shared by all three cloud calls.
 *  Returns the status code and fills body_out. Handles chunked transfer
 *  encoding PROPERLY: the chunk-size markers must be stripped, or they end
 *  up embedded inside the JSON body and the parse fails on long replies.
 * =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
 * on timeout / closed-and-empty. Every read below goes through this, because
 * a reasoning model can think for many seconds between the response headers
 * and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
  uint32_t t0 = millis();
  while (!c.available()) {
    if (!c.connected()) return false;
    if (millis() - t0 > ms) return false;
    delay(10);
  }
  return true;
}

static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
  body_out = "";

  if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);

  bool chunked = false;
  while (waitData(client, idle_ms)) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;         // blank line = end of headers
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
  }

  if (chunked) {
    int blanks = 0;
    while (true) {
      /* The chunk-size line may not arrive for a long time while the model
       * reasons. Waiting here - instead of letting read() time out - is the
       * whole fix: a timed-out read looks exactly like "0" (final chunk),
       * which silently truncated the body to nothing. */
      if (!waitData(client, idle_ms)) {
        Serial.println("HTTP: timed out waiting for the next chunk");
        break;
      }
      String szline = client.readStringUntil('\n');
      szline.trim();
      if (szline.length() == 0) {                    // stray blank line
        if (++blanks > 4) break;
        continue;
      }
      blanks = 0;

      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;                            // genuine final chunk

      long got = 0;
      while (got < sz) {
        if (!waitData(client, idle_ms)) break;
        while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
      }
      if (waitData(client, 3000)) client.readStringUntil('\n');   // CRLF after chunk
      if (got < sz) { Serial.println("HTTP: short chunk"); break; }
    }
  } else {
    while (waitData(client, idle_ms))
      while (client.available()) body_out += (char)client.read();
  }
  return code;
}


/* ===========================================================================
 *  CLOUD CALL 1  —  Azure speech-to-text
 *  One POST, one header, plain WAV body. This is why Azure does the ears.
 * =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
  /* HTTPClient's one-shot POST fails on bodies this large (it attempts one
   * giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
   * function speaks HTTP directly and streams the WAV up in 4 KB chunks -
   * reliable, and if it ever stalls we know the exact byte it stopped at. */
  WiFiClientSecure client;
  client.setInsecure();                    // no cert bundle on-device; see notes
  client.setTimeout(15);                   // seconds, for reads

  writeWavHeader(wav_buf, audio_bytes);
  dumpWavToSD(audio_bytes);                // PC-playable copy of what we send
  size_t total = WAV_HEADER_LEN + audio_bytes;

  if (!client.connect(AZURE_STT_HOST, 443)) {
    snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
    Serial.println("STT: TLS connect failed");
    return false;
  }

  /* Two valid host forms use DIFFERENT URL paths - detect which one is in
   * secrets.h:  <resource>.cognitiveservices.azure.com -> /stt/speech/...
   *             <region>.stt.speech.microsoft.com      -> /speech/...     */
  bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
  String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
               "/speech/recognition/conversation/cognitiveservices/v1"
               "?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
               "Host: " AZURE_STT_HOST "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
               "Accept: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(total) + "\r\n\r\n";
  client.print(req);

  /* body, 4 KB at a time */
  size_t sent = 0;
  while (sent < total) {
    size_t n = min((size_t)4096, total - sent);
    size_t w = client.write(wav_buf + sent, n);
    if (w == 0) {
      delay(50);                           // brief stall - retry once
      w = client.write(wav_buf + sent, n);
      if (w == 0) {
        snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
                 (unsigned)(sent / 1024));
        Serial.printf("STT: upload stalled at %u/%u bytes\n",
                      (unsigned)sent, (unsigned)total);
        client.stop();
        return false;
      }
    }
    sent += w;
    yield();
  }
  Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);

  /* read the reply with proper de-chunking */
  String resp;
  int code = readHttpResponse(client, resp, 10000);
  client.stop();

  if (code != 200) {
    snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
    Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
    Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
    return false;
  }

  const char *status = doc["RecognitionStatus"];
  if (!status || strcmp(status, "Success") != 0) {
    /* The status names the exact failure:
     *   InitialSilenceTimeout = Azure heard silence (mic level too low)
     *   NoMatch               = heard sound but no recognisable words
     *   BabbleTimeout         = heard only noise                       */
    snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
    Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
    return false;
  }
  text_out = doc["DisplayText"].as<String>();
  if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
  return text_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 2  —  DeepSeek chat completion
 *  OpenAI-compatible format. Model name is deepseek-v4-flash - the old
 *  deepseek-chat name is dead, see secrets.h.
 * =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
  /* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
   * bodies (long answers + the model's hidden reasoning), which shows up as
   * "bad JSON reply". Reading until the server closes the connection is
   * reliable regardless of reply length. */
  JsonDocument req;
  req["model"] = DEEPSEEK_MODEL;
  req["max_tokens"] = LLM_MAX_TOKENS;
  JsonArray msgs = req["messages"].to<JsonArray>();
  JsonObject sys = msgs.add<JsonObject>();
  sys["role"] = "system";  sys["content"] = SYSTEM_PROMPT;
  JsonObject usr = msgs.add<JsonObject>();
  usr["role"] = "user";    usr["content"] = question;

  String body;
  serializeJson(req, body);

  WiFiClientSecure client;
  client.setInsecure();
  /* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
   * ~2 s, but anything that needs actual working-out (an Ohm's law problem,
   * say) can think for 10-30 s before sending a single byte. */
  client.setTimeout(60);

  if (!client.connect(DEEPSEEK_HOST, 443)) {
    snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
    Serial.println("LLM: TLS connect failed");
    return false;
  }

  client.print(String("POST /chat/completions HTTP/1.1\r\n"
               "Host: " DEEPSEEK_HOST "\r\n"
               "Authorization: Bearer " DEEPSEEK_KEY "\r\n"
               "Content-Type: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: ") + String(body.length()) + "\r\n\r\n");
  client.print(body);

  /* read the reply with proper de-chunking; generous window - long
   * questions make the model think for a while before it responds */
  String resp;
  int code = readHttpResponse(client, resp, 60000);
  client.stop();

  if (code != 200) {
    snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
    Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
    Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
    return false;
  }

  const char *content = doc["choices"][0]["message"]["content"];
  const char *finish  = doc["choices"][0]["finish_reason"];
  if (!content || !content[0]) {
    /* v4-flash is a reasoning model: if finish_reason is "length", the whole
     * token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
    if (finish && strcmp(finish, "length") == 0)
      snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
    else
      snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
    Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
    return false;
  }
  answer_out = String(content);
  answer_out.trim();
  if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
  return answer_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 3  —  Azure text-to-speech, streamed straight to the speaker
 *  We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
 *  the I2S peripheral eats. Skip the 44-byte header, forward the rest.
 *  No MP3 decoder, no audio library, no buffering the whole reply.
 * =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
  size_t got = 0;
  uint32_t t0 = millis();
  while (got < n && millis() - t0 < 10000) {
    int r = c.read(dst + got, n - got);
    if (r > 0) { got += r; t0 = millis(); }
    else if (!c.connected() && !c.available()) break;
    else delay(2);
  }
  return got;
}

bool azureTTSSpeak(const String &text) {
  // Escape the XML special characters for the SSML body
  String safe = text;
  safe.replace("&", "&amp;");
  safe.replace("<", "&lt;");
  safe.replace(">", "&gt;");

  String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
                "<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";

  /* Manual HTTP like the other two cloud calls - and for a hard reason:
   * Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
   * raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
   * into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
   * Here we parse the framing properly and keep only clean audio bytes. */
  WiFiClientSecure client;
  client.setInsecure();
  client.setTimeout(20);

  const char *host = AZURE_REGION ".tts.speech.microsoft.com";
  if (!client.connect(host, 443)) {
    Serial.println("TTS: TLS connect failed");
    return false;
  }

  client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
               "Host: ") + host + "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: application/ssml+xml\r\n"
               "X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
               "User-Agent: MaTouchRobojax\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(ssml.length()) + "\r\n\r\n");
  client.print(ssml);

  /* status + headers; note whether the body is chunked */
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
  bool chunked = false;
  long content_len = -1;
  while (client.connected() || client.available()) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
    if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
  }
  if (code != 200) {
    Serial.printf("TTS HTTP %d\n", code);
    client.stop();
    return false;
  }

  const size_t AUDIO_CAP = 1200 * 1024;    // ~37 s of speech
  uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
  if (!audio) { client.stop(); return false; }
  size_t alen = 0;

  if (chunked) {
    /* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
    while (true) {
      String szline = client.readStringUntil('\n');
      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;
      if (alen + sz > AUDIO_CAP) break;
      size_t got = readExact(client, audio + alen, sz);
      alen += got;
      client.readStringUntil('\n');        // trailing CRLF after each chunk
      if (got < (size_t)sz) break;
    }
  } else if (content_len > 0) {
    alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
  } else {
    /* no framing info: read until the server closes */
    uint32_t idle = millis();
    while ((client.connected() || client.available()) && millis() - idle < 5000) {
      int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
      if (r > 0) { alen += r; idle = millis(); }
      else delay(5);
    }
  }
  client.stop();
  Serial.printf("TTS: %u KB clean audio (%s), playing\n",
                (unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");

  bool ok = (alen > WAV_HEADER_LEN);
  if (ok) {
    /* NOW the audio actually starts - this is the honest moment to go green */
    LED_SPEAK();
    drawBar("SPEAKING...", gfx->color565(0, 130, 40));

    /* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
    static const uint8_t lead_in[640] = {0};             // 20 ms of silence
    size_t w = 0;
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
  }
  free(audio);

  // let the DMA buffers drain so the last word is not cut off
  delay(150);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
  return ok;
}


/* ===========================================================================
 *  SETUP
 * =========================================================================== */
void setup() {
  Serial.begin(115200);
  delay(400);
  Serial.println("\n=== 04 Voice Assistant  |  Robojax.com ===");
  Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");

  pinMode(TFT_BLK, OUTPUT);
  digitalWrite(TFT_BLK, LOW);
  pinMode(SD_CS, OUTPUT);
  digitalWrite(SD_CS, HIGH);

  // one shared SPI bus for TFT + SD (started before either device)
  SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);

  gfx->begin();
  gfx->fillScreen(BLACK);
  digitalWrite(TFT_BLK, HIGH);

  // SD is optional here - it only stores the /stt_debug.wav diagnostic copy
  ok_sd = SD.begin(SD_CS, SPI, 20000000);
  digitalWrite(SD_CS, HIGH);
  Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
                       : "SD not found - debug WAV dump disabled (not fatal)");

  bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
  delay(50);

  rgb.begin();
  rgb.setBrightness(LED_BRIGHTNESS);
  LED_IDLE();

  /* One recording buffer for the whole session, in PSRAM. This is the 8 MB
   * that makes the board worth buying. */
  wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
  if (!wav_buf) {
    gfx->setTextColor(RED);
    gfx->setTextSize(2);
    gfx->setCursor(10, 100);
    gfx->print("PSRAM alloc failed!");
    gfx->setTextSize(1);
    gfx->setCursor(10, 130);
    gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
    while (1) delay(1000);
  }

  micInit();
  spkInit();

  gfx->setTextSize(1);
  gfx->setTextColor(YELLOW);
  gfx->setCursor(4, 4);
  gfx->printf("Connecting to %s ...", WIFI_SSID);
  Serial.printf("Connecting to %s ", WIFI_SSID);

  WiFi.mode(WIFI_STA);
  WiFi.begin(WIFI_SSID, WIFI_PASS);
  uint32_t t0 = millis();
  while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
    delay(300);
    Serial.print(".");
  }
  Serial.println();

  clearChat();
  if (WiFi.status() == WL_CONNECTED) {
    Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
    chatBubble("Hold SPEAK and ask me anything.", false);
  } else {
    chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
    LED_ERROR();
  }
  drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}


/* ===========================================================================
 *  LOOP  —  one full conversation turn per button press
 * =========================================================================== */
void loop() {
  /* CLEAR button: edge-detected so one tap wipes once. Reading the panel
   * twice per loop (here and in speakButtonHeld) is fine - the GT911 just
   * reports its current state. */
  static bool tap_latch = false;
  static uint8_t tap_release = 0;
  if (state == ST_IDLE) {
    uint16_t cx, cy;
    if (getTouch(&cx, &cy)) {
      tap_release = 0;
      if (!tap_latch) {
        tap_latch = true;
        if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
          clearChat();
          chatBubble("Hold SPEAK and ask me anything.", false);
        }
      }
    } else if (tap_latch && ++tap_release >= 4) {
      tap_latch = false;
      tap_release = 0;
    }

    /* live WiFi signal indicator, refreshed every 2 s while idle */
    static uint32_t last_wifi = 0;
    if (millis() - last_wifi > 2000) {
      last_wifi = millis();
      drawWifi();
    }
  }

  if (state == ST_IDLE && speakButtonHeld()) {

    /* ---- record ---- */
    state = ST_RECORDING;
    LED_LISTEN();
    drawBar("LISTENING...", gfx->color565(0, 60, 200));
    uint32_t t_rec = millis();
    size_t audio_bytes = recordWhileHeld();
    t_rec = millis() - t_rec;
    Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);

    if (audio_bytes < SAMPLE_RATE / 2) {   // under a quarter second - a tap
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }

    /* ---- speech to text ---- */
    state = ST_STT;
    LED_THINK();
    drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
    uint32_t t_stt = millis();
    String question;
    if (!azureSTT(audio_bytes, question)) {
      /* Show the REAL cause on screen - no serial monitor needed. */
      char diag[96];
      snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
               g_stt_err, g_mic_peak_pct,
               ok_sd ? " | saved /stt_debug.wav" : "");
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_stt = millis() - t_stt;
    chatBubble(question.c_str(), true);
    Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());

    /* ---- think ---- */
    state = ST_LLM;
    drawBar("THINKING...", gfx->color565(150, 90, 0));
    uint32_t t_llm = millis();
    String answer;
    if (!deepseekChat(question, answer)) {
      char diag[96];
      snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_llm = millis() - t_llm;
    chatBubble(answer.c_str(), false);
    Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());

    /* ---- speak ----
     * Still amber here: the voice has to be synthesised and downloaded first
     * (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
     * at the exact moment audio starts coming out of the speaker. */
    state = ST_TTS;
    LED_THINK();
    drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
    uint32_t t_tts = millis();
    bool spoke = azureTTSSpeak(answer);
    t_tts = millis() - t_tts;

    /* Timing summary on serial - this feeds the "honest numbers" segment. */
    Serial.printf("TIMINGS  rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
                  audio_bytes / 32000.0, (unsigned long)t_stt,
                  (unsigned long)t_llm, (unsigned long)t_tts,
                  spoke ? "" : "  (TTS FAILED)");

    drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
    LED_IDLE();
    state = ST_IDLE;
  }

  delay(20);
}

مواردی که ممکن است به آن‌ها نیاز داشته باشید

منابع و مراجع

فایل‌ها📁

فایل مورد نیاز (.h)

  • secrets.h
    فایل برای ماجیول Makerfabs MaTouch AI ESP32S3 2.8" TFT Camera
    secrets.h 0.01 MB

سایر فایل‌ها

  • pins.h
    فایل پینها برای MaTouch AI ESP32S3 2.8" دوربین LCD صفحه لمسی.
    pins.h 0.01 MB

شماتیک

  • MaTouch_AI 2.8" MaTouch AI ESP32S3 2.8" TFT ST7789V شماتیک
    جدیدترین برد MaTouch AI دارای ورودی صوتی I2S / بلندگوی I2S / دوربین 3 میلیون پیکسلی OV3660 / نمایشگر با وضوح 320*240 است، با پردازنده قدرتمند ESP32S3 و قابلیت وایفای، این برد را به ابزار/پلتفرم خوبی برای توسعه هوش مصنوعی با ESP32 تبدیل میکند.
    MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF 0.15 MB