كود البحث

Makerfabs MaTouch ESP32-S3 2.8 بوصة كاميرا قم ببناء مساعد صوتي بالذكاء الاصطناعي على ESP32-S3 (Azure + DeepSeek)

Makerfabs MaTouch ESP32-S3 2.8 بوصة كاميرا قم ببناء مساعد صوتي بالذكاء الاصطناعي على ESP32-S3 (Azure + DeepSeek)

اضغط مع الاستمرار على زر، اطرح سؤالاً، وستجيب اللوحة بصوت عالٍ

مساعد صوتي كامل على لوحة صغيرة واحدة. اضغط مع الاستمرار على زر SPEAK واطرح سؤالاً. تسجّل اللوحة صوتك عبر الميكروفونين، وترسل الصوت إلى Microsoft Azure لتحويله إلى نص، ثم ترسل النص إلى DeepSeek للتفكير فيه، ثم ترسل الإجابة مرة أخرى إلى Azure لتحويلها إلى كلام، وتشغّلها عبر مكبر الصوت الخاص بها. يظهر الحوار كاملاً على الشاشة على شكل فقاعات محادثة.

مساعد صوتي ذكي يتحدث يعمل على لوحة MaTouch AI ESP32-S3 مساعد صوتي ذكي يتحدث يعمل على لوحة MaTouch AI ESP32-S3

ماذا يحدث من لحظة الضغط على الزر

هذه هي الرحلة الكاملة لسؤال واحد، خطوة بخطوة. من الجدير قراءتها مرة واحدة، لأن كل ما تراه على الشاشة وعلى مؤشر LED يطابق إحدى هذه المراحل.

  1. تضغط مع الاستمرار على زر SPEAK. إنه وضع اضغط وتحدث، وليس اضغط ثم تحدث: يستمر التسجيل طوال مدة بقاء إصبعك على الزر، حتى ست ثوانٍ. يتحول مؤشر LED إلى اللون الأزرق ويظهر على الزر نص LISTENING.

  2. يسجّل كلا الميكروفونين صوتك. تلتقط اللوحة 16,000 عينة في الثانية من الزوج المجسّمي، وتدمج القناتين في قناة واحدة، وتطبّق قدراً بسيطاً من تضخيم الإشارة، وتخزّن النتيجة في ذاكرة PSRAM. يتحرك شريط تقدم عبر الزر أثناء تحدثك. ثانيتان من الكلام تعادل حوالي 64 كيلوبايت.

  3. تحرّر الزر. يتوقف التسجيل. تكتب اللوحة ترويسة WAV من 44 بايت في مقدمة الصوت - هذه التسمية الصغيرة هي كل ما يحوّل العينات الخام إلى ملف سيقبله Azure.

  4. يُرسل الصوت إلى Azure لتحويل الكلام إلى نص. يتم رفعه على شكل أجزاء بحجم 4 كيلوبايت عبر اتصال آمن، ويعود كنص من سطر واحد. أصبحت 64 كيلوبايت من الصوت حوالي 25 بايتاً من الكتابة. يتحول مؤشر LED إلى اللون الكهرماني.

  5. يظهر سؤالك على الشاشة كفقاعة محادثة زرقاء، لتتمكن من رؤية ما سمعه بالضبط - وهذا مفيد، لأن الكلمات المفهومة خطأً تفسّر معظم الإجابات الغريبة.

  6. يُرسل النص إلى DeepSeek. ترسل اللوحة سؤالك مع تعليمات ثابتة بإبقاء الردود في جملتين قصيرتين. يفكّر النموذج - يفكر فعلياً، إنه نموذج استدلالي - ويعيد إجابة.

  7. تظهر الإجابة على الشاشة كفقاعة رمادية. يمكنك قراءتها قبل سماعها.

  8. تُعاد الإجابة إلى Azure ليتم نطقها. تطلب اللوحة صوت PCM خاماً بتردد 16 كيلوهرتز، وهو بالضبط التنسيق الذي يريده مضخم الصوت الخاص بها، لذلك لا يوجد أي مفكك ترميز MP3 في هذا المشروع. يظهر الآن على الزر نص GETTING VOICE ويبقى مؤشر LED كهرمانياً، لأن شيئاً غير مسموع بعد.

  9. يتم تنزيل المقطع كاملاً إلى ذاكرة PSRAM قبل تشغيل أي عينة واحدة. هذا مهم - انظر الملاحظة أدناه.

  10. التشغيل. في اللحظة التي يُسلَّم فيها الصوت إلى مكبر الصوت، يتحول مؤشر LED إلى اللون الأخضر ويظهر على الزر نص SPEAKING. تسمع الإجابة.

لماذا يتم تنزيل الصوت أولاً بدلاً من تشغيله أثناء وصوله. البث المباشر من الشبكة إلى مكبر الصوت يبدو كصوت طرق. سعة المخزن المؤقت لمكبر الصوت لا تتجاوز حوالي عُشر ثانية، وأي توقف في نقل WiFi أطول من ذلك يفرّغه، مما ينتج صوت طرق مسموعاً. تنزيل الرد كاملاً إلى ذاكرة PSRAM أولاً يكلف حوالي ثانية إضافية من الانتظار ويزيل كل الفجوات. لهذا السبب أيضاً تعرض الشاشة GETTING VOICE قبل أن تعرض SPEAKING - فالمرحلتان مختلفتان بصدق.

كم يستغرق كل مرحلة من الوقت

تم قياسه على أجهزة حقيقية، لسؤال بسيط:

المرحلة

الوقت النموذجي

التسجيل

طوال مدة الضغط على الزر

تحويل الكلام إلى نص (Azure)

حوالي 1.8 ثانية

التفكير (DeepSeek)

حوالي 1.8 ثانية لسؤال بسيط، وأطول بكثير لسؤال يتطلب حسابات فعلية

جلب الصوت (Azure)

حوالي 7 ثوانٍ - أكبر شريحة زمنية

الإجمالي، من التحرير إلى أول صوت

حوالي 11 ثانية

كل تبادل يطبع توقيته الخاص على شاشة المراقبة التسلسلية، لذا يمكنك قياس توقيتاتك الخاصة بدلاً من الوثوق بهذه الأرقام. إذا أردت تسريع العملية، فإن التغيير الأكثر فعالية هو طلب ردود أقصر في SYSTEM_PROMPT - نص أقل للنطق يعني صوتاً أقل لتوليده وتنزيله.

اللوحة نفسها لا تفهم أي شيء أبداً. إنها رسول بآذان جيدة وصوت جيد - الذكاء مستأجر بالثانية.

لماذا هذه الخدمات الثلاث

  • Azure يتعامل مع الكلام داخلاً وخارجاً. يمكن لتحويل النص إلى كلام الخاص به إرجاع صوت PCM خام بتردد 16 كيلوهرتز، وهو بالضبط ما يريده رقاقة مكبر الصوت، لذلك لا يوجد أي مفكك ترميز MP3 في هذا المشروع. وتحويل الكلام إلى نص الخاص به يقبل ملف WAV بسيطاً عبر طلب POST بسيط.

  • DeepSeek هو عقل المحادثة. إنه سريع وتكلفته جزء صغير من السنت لكل رد.

  • OpenAI غير مستخدم هنا - انظر المشروع 05، حيث يقوم بعمل الرؤية.

تغيرت أسماء نماذج DeepSeek. تم إيقاف الاسمين القديمين deepseek-chat وdeepseek-reasoner في يوليو 2026. معظم الدروس التعليمية على الإنترنت ما زالت تستخدمهما وستعيد خطأً. الأسماء الحالية هي deepseek-v4-flash وdeepseek-v4-pro. يستخدم هذا المشروع v4-flash.

فخ نموذج الاستدلال

يفكّر DeepSeek v4-flash قبل أن يجيب، وهذا التفكير يُحتسب ضمن حد الرموز الخاص بك. إذا ضبطت LLM_MAX_TOKENS على قيمة منخفضة جدًا، فسيُستهلك الميزان بأكمله في التفكير، ويعود الجواب فارغًا، ولا تقول اللوحة شيئًا. وهو مضبوط على 400 هنا لهذا السبب. الأسئلة الصعبة تستغرق وقتًا أطول أيضًا - فالحقيقة البسيطة تُجاب في حوالي ثانيتين، أما السؤال الذي يتطلب حلًا فعليًا فقد يستغرق وقتًا أطول بكثير.

قراءة ضوء الحالة

اللون

المعنى

أزرق

يستمع إليك

كهرماني

السحابة تفكّر، أو يتم جلب الصوت

أخضر

يتحدث - يتحول إلى الأخضر في اللحظة التي يبدأ فيها الصوت بالضبط

أحمر

حدث خطأ ما - تحقق من المراقب التسلسلي

عناصر التحكم على الشاشة

يتدفق الدردشة مثل محادثة هاتفية، حيث تتحرك الرسائل الأقدم للأعلى وتختفي. CLEAR يمسحها. يوجد مقياس إشارة WiFi مع قراءة dBm الفعلية في الزاوية، وهو مفيد عندما تتساءل عما إذا كان الرد البطيء بسبب الشبكة أم الخدمة.

حول لوحة MaTouch AI ESP32-S3 2.8"

كل مشروع في هذه الصفحة يعمل على MaTouch AI ESP32-S3 2.8" TFT ST7789V من Makerfabs. إنها لوحة متكاملة: شاشة لمس ملونة، كاميرا بدقة 3 ميجابكسل، ميكروفونان ومضخم صوت حقيقي، وكلها مدفوعة بوحدة ESP32-S3 مع ذاكرة PSRAM بسعة 8 ميجابايت. هذا المزيج هو ما يجعل مشاريع الذكاء الاصطناعي هذه ممكنة على لوحة واحدة دون أي شيء آخر متصل بها.

ذاكرة PSRAM بسعة 8 ميجابايت أهم من أي رقم آخر هنا. إنها ما يسمح للوحة باحتواء إطار كاميرا، أو بضع ثوانٍ من الصوت المسجل، أو صورة مشفرة بـ base64 في الذاكرة في نفس الوقت - ولا شيء من ذلك يتسع في ذاكرة ESP32 العادية.

وثائق الشركة المصنعة: صفحة Makerfabs wiki.

المواصفات الرئيسية

  • المعالج: ESP32-S3، ثنائي النواة بتردد 240 ميجاهرتز، WiFi 2.4 جيجاهرتز + Bluetooth 5.0

  • الذاكرة: فلاش بسعة 16 ميجابايت، PSRAM بسعة 8 ميجابايت (مطلوبة تقريبًا لكل مشروع هنا)

  • الشاشة: IPS بحجم 2.8 بوصة، دقة 320×240، مشغل ST7789V، SPI

  • اللمس: GT911 سعوي، يتتبع 5 أصابع في وقت واحد

  • الكاميرا: OV3660، بدقة 3 ميجابكسل، حتى 2048×1536

  • الميكروفونات: ميكروفونان رقميان INMP441 I2S (زوج استريو حقيقي)

  • مكبر الصوت: مضخم فئة D MAX98357A، بقوة 3.2 واط عند 4 أوم

  • التخزين: فتحة بطاقة microSD (وضع SPI)

  • الطاقة: USB-C، موصل بطارية JST، شاحن TP4056، مفتاح طاقة

  • أيضًا على اللوحة: LED RGB WS2812B، ساعة زمنية حقيقية PCF8563T مدعومة ببطارية، ومقياس شحن بطارية MAX17048 غير مدرج في المواصفات الرسمية

منفذا USB-C ليسا متطابقين. يشارك مكبر صوت اللوحة دبابيس الإشارة الخاصة به (IO19 وIO20) مع منفذ USB الأصلي، لأن تلك الدبابيس هي خطوط بيانات USB المثبتة في ESP32-S3. قم دائمًا بالرفع والتشغيل عبر منفذ USB-C CH340K (الموجود بجانب زر RESET)، واضبط USB CDC On Boot على Disabled. استخدم المنفذ الخاطئ وسيتعطل الصوت أو ستفشل عمليات الرفع.

إعدادات Arduino IDE

هذه الإعدادات مهمة. معظم المشكلات التي يبلغ عنها الأشخاص مع هذه اللوحة هي أحد هذه الإعدادات الخاطئة، وهي تُعاد ضبطها عند تغيير إصدار النواة، لذا تحقق منها مرة أخرى بعد أي تغيير.

الإعداد

القيمة

اللوحة

ESP32S3 Dev Module

إصدار نواة ESP32

2.0.17

PSRAM

OPI PSRAM

حجم الفلاش

16MB (128Mb)

مخطط التقسيم

16M Flash (3MB APP/9.9MB FATFS)

USB CDC On Boot

Disabled

سرعة الرفع

921600

مسح كل الفلاش قبل الرفع

Disabled

المنفذ

منفذ USB-C CH340K

استخدم نواة ESP32 2.0.17، وليس 3.x. أزالت Espressif نماذج اكتشاف الوجه على الجهاز في النواة 3، لذا لن تُترجم مشاريع الوجه هناك. تثبيت 2.0.17 يبقي كل مشروع في هذه الصفحة يعمل بإعداد واحد. في Boards Manager، تتيح لك القائمة المنسدلة للإصدارات التبديل ذهابًا وإيابًا متى شئت.

استخدم GFX Library for Arduino الإصدار 1.5.6، وليس 1.6.x. إصدارات 1.6 مبنية لنواة ESP32 3 ويمكن أن تعلق عند بدء التشغيل على النواة 2.0.17. إذا بقيت شاشتك سوداء بعد الرفع، فهذا أول شيء يجب التحقق منه.

المكتبات المطلوبة

ثبّت هذه عبر Tools → Manage Libraries في Arduino IDE. أرقام الإصدارات مهمة - يرجى استخدام تلك المدرجة.

المكتبة

الإصدار

المؤلف

GFX Library for Arduino

1.5.6

moononournation

bb_captouch

1.3.1

Larry Bank

ArduinoJson

7.x

Benoit Blanchon

Adafruit NeoPixel

أي إصدار حديث

Adafruit

إعداد secrets.h

تفاصيل الواي فاي الخاصة بك وأي مفاتيح API توضع في secrets.h، وهو مضمّن في ملف التحميل بقيم افتراضية. افتح تلك العلامة التبويبية في بيئة التطوير المتكاملة Arduino IDE واستبدلها بقيمك الخاصة.

يجب أن يكون الواي فاي بتردد 2.4 جيجاهرتز. لا يمكن لـ ESP32-S3 رؤية شبكة بتردد 5 جيجاهرتز على الإطلاق. إذا كان جهاز التوجيه الخاص بك يجمع بين النطاقين تحت اسم واحد (تسميه Asus Smart Connect)، فإما أن تقوم بإيقاف تشغيل ذلك أو تعطي نطاق 2.4 جيجاهرتز اسمًا خاصًا به واستخدم ذلك في secrets.h.

الحصول على مفاتيح API الخاصة بك

يتواصل هذا المشروع مع خدمة ذكاء اصطناعي سحابية، لذا تحتاج إلى مفتاح خاص بك. إذا لم تقم بذلك من قبل، فلا تقلق - إنها نفس فكرة كلمة المرور التي تحدد حسابك للخدمة. يستغرق الأمر بضع دقائق، مرة واحدة.

المفتاح ليس اشتراكًا في موقع ويب. على سبيل المثال، الدفع مقابل ChatGPT Plus لا يمنحك مفتاح API - فهما منتجان منفصلان بفوترة منفصلة. تحتاج إلى حساب على منصة المطورين، كما هو موضح أدناه.

Microsoft Azure Speech - للاستماع والتحدث

يحول Azure كلامك إلى نص ويحول الإجابة مرة أخرى إلى صوت. الطبقة المجانية سخية بما يكفي لكل شيء في هذه الصفحة.

  1. انتقل إلى portal.azure.com وسجّل الدخول بحساب Microsoft (الحساب المجاني جيد).

  2. إذا لم تستخدم Azure من قبل، فسترى شاشة مرحبًا بك في Azure تقدم ثلاثة خيارات. اختر ابدأ بالتجربة المجانية لـ Azure - تحتاج إلى اشتراك قبل أن يسمح لك Azure بإنشاء أي شيء. (يجب على الطلاب اختيار Azure for Students بدلاً من ذلك: نفس النتيجة، بدون بطاقة مطلوبة.) تجاهل إدارة Microsoft Entra ID، وهو شيء مختلف تمامًا.

  3. انقر فوق إنشاء مورد، وابحث عن Speech، واختر خدمة Speech المنشورة بواسطة Microsoft.

  4. املأ النموذج: أي مجموعة موارد، وأي اسم، واختر منطقة قريبة منك - اكتب تلك المنطقة تمامًا كما تظهر، على سبيل المثال eastus.

  5. بالنسبة لـ طبقة التسعير اختر F0 (مجاني). يسمح ذلك بحوالي خمس ساعات من تحويل الكلام إلى نص ونصف مليون حرف من تحويل النص إلى كلام كل شهر.

  6. انقر فوق مراجعة + إنشاء، ثم إنشاء. انتظر حوالي دقيقة، ثم انقر فوق الانتقال إلى المورد.

  7. في القائمة اليسرى، افتح المفاتيح ونقطة النهاية. انسخ KEY 1 والموقع/المنطقة.

ضع تلك في secrets.h كـ AZURE_SPEECH_KEY و AZURE_REGION. بالنسبة لـ AZURE_STT_HOST، استخدم <region>.stt.speech.microsoft.com - لذا مع المنطقة eastus يكون ذلك eastus.stt.speech.microsoft.com.

بخصوص بطاقة الائتمان. تطلب التجربة المجانية لـ Azure بطاقة للتحقق من هويتك. لا تفرض عليك رسومًا. تحصل على رصيد بقيمة 200 دولار لمدة 30 يومًا، وبعد ذلك ينتقل الحساب إلى الدفع حسب الاستخدام - لكن طبقة F0 Speech تبقى مجانية، شهرًا بعد شهر، وكل شيء في هذه المشاريع يناسبها بشكل مريح. إذا كنت تفضل عدم تقديم بطاقة على الإطلاق وكنت طالبًا، فإن خيار Azure for Students يمنحك رصيدًا بدونها.

يجب أن يكون مورد "خدمة Speech". المفتاح من مورد Translator أو Language أو Cognitive Services العام يبدو متطابقًا وهو صالح تمامًا - لكن كل طلب كلام يُرجع خطأ 401. هذا ما أمسك بنا أثناء الاختبار وكلفنا ساعة. إذا فشل الكلام مع 401 بينما يبدو المفتاح صحيحًا، تحقق من نوع المورد الذي أنشأته.

DeepSeek - الجزء المفكر

DeepSeek هو نموذج اللغة الذي يجيب فعليًا على سؤالك. إنه غير مكلف - بضعة دولارات من الرصيد تغطي آلاف الردود.

  1. انتقل إلى platform.deepseek.com وأنشئ حسابًا.

  2. افتح مفاتيح API في القائمة وانقر فوق إنشاء مفتاح API جديد.

  3. انسخه فورًا. يتم عرضه مرة واحدة فقط ولا يُعرض مرة أخرى أبدًا - إذا فقدته، احذف هذا المفتاح وأنشئ آخر.

  4. أضف مبلغًا صغيرًا من الرصيد تحت الشحن. لا توجد طبقة مجانية، لكن أصغر شحنة تدوم لفترة طويلة جدًا مع هذا الاستخدام.

ضع المفتاح في secrets.h كـ DEEPSEEK_KEY. يبدأ بـ sk-.

تغيرت أسماء النماذج في يوليو 2026. تم إيقاف deepseek-chat و deepseek-reasoner القديمين، لذا فإن معظم الدروس التعليمية التي تجدها عبر الإنترنت ستفشل مع خطأ 400. استخدم deepseek-v4-flash، وهو ما تم ضبطه بالفعل في هذه المشاريع.

كم تكلفة التشغيل

قليلة جدًا، لكنها ليست مجانية، ويجب أن تعرف تقريبًا ما تنفقه قبل ترك مشروع يعمل.

الخدمة

التكلفة التقريبية

Azure Speech

الطبقة المجانية تغطي حوالي 5 ساعات من الاستماع و0.5 مليون حرف من التحدث شهريًا

DeepSeek

جزء من السنت لكل إجابة - آلاف الردود مقابل بضعة دولارات

OpenAI vision

حوالي سنت أو سنتان لكل صورة، حسب النموذج

تتغير الأسعار، لذا تعامل مع هذه كدليل وليس كعرض سعر. كل واحدة من هذه الخدمات لديها صفحة استخدام حيث يمكنك مراقبة ما أنفقته، وكلها تتيح لك تعيين حد إنفاق - وهو أمر يستحق القيام به من اليوم الأول.

حافظ على مفاتيحك خاصة. أي شخص يمتلكها يمكنه إنفاق أموالك. لا تضعها في فيديو أو لقطة شاشة أو منشور في منتدى أو مستودع كود عام. إذا تم كشف مفتاح في أي وقت، احذفه على موقع المزود وأنشئ مفتاحًا جديدًا - يستغرق ذلك ثوانٍ، وهو الحل الحقيقي الوحيد.

استكشاف الأخطاء وإصلاحها

العَرَض

السبب والحل

تبقى الشاشة سوداء

إصدار مكتبة GFX خاطئ (استخدم 1.5.6) أو إعدادات اللوحة خاطئة.

فشل تخصيص PSRAM أو خطأ الكاميرا 0xffffffff

أدوات ← PSRAM غير مضبوط على OPI PSRAM.

لا يتم رفع أي شيء / لا يوجد منفذ COM

منفذ USB-C خاطئ، أو برنامج تشغيل CH340 غير مثبت.

تفشل الكاميرا ولا تتعافى أبدًا

خط إعادة تعيين الكاميرا مرتبط بزر RESET على اللوحة، لذا لا يمكن للبرنامج إعادة تشغيلها. اضغط على RESET. إذا استمر الفشل، أعد تركيب كابل الشريط الخاص بالكاميرا.

تنزيل الكود

مخطط Arduino الكامل لهذا المشروع، مع pins.h وكل ما يحتاجه الآخر، متاح للتنزيل مجانًا.

تنزيل 04_Voice_Assistant

فك ضغطه، وافتح ملف .ino في بيئة Arduino IDE، وتحقق من الإعدادات أعلاه، وارفعه عبر منفذ CH340K USB-C.

884-Arduin code for MaTouch AI ESP32S3 2.8in AI Camera: AI voice assistant
اللغة: C++
/*
 * ===========================================================================
 *  04_Voice_Assistant  —  MaTouch AI ESP32-S3 2.8" TFT ST7789V
 * ===========================================================================
 *
 ----------
 *  ROBOJAX.COM  -  MaTouch AI ESP32-S3 2.8" project series
 *
 *    WATCH THE VIDEO
 *        https://youtu.be/6AL3g3tC_Hk
 *
 *    WRITTEN TUTORIALS - every project, with photos and full explanation
 *        Camera and touchscreen.... https://robojax.com/RTJ849
 *        Offline face recognition.. https://robojax.com/RTJ850
 *        AI voice assistant........ https://robojax.com/RTJ851
 *        AI vision................. https://robojax.com/RTJ852
 *
 *    GET THE BOARD - SAVE $5 with coupon code:  Robojax_Makerfab
 *        https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
 *        (enter the code at checkout)
 *
 *  All of this code is free. If it helped you, a subscribe on YouTube is
 *  the best way to support more of it.
 *  
 *  ---------------------------------------------------------------------------
 *
 *  A complete voice assistant on a $40 board:
 *
 *      hold SPEAK  ->  both INMP441 microphones record you
 *                  ->  Azure Speech turns the audio into text
 *                  ->  DeepSeek v4-flash thinks of an answer
 *                  ->  Azure Speech turns the answer into audio
 *                  ->  the MAX98357 speaker says it out loud
 *
 *  and the whole conversation is drawn as chat bubbles on the touchscreen.
 *
 *  WHY THIS COMBINATION OF SERVICES (each is used where it is best):
 *   - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
 *     transcription endpoint wants multipart/form-data - miserable on an MCU.
 *   - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
 *     which streams straight into the I2S speaker with NO MP3 decoder at all.
 *   - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
 *     model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
 *     Most tutorials online still use them and are broken. See secrets.h.
 *
 *  A detail the vendor examples get wrong: this board has TWO microphones on
 *  one I2S bus (left + right), but every Makerfabs demo records left-only and
 *  throws one away. This sketch records both and averages them.
 *
 *  ---------------------------------------------------------------------------
 *  *** WHICH USB PORT - THIS MATTERS ***
 *  The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
 *  through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
 *  If you use the wrong port the audio will be garbage or uploads will fail.
 *  ---------------------------------------------------------------------------
 *
 *  FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
 *
 *  BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
 *  or compile errors. These reset when you switch cores - recheck them!)
 *
 *      Board            : ESP32S3 Dev Module
 *      ESP32 core       : 2.0.17
 *      PSRAM            : OPI PSRAM        <-- required, audio buffer lives there
 *      Flash Size       : 16MB (128Mb)
 *      Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
 *      USB CDC On Boot  : Disabled         <-- required, see USB note above
 *      Upload Speed     : 921600
 *      Port             : the CH340K USB-C port (the one near RESET)
 *
 *  LIBRARIES
 *      GFX Library for Arduino   v1.5.6   (NOT 1.6.x - that pairs with core 3)
 *      bb_captouch               v1.3.1
 *      ArduinoJson               v7.x
 *      Adafruit NeoPixel         any recent
 *
 *  ---------------------------------------------------------------------------
 *  FUNCTIONS IN THIS SKETCH
 *      led(r,g,b) + LED_* macros   RGB status colours (blue/amber/green/red)
 *      getTouch(&x,&y)      read the touch panel, mapped to screen coordinates
 *      speakButtonHeld()    true while a finger is on the SPEAK button
 *      bubbleLines(t)       how many lines a message wraps to
 *      drawOneBubble(m,y)   draw a single chat bubble
 *      redrawChat()         rebuild the chat area from history, newest at bottom
 *      clearChat()          wipe the chat history (CLEAR button)
 *      chatBubble(t,user)   add a message to history and redraw
 *      drawWifi()           WiFi signal bars + dBm readout
 *      drawBar(label,col)   bottom bar: SPEAK button + CLEAR + WiFi meter
 *      micInit()            I2S input - BOTH INMP441 mics, stereo
 *      spkInit()            I2S output - MAX98357 speaker
 *      recordWhileHeld()    record while SPEAK held, downmix stereo->mono
 *      writeWavHeader(...)  prepend the 44-byte RIFF/WAVE header
 *      dumpWavToSD(...)     save the exact upload to SD (/stt_debug.wav)
 *      readHttpResponse()   read an HTTPS reply, de-chunking it properly
 *      azureSTT(...)        chunked upload of the WAV -> recognised text
 *      deepseekChat(...)    question -> deepseek-v4-flash -> answer text
 *      azureTTSSpeak(text)  answer -> Azure voice -> PSRAM -> speaker
 *      setup() / loop()     boot + WiFi / one conversation turn per press
 *
 *  Robojax.com
 * ===========================================================================
 */

#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"

/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE     16000
#define RECORD_MAX_S    6                       // hard cap on one question
#define WAV_HEADER_LEN  44
#define REC_BUF_BYTES   (SAMPLE_RATE * RECORD_MAX_S * 2)   // 16-bit mono

/* Software gain applied to the recording. The INMP441 capture is quiet at
 * 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
 * speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
 * lower it. */
#define MIC_GAIN        6

/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
 * 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
 *     back silent or garbled in stereo mode. */
#define USE_BOTH_MICS   1

/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
 * uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS  25

/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
 * shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
    TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);

BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);

/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr;      // WAV_HEADER_LEN + up to REC_BUF_BYTES

/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;

/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool  ok_sd = false;
float g_mic_peak_pct = 0;          // last recording's raw peak, % of full scale
char  g_stt_err[64] = "";          // last STT failure cause, verbatim
char  g_llm_err[64] = "";          // last DeepSeek failure cause, verbatim

/* --- layout --------------------------------------------------------------- */
#define CHAT_H     200              // chat area: y 0..199
#define BAR_Y      202              // button bar below it
#define BTN_SPEAK_X   4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W  58
#define WIFI_X      236             // signal indicator, right end of the bar
#define BTN_H        36

/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
 * This is what prevents new text printing over old - the whole area is
 * rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int     chat_count = 0;

/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
  rgb.setPixelColor(0, rgb.Color(r, g, b));
  rgb.show();
}
#define LED_IDLE()    led(0, 0, 0)
#define LED_LISTEN()  led(0, 60, 255)     // blue   - recording
#define LED_THINK()   led(255, 120, 0)    // amber  - waiting on the cloud
#define LED_SPEAK()   led(0, 255, 40)     // green  - talking
#define LED_ERROR()   led(255, 0, 0)      // red


/* ===========================================================================
 *  Touch
 * =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
  TOUCHINFO ti;
  if (!bbct.getSamples(&ti)) return false;
  if (ti.count < 1) return false;
  *x = ti.y[0];
  *y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
  return true;
}

bool speakButtonHeld() {
  uint16_t x, y;
  if (!getTouch(&x, &y)) return false;
  return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}


/* ===========================================================================
 *  Chat UI  —  word-wrapped bubbles, user right/blue, assistant left/grey
 * =========================================================================== */
#define CHAT_CHARS 42                          // chars per line at textsize 1

static int bubbleLines(const char *t) {
  int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
  return l < 1 ? 1 : l;
}

void drawOneBubble(const ChatMsg &m, int y) {
  int len = strlen(m.text);
  int lines = bubbleLines(m.text);
  int h = lines * 10 + 8;
  uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
  int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
  if (w < 30) w = 30;
  int x = m.from_user ? (316 - w) : 4;

  gfx->fillRoundRect(x, y, w, h, 5, bg);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  for (int i = 0; i < lines; i++) {
    char line[CHAT_CHARS + 1] = {0};
    strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
    gfx->setCursor(x + 5, y + 5 + i * 10);
    gfx->print(line);
  }
}

/* Rebuild the whole chat area from history: newest message anchored at the
 * bottom, older ones stacked upward until the area is full. */
void redrawChat() {
  gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
  int shown = min(chat_count, CHAT_HISTORY);
  int y = CHAT_H - 2;
  for (int i = 0; i < shown; i++) {
    ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
    int h = bubbleLines(m.text) * 10 + 8;
    y -= h;
    if (y < 0) break;                          // area full - older ones drop off
    drawOneBubble(m, y);
    y -= 4;
  }
}

void clearChat() {
  chat_count = 0;
  redrawChat();
}

void chatBubble(const char *text, bool from_user) {
  ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
  strncpy(m.text, text, sizeof(m.text) - 1);
  m.text[sizeof(m.text) - 1] = 0;
  m.from_user = from_user;
  chat_count++;
  redrawChat();
}

/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
  gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
  bool up = (WiFi.status() == WL_CONNECTED);
  long rssi = up ? WiFi.RSSI() : -100;
  // -55 dBm or better = full bars; each 10 dB drops one
  int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;

  for (int b = 0; b < 4; b++) {
    int bh = 6 + b * 6;                       // heights 6,12,18,24
    uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
    gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
  }
  gfx->setTextSize(1);
  gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
  if (up) {
    gfx->setTextColor(CYAN);
    gfx->printf("%lddBm", rssi);
  } else {
    gfx->setTextColor(RED);
    gfx->print("DOWN");
  }
  gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
  gfx->setTextColor(gfx->color565(120, 120, 120));
  gfx->print("WiFi");
}

void drawBar(const char *label, uint16_t colour) {
  gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
  gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
  gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
  gfx->setTextSize(2);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
  gfx->print(label);

  // CLEAR wipes the chat history
  gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
  gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
  gfx->print("CLEAR");

  drawWifi();
}


/* ===========================================================================
 *  I2S  —  microphones on port 0, speaker on port 1. Separate hardware
 *  ports, so recording and playback can never fight over a bus.
 * =========================================================================== */
void micInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    /* BOTH channels - this is the two-microphone fix. The vendor examples
     * use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
     * falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
    .channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = false,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_MIC_SCK,
    .ws_io_num = I2S_MIC_WS,
    .data_out_num = I2S_PIN_NO_CHANGE,
    .data_in_num = I2S_MIC_SD
  };
  i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_MIC_PORT, &pins);
}

void spkInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,   // mono - the amp downmixes anyway
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = true,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_SPK_BCLK,
    .ws_io_num = I2S_SPK_LRC,
    .data_out_num = I2S_SPK_DOUT,
    .data_in_num = I2S_PIN_NO_CHANGE
  };
  i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_SPK_PORT, &pins);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
}


/* ===========================================================================
 *  Recording  —  runs while the SPEAK button is held (up to RECORD_MAX_S).
 *  Reads stereo pairs, averages L+R into one mono stream, applies a little
 *  software gain, and fills wav_buf after the 44-byte header slot.
 *  Returns the number of audio bytes recorded.
 * =========================================================================== */
size_t recordWhileHeld() {
  int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
  size_t   mono_samples = 0;
  const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;

  int16_t chunk[512];
  uint32_t last_touch_ok = millis();
  int32_t  peak = 0;                        // loudest raw sample - mic health check

  i2s_zero_dma_buffer(I2S_MIC_PORT);

  while (mono_samples < max_samples) {
    size_t got = 0;
    i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);

#if USE_BOTH_MICS
    size_t n = got / 4;                     // 4 bytes = one L+R pair
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
    size_t n = got / 2;                     // 2 bytes = one mono sample
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = chunk[i];
#endif
      if (abs(raw) > peak) peak = abs(raw);
      int32_t mixed = raw * MIC_GAIN;
      if (mixed >  32767) mixed =  32767;
      if (mixed < -32768) mixed = -32768;
      mono[mono_samples++] = (int16_t)mixed;
    }

    /* The GT911 is polled between I2S reads. A 250 ms grace period stops a
     * momentary missed touch sample from cutting the recording short. */
    if (speakButtonHeld()) last_touch_ok = millis();
    else if (millis() - last_touch_ok > 250) break;

    // live progress on the button
    static uint32_t last_draw = 0;
    if (millis() - last_draw > 200) {
      last_draw = millis();
      gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
                    (int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
    }
  }

  /* Mic health line: peak as % of full scale BEFORE gain.
   *   0%       = the mic is not being read at all (config/pin problem)
   *   under 3% = too quiet - speak closer or raise MIC_GAIN
   *   3-40%    = healthy speech level
   */
  g_mic_peak_pct = peak * 100.0 / 32768.0;
  Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
                g_mic_peak_pct, MIC_GAIN,
                peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");

  return mono_samples * 2;
}

/* Dump the exact WAV we are about to POST onto the SD card, so it can be
 * played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
  if (!ok_sd) return;
  SD.remove("/stt_debug.wav");
  File f = SD.open("/stt_debug.wav", FILE_WRITE);
  if (!f) return;
  f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
  f.close();
  Serial.println("debug copy saved to SD as /stt_debug.wav");
}

/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
  uint32_t file_len = data_bytes + 36;
  uint32_t byte_rate = SAMPLE_RATE * 2;
  memcpy(h, "RIFF", 4);           memcpy(h + 4, &file_len, 4);
  memcpy(h + 8, "WAVEfmt ", 8);
  uint32_t fmt_len = 16;          memcpy(h + 16, &fmt_len, 4);
  uint16_t fmt = 1, ch = 1;       memcpy(h + 20, &fmt, 2);   memcpy(h + 22, &ch, 2);
  uint32_t rate = SAMPLE_RATE;    memcpy(h + 24, &rate, 4);  memcpy(h + 28, &byte_rate, 4);
  uint16_t align = 2, bits = 16;  memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
  memcpy(h + 36, "data", 4);      memcpy(h + 40, &data_bytes, 4);
}


/* ===========================================================================
 *  HTTP response reader  —  shared by all three cloud calls.
 *  Returns the status code and fills body_out. Handles chunked transfer
 *  encoding PROPERLY: the chunk-size markers must be stripped, or they end
 *  up embedded inside the JSON body and the parse fails on long replies.
 * =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
 * on timeout / closed-and-empty. Every read below goes through this, because
 * a reasoning model can think for many seconds between the response headers
 * and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
  uint32_t t0 = millis();
  while (!c.available()) {
    if (!c.connected()) return false;
    if (millis() - t0 > ms) return false;
    delay(10);
  }
  return true;
}

static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
  body_out = "";

  if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);

  bool chunked = false;
  while (waitData(client, idle_ms)) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;         // blank line = end of headers
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
  }

  if (chunked) {
    int blanks = 0;
    while (true) {
      /* The chunk-size line may not arrive for a long time while the model
       * reasons. Waiting here - instead of letting read() time out - is the
       * whole fix: a timed-out read looks exactly like "0" (final chunk),
       * which silently truncated the body to nothing. */
      if (!waitData(client, idle_ms)) {
        Serial.println("HTTP: timed out waiting for the next chunk");
        break;
      }
      String szline = client.readStringUntil('\n');
      szline.trim();
      if (szline.length() == 0) {                    // stray blank line
        if (++blanks > 4) break;
        continue;
      }
      blanks = 0;

      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;                            // genuine final chunk

      long got = 0;
      while (got < sz) {
        if (!waitData(client, idle_ms)) break;
        while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
      }
      if (waitData(client, 3000)) client.readStringUntil('\n');   // CRLF after chunk
      if (got < sz) { Serial.println("HTTP: short chunk"); break; }
    }
  } else {
    while (waitData(client, idle_ms))
      while (client.available()) body_out += (char)client.read();
  }
  return code;
}


/* ===========================================================================
 *  CLOUD CALL 1  —  Azure speech-to-text
 *  One POST, one header, plain WAV body. This is why Azure does the ears.
 * =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
  /* HTTPClient's one-shot POST fails on bodies this large (it attempts one
   * giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
   * function speaks HTTP directly and streams the WAV up in 4 KB chunks -
   * reliable, and if it ever stalls we know the exact byte it stopped at. */
  WiFiClientSecure client;
  client.setInsecure();                    // no cert bundle on-device; see notes
  client.setTimeout(15);                   // seconds, for reads

  writeWavHeader(wav_buf, audio_bytes);
  dumpWavToSD(audio_bytes);                // PC-playable copy of what we send
  size_t total = WAV_HEADER_LEN + audio_bytes;

  if (!client.connect(AZURE_STT_HOST, 443)) {
    snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
    Serial.println("STT: TLS connect failed");
    return false;
  }

  /* Two valid host forms use DIFFERENT URL paths - detect which one is in
   * secrets.h:  <resource>.cognitiveservices.azure.com -> /stt/speech/...
   *             <region>.stt.speech.microsoft.com      -> /speech/...     */
  bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
  String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
               "/speech/recognition/conversation/cognitiveservices/v1"
               "?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
               "Host: " AZURE_STT_HOST "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
               "Accept: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(total) + "\r\n\r\n";
  client.print(req);

  /* body, 4 KB at a time */
  size_t sent = 0;
  while (sent < total) {
    size_t n = min((size_t)4096, total - sent);
    size_t w = client.write(wav_buf + sent, n);
    if (w == 0) {
      delay(50);                           // brief stall - retry once
      w = client.write(wav_buf + sent, n);
      if (w == 0) {
        snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
                 (unsigned)(sent / 1024));
        Serial.printf("STT: upload stalled at %u/%u bytes\n",
                      (unsigned)sent, (unsigned)total);
        client.stop();
        return false;
      }
    }
    sent += w;
    yield();
  }
  Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);

  /* read the reply with proper de-chunking */
  String resp;
  int code = readHttpResponse(client, resp, 10000);
  client.stop();

  if (code != 200) {
    snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
    Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
    Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
    return false;
  }

  const char *status = doc["RecognitionStatus"];
  if (!status || strcmp(status, "Success") != 0) {
    /* The status names the exact failure:
     *   InitialSilenceTimeout = Azure heard silence (mic level too low)
     *   NoMatch               = heard sound but no recognisable words
     *   BabbleTimeout         = heard only noise                       */
    snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
    Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
    return false;
  }
  text_out = doc["DisplayText"].as<String>();
  if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
  return text_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 2  —  DeepSeek chat completion
 *  OpenAI-compatible format. Model name is deepseek-v4-flash - the old
 *  deepseek-chat name is dead, see secrets.h.
 * =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
  /* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
   * bodies (long answers + the model's hidden reasoning), which shows up as
   * "bad JSON reply". Reading until the server closes the connection is
   * reliable regardless of reply length. */
  JsonDocument req;
  req["model"] = DEEPSEEK_MODEL;
  req["max_tokens"] = LLM_MAX_TOKENS;
  JsonArray msgs = req["messages"].to<JsonArray>();
  JsonObject sys = msgs.add<JsonObject>();
  sys["role"] = "system";  sys["content"] = SYSTEM_PROMPT;
  JsonObject usr = msgs.add<JsonObject>();
  usr["role"] = "user";    usr["content"] = question;

  String body;
  serializeJson(req, body);

  WiFiClientSecure client;
  client.setInsecure();
  /* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
   * ~2 s, but anything that needs actual working-out (an Ohm's law problem,
   * say) can think for 10-30 s before sending a single byte. */
  client.setTimeout(60);

  if (!client.connect(DEEPSEEK_HOST, 443)) {
    snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
    Serial.println("LLM: TLS connect failed");
    return false;
  }

  client.print(String("POST /chat/completions HTTP/1.1\r\n"
               "Host: " DEEPSEEK_HOST "\r\n"
               "Authorization: Bearer " DEEPSEEK_KEY "\r\n"
               "Content-Type: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: ") + String(body.length()) + "\r\n\r\n");
  client.print(body);

  /* read the reply with proper de-chunking; generous window - long
   * questions make the model think for a while before it responds */
  String resp;
  int code = readHttpResponse(client, resp, 60000);
  client.stop();

  if (code != 200) {
    snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
    Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
    Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
    return false;
  }

  const char *content = doc["choices"][0]["message"]["content"];
  const char *finish  = doc["choices"][0]["finish_reason"];
  if (!content || !content[0]) {
    /* v4-flash is a reasoning model: if finish_reason is "length", the whole
     * token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
    if (finish && strcmp(finish, "length") == 0)
      snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
    else
      snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
    Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
    return false;
  }
  answer_out = String(content);
  answer_out.trim();
  if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
  return answer_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 3  —  Azure text-to-speech, streamed straight to the speaker
 *  We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
 *  the I2S peripheral eats. Skip the 44-byte header, forward the rest.
 *  No MP3 decoder, no audio library, no buffering the whole reply.
 * =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
  size_t got = 0;
  uint32_t t0 = millis();
  while (got < n && millis() - t0 < 10000) {
    int r = c.read(dst + got, n - got);
    if (r > 0) { got += r; t0 = millis(); }
    else if (!c.connected() && !c.available()) break;
    else delay(2);
  }
  return got;
}

bool azureTTSSpeak(const String &text) {
  // Escape the XML special characters for the SSML body
  String safe = text;
  safe.replace("&", "&amp;");
  safe.replace("<", "&lt;");
  safe.replace(">", "&gt;");

  String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
                "<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";

  /* Manual HTTP like the other two cloud calls - and for a hard reason:
   * Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
   * raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
   * into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
   * Here we parse the framing properly and keep only clean audio bytes. */
  WiFiClientSecure client;
  client.setInsecure();
  client.setTimeout(20);

  const char *host = AZURE_REGION ".tts.speech.microsoft.com";
  if (!client.connect(host, 443)) {
    Serial.println("TTS: TLS connect failed");
    return false;
  }

  client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
               "Host: ") + host + "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: application/ssml+xml\r\n"
               "X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
               "User-Agent: MaTouchRobojax\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(ssml.length()) + "\r\n\r\n");
  client.print(ssml);

  /* status + headers; note whether the body is chunked */
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
  bool chunked = false;
  long content_len = -1;
  while (client.connected() || client.available()) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
    if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
  }
  if (code != 200) {
    Serial.printf("TTS HTTP %d\n", code);
    client.stop();
    return false;
  }

  const size_t AUDIO_CAP = 1200 * 1024;    // ~37 s of speech
  uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
  if (!audio) { client.stop(); return false; }
  size_t alen = 0;

  if (chunked) {
    /* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
    while (true) {
      String szline = client.readStringUntil('\n');
      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;
      if (alen + sz > AUDIO_CAP) break;
      size_t got = readExact(client, audio + alen, sz);
      alen += got;
      client.readStringUntil('\n');        // trailing CRLF after each chunk
      if (got < (size_t)sz) break;
    }
  } else if (content_len > 0) {
    alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
  } else {
    /* no framing info: read until the server closes */
    uint32_t idle = millis();
    while ((client.connected() || client.available()) && millis() - idle < 5000) {
      int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
      if (r > 0) { alen += r; idle = millis(); }
      else delay(5);
    }
  }
  client.stop();
  Serial.printf("TTS: %u KB clean audio (%s), playing\n",
                (unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");

  bool ok = (alen > WAV_HEADER_LEN);
  if (ok) {
    /* NOW the audio actually starts - this is the honest moment to go green */
    LED_SPEAK();
    drawBar("SPEAKING...", gfx->color565(0, 130, 40));

    /* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
    static const uint8_t lead_in[640] = {0};             // 20 ms of silence
    size_t w = 0;
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
  }
  free(audio);

  // let the DMA buffers drain so the last word is not cut off
  delay(150);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
  return ok;
}


/* ===========================================================================
 *  SETUP
 * =========================================================================== */
void setup() {
  Serial.begin(115200);
  delay(400);
  Serial.println("\n=== 04 Voice Assistant  |  Robojax.com ===");
  Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");

  pinMode(TFT_BLK, OUTPUT);
  digitalWrite(TFT_BLK, LOW);
  pinMode(SD_CS, OUTPUT);
  digitalWrite(SD_CS, HIGH);

  // one shared SPI bus for TFT + SD (started before either device)
  SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);

  gfx->begin();
  gfx->fillScreen(BLACK);
  digitalWrite(TFT_BLK, HIGH);

  // SD is optional here - it only stores the /stt_debug.wav diagnostic copy
  ok_sd = SD.begin(SD_CS, SPI, 20000000);
  digitalWrite(SD_CS, HIGH);
  Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
                       : "SD not found - debug WAV dump disabled (not fatal)");

  bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
  delay(50);

  rgb.begin();
  rgb.setBrightness(LED_BRIGHTNESS);
  LED_IDLE();

  /* One recording buffer for the whole session, in PSRAM. This is the 8 MB
   * that makes the board worth buying. */
  wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
  if (!wav_buf) {
    gfx->setTextColor(RED);
    gfx->setTextSize(2);
    gfx->setCursor(10, 100);
    gfx->print("PSRAM alloc failed!");
    gfx->setTextSize(1);
    gfx->setCursor(10, 130);
    gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
    while (1) delay(1000);
  }

  micInit();
  spkInit();

  gfx->setTextSize(1);
  gfx->setTextColor(YELLOW);
  gfx->setCursor(4, 4);
  gfx->printf("Connecting to %s ...", WIFI_SSID);
  Serial.printf("Connecting to %s ", WIFI_SSID);

  WiFi.mode(WIFI_STA);
  WiFi.begin(WIFI_SSID, WIFI_PASS);
  uint32_t t0 = millis();
  while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
    delay(300);
    Serial.print(".");
  }
  Serial.println();

  clearChat();
  if (WiFi.status() == WL_CONNECTED) {
    Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
    chatBubble("Hold SPEAK and ask me anything.", false);
  } else {
    chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
    LED_ERROR();
  }
  drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}


/* ===========================================================================
 *  LOOP  —  one full conversation turn per button press
 * =========================================================================== */
void loop() {
  /* CLEAR button: edge-detected so one tap wipes once. Reading the panel
   * twice per loop (here and in speakButtonHeld) is fine - the GT911 just
   * reports its current state. */
  static bool tap_latch = false;
  static uint8_t tap_release = 0;
  if (state == ST_IDLE) {
    uint16_t cx, cy;
    if (getTouch(&cx, &cy)) {
      tap_release = 0;
      if (!tap_latch) {
        tap_latch = true;
        if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
          clearChat();
          chatBubble("Hold SPEAK and ask me anything.", false);
        }
      }
    } else if (tap_latch && ++tap_release >= 4) {
      tap_latch = false;
      tap_release = 0;
    }

    /* live WiFi signal indicator, refreshed every 2 s while idle */
    static uint32_t last_wifi = 0;
    if (millis() - last_wifi > 2000) {
      last_wifi = millis();
      drawWifi();
    }
  }

  if (state == ST_IDLE && speakButtonHeld()) {

    /* ---- record ---- */
    state = ST_RECORDING;
    LED_LISTEN();
    drawBar("LISTENING...", gfx->color565(0, 60, 200));
    uint32_t t_rec = millis();
    size_t audio_bytes = recordWhileHeld();
    t_rec = millis() - t_rec;
    Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);

    if (audio_bytes < SAMPLE_RATE / 2) {   // under a quarter second - a tap
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }

    /* ---- speech to text ---- */
    state = ST_STT;
    LED_THINK();
    drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
    uint32_t t_stt = millis();
    String question;
    if (!azureSTT(audio_bytes, question)) {
      /* Show the REAL cause on screen - no serial monitor needed. */
      char diag[96];
      snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
               g_stt_err, g_mic_peak_pct,
               ok_sd ? " | saved /stt_debug.wav" : "");
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_stt = millis() - t_stt;
    chatBubble(question.c_str(), true);
    Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());

    /* ---- think ---- */
    state = ST_LLM;
    drawBar("THINKING...", gfx->color565(150, 90, 0));
    uint32_t t_llm = millis();
    String answer;
    if (!deepseekChat(question, answer)) {
      char diag[96];
      snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_llm = millis() - t_llm;
    chatBubble(answer.c_str(), false);
    Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());

    /* ---- speak ----
     * Still amber here: the voice has to be synthesised and downloaded first
     * (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
     * at the exact moment audio starts coming out of the speaker. */
    state = ST_TTS;
    LED_THINK();
    drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
    uint32_t t_tts = millis();
    bool spoke = azureTTSSpeak(answer);
    t_tts = millis() - t_tts;

    /* Timing summary on serial - this feeds the "honest numbers" segment. */
    Serial.printf("TIMINGS  rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
                  audio_bytes / 32000.0, (unsigned long)t_stt,
                  (unsigned long)t_llm, (unsigned long)t_tts,
                  spoke ? "" : "  (TTS FAILED)");

    drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
    LED_IDLE();
    state = ST_IDLE;
  }

  delay(20);
}

الأشياء التي قد تحتاجها

ملفات📁

الملف المطلوب (.h)

  • secrets.h
    ملف لـ Makerfabs MaTouch AI ESP32S3 2.8" وحدة كاميرا TFT
    secrets.h 0.01 MB

ملفات أخرى

  • pins.h
    ملف دبابيس لـ MaTouch AI ESP32S3 2.8" شاشة LCD تعمل باللمس مع كاميرا.
    pins.h 0.01 MB

مخطط

  • MaTouch_AI 2.8" MaTouch AI ESP32S3 2.8" TFT ST7789V مخطط
    أحدث لوحة MaTouch AI تدمج إدخال صوت I2S/مكبر صوت I2S/ كاميرا 3 ملايين بكسل OV3660/ شاشة بدقة 320*240، مع معالج ESP32S3 القوي وقدرة WiFi، لجعل هذه اللوحة أداة/منصة جيدة لتطوير الذكاء الاصطناعي مع ESP32.
    MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF 0.15 MB