Search Code

میکر فیبس ما ٹچ ESP32-S3 2.8 انچ کیمرہ ESP32-S3 پر AI وائس اسسٹنٹ بنائیں (Azure + DeepSeek)

میکر فیبس ما ٹچ ESP32-S3 2.8 انچ کیمرہ ESP32-S3 پر AI وائس اسسٹنٹ بنائیں (Azure + DeepSeek)

ایک بٹن دبائیں، سوال پوچھیں، اور بورڈ زور سے جواب دے گا

ایک مکمل وائس اسسٹنٹ ایک چھوٹے بورڈ پر۔ SPEAK بٹن دبائے رکھیں اور سوال پوچھیں۔ بورڈ آپ کو دونوں مائیکروفونز سے ریکارڈ کرتا ہے، آڈیو کو Microsoft Azure کو بھیجتا ہے تاکہ اسے متن میں تبدیل کیا جائے، اس متن کو DeepSeek کو بھیجتا ہے تاکہ وہ سوچے، جواب کو واپس Azure کو بھیجتا ہے تاکہ اسے تقریر میں تبدیل کیا جائے، اور اسے اپنے اسپیکر کے ذریعے چلاتا ہے۔ پوری گفتگو اسکرین پر چیٹ بلبلز کے طور پر ظاہر ہوتی ہے۔

Talking AI Voice Assistant running on the MaTouch AI ESP32-S3 board MaTouch AI ESP32-S3 بورڈ پر چلنے والا بولنے والا AI وائس اسسٹنٹ

بٹن دبانے کے لمحے سے کیا ہوتا ہے

یہ ہے ایک سوال کا پورا سفر، قدم بہ قدم۔ اسے ایک بار پڑھنا ضروری ہے، کیونکہ اسکرین اور LED پر جو کچھ بھی آپ دیکھتے ہیں وہ ان میں سے کسی ایک مرحلے سے مطابقت رکھتا ہے۔

  1. آپ SPEAK بٹن دبائے رکھتے ہیں۔ یہ دبائے رکھنے سے بولنا ہے، ٹیپ کرنے سے نہیں: ریکارڈنگ بالکل اسی وقت تک چلتی ہے جب تک آپ کی انگلی نیچے رہتی ہے، زیادہ سے زیادہ چھ سیکنڈ تک۔ اسٹیٹس LED نیلی ہو جاتی ہے اور بٹن LISTENING دکھاتا ہے۔

  2. دونوں مائیکروفون آپ کو ریکارڈ کرتے ہیں۔ بورڈ سٹیریو جوڑی سے ہر سیکنڈ میں 16,000 بار نمونے لیتا ہے، دونوں چینلز کو ایک میں ملا دیتا ہے، تھوڑا سا گین لگاتا ہے، اور نتیجہ PSRAM میں محفوظ کرتا ہے۔ جب آپ بولتے ہیں تو بٹن پر ایک پیش رفت بار پھیلتی ہے۔ دو سیکنڈ کی تقریر تقریباً 64 KB ہوتی ہے۔

  3. آپ بٹن چھوڑ دیتے ہیں۔ ریکارڈنگ رک جاتی ہے۔ بورڈ آڈیو کے شروع میں 44 بائٹ کا WAV ہیڈر لکھتا ہے - یہ چھوٹا سا لیبل ہی وہ چیز ہے جو خام نمونوں کو ایک فائل میں بدل دیتا ہے جسے Azure قبول کرے گا۔

  4. آڈیو Azure Speech-to-Text کو جاتا ہے۔ اسے ایک محفوظ کنکشن پر 4 KB کے ٹکڑوں میں اپ لوڈ کیا جاتا ہے، اور ایک متن کی ایک لائن کے طور پر واپس آتا ہے۔ آپ کی 64 KB آواز تقریباً 25 بائٹ تحریر بن گئی ہے۔ LED عنبری ہو جاتی ہے۔

  5. آپ کا سوال اسکرین پر ظاہر ہوتا ہے ایک نیلے چیٹ بلبل کے طور پر، تاکہ آپ بالکل دیکھ سکیں کہ اس نے کیا سنا - جو مفید ہے، کیونکہ غلط سنے گئے الفاظ زیادہ تر عجیب جوابوں کی وضاحت کرتے ہیں۔

  6. متن DeepSeek کو جاتا ہے۔ بورڈ آپ کا سوال اور ایک مستقل ہدایت بھیجتا ہے کہ جواب دو چھوٹے جملوں میں رکھیں۔ ماڈل سوچتا ہے - واقعی سوچتا ہے، یہ ایک استدلال کرنے والا ماڈل ہے - اور جواب واپس کرتا ہے۔

  7. جواب اسکرین پر ظاہر ہوتا ہے ایک سرمئی بلبل کے طور پر۔ آپ اسے سننے سے پہلے پڑھ سکتے ہیں۔

  8. جواب بولنے کے لیے واپس Azure کو جاتا ہے۔ بورڈ خام 16 kHz PCM آڈیو مانگتا ہے، جو بالکل وہی فارمیٹ ہے جو اس کے ایمپلیفائر کو چاہیے، اس لیے اس پروجیکٹ میں کہیں بھی MP3 ڈیکوڈر نہیں ہے۔ بٹن اب GETTING VOICE پڑھتا ہے اور LED عنبری رہتی ہے، کیونکہ ابھی کچھ سنائی نہیں دیتا۔

  9. پوری کلپ PSRAM میں ڈاؤن لوڈ ہوتی ہے اس سے پہلے کہ ایک بھی نمونہ چلایا جائے۔ یہ اہم ہے - نیچے دیا گیا نوٹ دیکھیں۔

  10. پلے بیک۔ جس لمحے آڈیو اسپیکر کو دی جاتی ہے، LED سبز ہو جاتی ہے اور بٹن SPEAKING پڑھتا ہے۔ آپ جواب سنتے ہیں۔

آڈیو کو پہلے ڈاؤن لوڈ کیوں کیا جاتا ہے بجائے اس کے کہ آتے ہی چلایا جائے۔ نیٹ ورک سے براہ راست اسپیکر میں سٹریم کرنا کھٹکھٹانے جیسا لگتا ہے۔ اسپیکر کا بفر صرف ایک سیکنڈ کے دسویں حصے کے برابر رکھتا ہے، اور WiFi منتقلی میں ہر وقفہ جو اس سے زیادہ طویل ہو اسے خالی کر دیتا ہے، جس سے سنائی دینے والی کھٹک پیدا ہوتی ہے۔ پورے جواب کو پہلے PSRAM میں ڈاؤن لوڈ کرنے میں تقریباً ایک سیکنڈ اضافی انتظار لگتا ہے اور ہر خلا کو ختم کر دیتا ہے۔ یہی وجہ ہے کہ ڈسپلے SPEAKING کہنے سے پہلے GETTING VOICE کہتا ہے - یہ دونوں واقعی الگ الگ مراحل ہیں۔

ہر مرحلے میں کتنا وقت لگتا ہے

حقیقی ہارڈویئر پر ناپا گیا، ایک سادہ سوال کے لیے:

مرحلہ

عام وقت

ریکارڈنگ

جتنی دیر آپ بٹن دبائے رکھیں

تقریر سے متن (Azure)

تقریباً 1.8 سیکنڈ

سوچنا (DeepSeek)

سادہ سوال کے لیے تقریباً 1.8 سیکنڈ، حقیقی حساب کتاب کی ضرورت والے سوال کے لیے بہت زیادہ

آواز حاصل کرنا (Azure)

تقریباً 7 سیکنڈ - سب سے بڑا حصہ

کل، چھوڑنے سے پہلی آواز تک

تقریباً 11 سیکنڈ

ہر تبادلہ اپنے اوقات سیریل مانیٹر پر پرنٹ کرتا ہے، تاکہ آپ ان پر بھروسہ کرنے کے بجائے اپنے خود ناپ سکیں۔ اگر آپ اسے تیز چاہتے ہیں، تو سب سے مؤثر تبدیلی SYSTEM_PROMPT میں چھوٹے جوابات مانگنا ہے - بولنے کے لیے کم متن کا مطلب ہے ترکیب اور ڈاؤن لوڈ کرنے کے لیے کم آڈیو۔

بورڈ خود کبھی کچھ نہیں سمجھتا۔ یہ ایک قاصد ہے جس کے کان اچھے ہیں اور آواز اچھی ہے - ذہانت سیکنڈ کے حساب سے کرائے پر لی گئی ہے۔

یہ تین خدمات کیوں

  • Azure اندر اور باہر تقریر سنبھالتا ہے۔ اس کی متن سے تقریر خام 16 kHz PCM واپس کر سکتی ہے، جو بالکل وہی ہے جو اسپیکر چپ چاہتا ہے، اس لیے اس پروجیکٹ میں کہیں بھی MP3 ڈیکوڈر نہیں ہے۔ اس کی تقریر سے متن ایک سادہ POST میں ایک سادہ WAV لیتا ہے۔

  • DeepSeek گفتگو کا دماغ ہے۔ یہ تیز ہے اور ہر جواب پر ایک سینٹ کا ایک حصہ خرچ کرتا ہے۔

  • OpenAI یہاں استعمال نہیں ہوتا - پروجیکٹ 05 دیکھیں، جہاں یہ بصری کام کرتا ہے۔

DeepSeek ماڈل کے نام بدل گئے۔ پرانے deepseek-chat اور deepseek-reasoner نام جولائی 2026 میں ریٹائر ہو گئے۔ زیادہ تر آن لائن ٹیوٹوریل اب بھی انہیں استعمال کرتے ہیں اور ایک خرابی واپس کریں گے۔ موجودہ نام deepseek-v4-flash اور deepseek-v4-pro ہیں۔ یہ پروجیکٹ v4-flash استعمال کرتا ہے۔

استدلال کرنے والے ماڈل کا جال

DeepSeek v4-flash جواب دینے سے پہلے سوچتا ہے، اور یہ سوچ آپ کی ٹوکن حد میں شمار ہوتی ہے۔ LLM_MAX_TOKENS بہت کم سیٹ کریں تو پورا بجٹ سوچنے میں خرچ ہو جاتا ہے، جواب خالی واپس آتا ہے، اور بورڈ کچھ نہیں کہتا۔ یہی وجہ ہے کہ یہاں 400 سیٹ کیا گیا ہے۔ مشکل سوال بھی زیادہ وقت لیتے ہیں - ایک سادہ حقیقت کا جواب تقریباً دو سیکنڈ میں آتا ہے، جس سوال میں حقیقی حساب کتاب درکار ہو اس میں بہت زیادہ وقت لگ سکتا ہے۔

اسٹیٹس لائٹ پڑھنا

رنگ

مطلب

نیلا

آپ کی بات سن رہا ہے

عنبری

کلاؤڈ سوچ رہا ہے، یا آواز لائی جا رہی ہے

سبز

بول رہا ہے - یہ بالکل اسی لمحے سبز ہوتا ہے جب آواز شروع ہوتی ہے

سرخ

کچھ ناکام ہو گیا - سیریل مانیٹر چیک کریں

اسکرین پر کنٹرولز

چیٹ فون گفتگو کی طرح اسکرول ہوتی ہے، پرانے پیغامات اوپر جاتے ہوئے غائب ہو جاتے ہیں۔ CLEAR اسے صاف کر دیتا ہے۔ کونے میں اصل dBm ریڈنگ کے ساتھ وائی فائی سگنل میٹر موجود ہے، جو اس وقت مفید ہے جب آپ سوچ رہے ہوں کہ سست جواب نیٹ ورک کی وجہ سے ہے یا سروس کی۔

MaTouch AI ESP32-S3 2.8" بورڈ کے بارے میں

اس صفحے کا ہر پروجیکٹ Makerfabs کے MaTouch AI ESP32-S3 2.8" TFT ST7789V پر چلتا ہے۔ یہ ایک آل ان ون بورڈ ہے: ایک رنگین ٹچ اسکرین، ایک 3 میگاپکسل کیمرہ، دو مائیکروفون اور ایک حقیقی اسپیکر ایمپلیفائر، سب کچھ ESP32-S3 کے ذریعے چلایا جاتا ہے جس میں 8 MB PSRAM ہے۔ یہی مجموعہ ان AI پروجیکٹس کو ایک ہی بورڈ پر ممکن بناتا ہے جس میں کچھ اور منسلک نہ ہو۔

8 MB PSRAM یہاں کسی بھی دوسرے نمبر سے زیادہ اہم ہے۔ یہی بورڈ کو ایک ہی وقت میں میموری میں کیمرہ فریم، چند سیکنڈ کی ریکارڈ شدہ آڈیو، یا بیس64-انکوڈڈ تصویر رکھنے کی اجازت دیتا ہے - ان میں سے کوئی بھی ESP32 کی عام RAM میں فٹ نہیں ہوتا۔

مینوفیکچرر دستاویزات: Makerfabs ویکی صفحہ۔

اہم خصوصیات

  • پروسیسر: ESP32-S3، ڈوئل کور 240 MHz، وائی فائی 2.4 GHz + بلوٹوتھ 5.0

  • میموری: 16 MB فلیش، 8 MB PSRAM (یہاں تقریباً ہر پروجیکٹ کے لیے ضروری)

  • ڈسپلے: 2.8" IPS، 320×240، ST7789V ڈرائیور، SPI

  • ٹچ: GT911 کیپسیٹو، ایک ساتھ 5 انگلیوں کو ٹریک کرتا ہے

  • کیمرہ: OV3660، 3 میگاپکسل، 2048×1536 تک

  • مائیکروفون: دو INMP441 I2S ڈیجیٹل مائکس (ایک حقیقی سٹیریو جوڑا)

  • اسپیکر: MAX98357A کلاس-D ایمپلیفائر، 4 Ω پر 3.2 W

  • اسٹوریج: مائیکرو ایس ڈی کارڈ سلاٹ (SPI موڈ)

  • پاور: USB-C، JST بیٹری کنیکٹر، TP4056 چارجر، پاور سوئچ

  • بورڈ پر بھی: WS2812B RGB LED، PCF8563T بیٹری بیکڈ ریئل ٹائم کلاک، اور ایک MAX17048 بیٹری فیول گیج جو سرکاری خصوصیات میں درج نہیں ہے

دونوں USB-C پورٹس ایک جیسے نہیں ہیں۔ بورڈ کا اسپیکر اپنے سگنل پن (IO19 اور IO20) کو نیٹیو USB پورٹ کے ساتھ شیئر کرتا ہے، کیونکہ وہ پن ESP32-S3 کی ہارڈوائرڈ USB ڈیٹا لائنیں ہیں۔ ہمیشہ CH340K USB-C پورٹ (RESET بٹن کے پاس والا) کے ذریعے اپ لوڈ اور پاور کریں، اور USB CDC On Boot کو Disabled سیٹ کریں۔ غلط پورٹ استعمال کریں تو آڈیو خراب ہو گی یا اپ لوڈ ناکام ہو جائیں گے۔

Arduino IDE سیٹنگز

یہ سیٹنگز اہم ہیں۔ اس بورڈ کے بارے میں لوگ جو زیادہ تر مسائل رپورٹ کرتے ہیں ان میں سے ایک یہی غلط ہوتی ہے، اور یہ کور ورژن تبدیل کرنے پر ری سیٹ ہو جاتی ہیں، لہٰذا کسی بھی تبدیلی کے بعد انہیں دوبارہ چیک کریں۔

سیٹنگ

قدر

بورڈ

ESP32S3 Dev Module

ESP32 کور ورژن

2.0.17

PSRAM

OPI PSRAM

فلیش سائز

16MB (128Mb)

پارٹیشن اسکیم

16M Flash (3MB APP/9.9MB FATFS)

USB CDC On Boot

Disabled

اپ لوڈ اسپیڈ

921600

اپ لوڈ سے پہلے تمام فلیش مٹائیں

Disabled

پورٹ

CH340K USB-C پورٹ

ESP32 کور 2.0.17 استعمال کریں، 3.x نہیں۔ Espressif نے کور 3 میں ڈیوائس پر چہرہ شناخت کے ماڈل ہٹا دیے ہیں، لہٰذا چہرے کے پروجیکٹ وہاں کمپائل نہیں ہوں گے۔ 2.0.17 کو پن کرنے سے اس صفحے کا ہر پروجیکٹ ایک کنفیگریشن کے ساتھ کام کرتا رہتا ہے۔ Boards Manager میں، ورژن ڈراپ ڈاؤن آپ کو جب چاہیں آگے پیچھے سوئچ کرنے دیتا ہے۔

GFX Library for Arduino ورژن 1.5.6 استعمال کریں، 1.6.x نہیں۔ 1.6 ریلیزز ESP32 کور 3 کے لیے بنائی گئی ہیں اور کور 2.0.17 پر شروع ہوتے وقت پھنس سکتی ہیں۔ اگر اپ لوڈ کے بعد آپ کی اسکرین سیاہ رہتی ہے، تو یہ سب سے پہلے چیک کرنے والی چیز ہے۔

مطلوبہ لائبریریاں

انہیں Arduino IDE میں Tools → Manage Libraries کے ذریعے انسٹال کریں۔ ورژن نمبر اہم ہیں - براہ کرم درج کردہ استعمال کریں۔

لائبریری

ورژن

مصنف

GFX Library for Arduino

1.5.6

moononournation

bb_captouch

1.3.1

Larry Bank

ArduinoJson

7.x

Benoit Blanchon

Adafruit NeoPixel

کوئی بھی حالیہ

Adafruit

secrets.h ترتیب دینا

آپ کی WiFi تفصیلات اور کوئی بھی API کلید secrets.h میں ڈالی جاتی ہیں، جو ڈاؤن لوڈ میں پلیس ہولڈر اقدار کے ساتھ شامل ہوتی ہے۔ Arduino IDE میں اس ٹیب کو کھولیں اور انہیں اپنی اپنی اقدار سے تبدیل کریں۔

WiFi کا 2.4 GHz ہونا ضروری ہے۔ ESP32-S3 5 GHz نیٹ ورک کو بالکل نہیں دیکھ سکتا۔ اگر آپ کا راؤٹر دونوں بینڈز کو ایک نام کے تحت یکجا کرتا ہے (Asus اسے Smart Connect کہتا ہے)، تو یا تو اسے بند کریں یا 2.4 GHz بینڈ کو اپنا الگ نام دیں اور اسے secrets.h میں استعمال کریں۔

اپنی API کلیدیں حاصل کرنا

یہ پروجیکٹ کلاؤڈ AI سروس سے بات کرتا ہے، اس لیے آپ کو اپنی کلید کی ضرورت ہے۔ اگر آپ نے پہلے کبھی ایسا نہیں کیا تو پریشان نہ ہوں - یہ ایک پاس ورڈ کی طرح ہے جو آپ کے اکاؤنٹ کو سروس سے پہچانتا ہے۔ اس میں ایک بار چند منٹ لگتے ہیں۔

کلید کسی ویب سائٹ کی سبسکرپشن نہیں ہے۔ مثال کے طور پر، ChatGPT Plus کے لیے ادائیگی کرنے سے آپ کو API کلید نہیں ملتی - یہ دو الگ مصنوعات ہیں جن کا بلنگ الگ ہے۔ آپ کو ڈیولپر پلیٹ فارم پر اکاؤنٹ کی ضرورت ہے، جیسا کہ نیچے بیان کیا گیا ہے۔

Microsoft Azure Speech - سننے اور بولنے کے لیے

Azure آپ کی تقریر کو متن میں بدلتا ہے اور جواب کو دوبارہ آواز میں بدل دیتا ہے۔ مفت درجہ اس صفحے کی ہر چیز کے لیے کافی فراخدلانہ ہے۔

  1. portal.azure.com پر جائیں اور Microsoft اکاؤنٹ سے سائن ان کریں (مفت اکاؤنٹ بھی ٹھیک ہے)۔

  2. اگر آپ نے پہلے کبھی Azure استعمال نہیں کیا تو آپ کو Welcome to Azure اسکرین نظر آئے گی جس میں تین اختیارات ہوں گے۔ Start with an Azure free trial منتخب کریں - Azure کو کچھ بنانے کی اجازت دینے سے پہلے آپ کو سبسکرپشن کی ضرورت ہے۔ (طلباء کو اس کے بجائے Azure for Students منتخب کرنا چاہیے: ایک ہی نتیجہ، کارڈ کی ضرورت نہیں۔) Manage Microsoft Entra ID کو نظر انداز کریں، جو بالکل مختلف چیز ہے۔

  3. Create a resource پر کلک کریں، Speech تلاش کریں، اور Microsoft کی شائع کردہ Speech service منتخب کریں۔

  4. فارم بھریں: کوئی بھی ریسورس گروپ، کوئی بھی نام، اور اپنے قریب ایک Region منتخب کریں - اس علاقے کو بالکل ویسا ہی لکھیں جیسا ظاہر ہوتا ہے، مثال کے طور پر eastus۔

  5. Pricing tier کے لیے F0 (Free) منتخب کریں۔ یہ ہر ماہ تقریباً پانچ گھنٹے تقریر سے متن اور پانچ لاکھ حروف متن سے تقریر کی اجازت دیتا ہے۔

  6. Review + create پر کلک کریں، پھر Create پر۔ تقریباً ایک منٹ انتظار کریں، پھر Go to resource پر کلک کریں۔

  7. بائیں مینو میں Keys and Endpoint کھولیں۔ KEY 1 اور Location/Region کاپی کریں۔

انہیں secrets.h میں AZURE_SPEECH_KEY اور AZURE_REGION کے طور پر ڈالیں۔ AZURE_STT_HOST کے لیے، <region>.stt.speech.microsoft.com استعمال کریں - لہذا علاقے eastus کے ساتھ یہ eastus.stt.speech.microsoft.com ہوگا۔

کریڈٹ کارڈ کے بارے میں۔ Azure مفت آزمائش آپ کی شناخت کی تصدیق کے لیے کارڈ مانگتی ہے۔ یہ آپ سے کوئی رقم نہیں لیتا۔ آپ کو 30 دنوں کے لیے $200 کا کریڈٹ ملتا ہے، اور اس کے بعد اکاؤنٹ Pay-As-You-Go پر منتقل ہو جاتا ہے - لیکن F0 Speech درجہ مفت رہتا ہے، مہینے بہ مہینے، اور اس صفحے کی ہر چیز اس میں آرام سے فٹ بیٹھتی ہے۔ اگر آپ بالکل کارڈ نہیں دینا چاہتے اور آپ طالب علم ہیں، تو Azure for Students آپشن آپ کو بغیر کارڈ کے کریڈٹ دیتا ہے۔

یہ "Speech service" ریسورس ہونا ضروری ہے۔ Translator، Language، یا عام Cognitive Services ریسورس کی کلید ایک جیسی نظر آتی ہے اور بالکل درست ہوتی ہے - لیکن ہر تقریر کی درخواست پر 401 غلطی آتی ہے۔ یہ ہمیں جانچ کے دوران ملا اور ایک گھنٹہ ضائع کیا۔ اگر تقریر 401 کے ساتھ ناکام ہو جبکہ کلید درست نظر آتی ہے، تو چیک کریں کہ آپ نے کس قسم کا ریسورس بنایا ہے۔

DeepSeek - سوچنے والا حصہ

DeepSeek وہ زبان کا ماڈل ہے جو دراصل آپ کے سوال کا جواب دیتا ہے۔ یہ سستا ہے - چند ڈالر کا کریڈٹ ہزاروں جوابات کا احاطہ کرتا ہے۔

  1. platform.deepseek.com پر جائیں اور اکاؤنٹ بنائیں۔

  2. مینو میں API keys کھولیں اور Create new API key پر کلک کریں۔

  3. اسے فوراً کاپی کریں۔ یہ صرف ایک بار دکھائی جاتی ہے اور پھر کبھی نہیں - اگر آپ اسے کھو دیں تو اس کلید کو حذف کریں اور دوسری بنائیں۔

  4. Top up کے تحت تھوڑی سی رقم شامل کریں۔ کوئی مفت درجہ نہیں ہے، لیکن سب سے چھوٹی ٹاپ اپ اس استعمال پر بہت طویل عرصہ چلتی ہے۔

کلید کو secrets.h میں DEEPSEEK_KEY کے طور پر ڈالیں۔ یہ sk- سے شروع ہوتی ہے۔

جولائی 2026 میں ماڈل کے نام تبدیل ہوئے۔ پرانے deepseek-chat اور deepseek-reasoner کو ریٹائر کر دیا گیا، اس لیے آن لائن ملنے والے زیادہ تر ٹیوٹوریلز 400 غلطی کے ساتھ ناکام ہوں گے۔ deepseek-v4-flash استعمال کریں، جو ان پروجیکٹس میں پہلے سے سیٹ ہے۔

چلانے کی لاگت کیا ہے

بہت کم، لیکن یہ مفت نہیں ہے، اور آپ کو تقریباً معلوم ہونا چاہیے کہ پروجیکٹ چھوڑنے سے پہلے آپ کتنا خرچ کر رہے ہیں۔

سروس

تقریبی لاگت

Azure Speech

مفت درجہ ہر ماہ تقریباً 5 گھنٹے سننے اور 0.5 M حروف بولنے کا احاطہ کرتا ہے

DeepSeek

فی جواب ایک سینٹ کا ایک حصہ - چند ڈالر میں ہزاروں جوابات

OpenAI vision

ماڈل کے لحاظ سے تقریباً ایک یا دو سینٹ فی تصویر

قیمتیں بدلتی رہتی ہیں، اس لیے انہیں حوالے کے بجائے رہنما کے طور پر لیں۔ ان میں سے ہر سروس کا ایک استعمال کا صفحہ ہے جہاں آپ اپنا خرچ دیکھ سکتے ہیں، اور سب آپ کو خرچ کی حد مقرر کرنے دیتی ہیں - جو پہلے دن کرنا قابل قدر ہے۔

اپنی چابیاں نجی رکھیں۔ جس کے پاس بھی یہ ہوں گی وہ آپ کے پیسے خرچ کر سکتا ہے۔ انہیں ویڈیو، اسکرین شاٹ، فورم پوسٹ یا عوامی کوڈ ریپوزٹری میں نہ ڈالیں۔ اگر کوئی چابی کبھی ظاہر ہو جائے، تو اسے فراہم کنندہ کی ویب سائٹ پر حذف کریں اور نئی بنائیں - اس میں سیکنڈ لگتے ہیں، اور یہی واحد حقیقی حل ہے۔

مسائل کا حل

علامت

وجہ اور حل

اسکرین سیاہ رہتی ہے

غلط GFX لائبریری ورژن (1.5.6 استعمال کریں) یا غلط بورڈ سیٹنگز۔

PSRAM alloc failed یا کیمرہ خرابی 0xffffffff

Tools → PSRAM کو OPI PSRAM پر سیٹ نہیں کیا گیا۔

کچھ اپ لوڈ نہیں ہوتا / کوئی COM پورٹ نہیں

غلط USB-C پورٹ، یا CH340 ڈرائیور انسٹال نہیں ہے۔

کیمرہ ناکام ہو جاتا ہے اور کبھی بحال نہیں ہوتا

کیمرے کی ری سیٹ لائن بورڈ کے RESET بٹن سے منسلک ہے، لہذا سافٹ ویئر اسے دوبارہ شروع نہیں کر سکتا۔ RESET دبائیں۔ اگر پھر بھی ناکام ہو، تو کیمرے کی ربن کیبل کو دوبارہ لگائیں۔

کوڈ ڈاؤن لوڈ کریں

اس پروجیکٹ کے لیے مکمل Arduino اسکیچ، pins.h اور باقی سب کچھ سمیت، مفت ڈاؤن لوڈ کے لیے دستیاب ہے۔

04_Voice_Assistant ڈاؤن لوڈ کریں

اسے ان زپ کریں، Arduino IDE میں .ino فائل کھولیں، اوپر دی گئی سیٹنگز چیک کریں، اور CH340K USB-C پورٹ کے ذریعے اپ لوڈ کریں۔

884-Arduin code for MaTouch AI ESP32S3 2.8in AI Camera: AI voice assistant
زبان: C++
/*
 * ===========================================================================
 *  04_Voice_Assistant  —  MaTouch AI ESP32-S3 2.8" TFT ST7789V
 * ===========================================================================
 *
 ----------
 *  ROBOJAX.COM  -  MaTouch AI ESP32-S3 2.8" project series
 *
 *    WATCH THE VIDEO
 *        https://youtu.be/6AL3g3tC_Hk
 *
 *    WRITTEN TUTORIALS - every project, with photos and full explanation
 *        Camera and touchscreen.... https://robojax.com/RTJ849
 *        Offline face recognition.. https://robojax.com/RTJ850
 *        AI voice assistant........ https://robojax.com/RTJ851
 *        AI vision................. https://robojax.com/RTJ852
 *
 *    GET THE BOARD - SAVE $5 with coupon code:  Robojax_Makerfab
 *        https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
 *        (enter the code at checkout)
 *
 *  All of this code is free. If it helped you, a subscribe on YouTube is
 *  the best way to support more of it.
 *  
 *  ---------------------------------------------------------------------------
 *
 *  A complete voice assistant on a $40 board:
 *
 *      hold SPEAK  ->  both INMP441 microphones record you
 *                  ->  Azure Speech turns the audio into text
 *                  ->  DeepSeek v4-flash thinks of an answer
 *                  ->  Azure Speech turns the answer into audio
 *                  ->  the MAX98357 speaker says it out loud
 *
 *  and the whole conversation is drawn as chat bubbles on the touchscreen.
 *
 *  WHY THIS COMBINATION OF SERVICES (each is used where it is best):
 *   - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
 *     transcription endpoint wants multipart/form-data - miserable on an MCU.
 *   - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
 *     which streams straight into the I2S speaker with NO MP3 decoder at all.
 *   - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
 *     model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
 *     Most tutorials online still use them and are broken. See secrets.h.
 *
 *  A detail the vendor examples get wrong: this board has TWO microphones on
 *  one I2S bus (left + right), but every Makerfabs demo records left-only and
 *  throws one away. This sketch records both and averages them.
 *
 *  ---------------------------------------------------------------------------
 *  *** WHICH USB PORT - THIS MATTERS ***
 *  The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
 *  through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
 *  If you use the wrong port the audio will be garbage or uploads will fail.
 *  ---------------------------------------------------------------------------
 *
 *  FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
 *
 *  BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
 *  or compile errors. These reset when you switch cores - recheck them!)
 *
 *      Board            : ESP32S3 Dev Module
 *      ESP32 core       : 2.0.17
 *      PSRAM            : OPI PSRAM        <-- required, audio buffer lives there
 *      Flash Size       : 16MB (128Mb)
 *      Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
 *      USB CDC On Boot  : Disabled         <-- required, see USB note above
 *      Upload Speed     : 921600
 *      Port             : the CH340K USB-C port (the one near RESET)
 *
 *  LIBRARIES
 *      GFX Library for Arduino   v1.5.6   (NOT 1.6.x - that pairs with core 3)
 *      bb_captouch               v1.3.1
 *      ArduinoJson               v7.x
 *      Adafruit NeoPixel         any recent
 *
 *  ---------------------------------------------------------------------------
 *  FUNCTIONS IN THIS SKETCH
 *      led(r,g,b) + LED_* macros   RGB status colours (blue/amber/green/red)
 *      getTouch(&x,&y)      read the touch panel, mapped to screen coordinates
 *      speakButtonHeld()    true while a finger is on the SPEAK button
 *      bubbleLines(t)       how many lines a message wraps to
 *      drawOneBubble(m,y)   draw a single chat bubble
 *      redrawChat()         rebuild the chat area from history, newest at bottom
 *      clearChat()          wipe the chat history (CLEAR button)
 *      chatBubble(t,user)   add a message to history and redraw
 *      drawWifi()           WiFi signal bars + dBm readout
 *      drawBar(label,col)   bottom bar: SPEAK button + CLEAR + WiFi meter
 *      micInit()            I2S input - BOTH INMP441 mics, stereo
 *      spkInit()            I2S output - MAX98357 speaker
 *      recordWhileHeld()    record while SPEAK held, downmix stereo->mono
 *      writeWavHeader(...)  prepend the 44-byte RIFF/WAVE header
 *      dumpWavToSD(...)     save the exact upload to SD (/stt_debug.wav)
 *      readHttpResponse()   read an HTTPS reply, de-chunking it properly
 *      azureSTT(...)        chunked upload of the WAV -> recognised text
 *      deepseekChat(...)    question -> deepseek-v4-flash -> answer text
 *      azureTTSSpeak(text)  answer -> Azure voice -> PSRAM -> speaker
 *      setup() / loop()     boot + WiFi / one conversation turn per press
 *
 *  Robojax.com
 * ===========================================================================
 */

#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"

/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE     16000
#define RECORD_MAX_S    6                       // hard cap on one question
#define WAV_HEADER_LEN  44
#define REC_BUF_BYTES   (SAMPLE_RATE * RECORD_MAX_S * 2)   // 16-bit mono

/* Software gain applied to the recording. The INMP441 capture is quiet at
 * 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
 * speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
 * lower it. */
#define MIC_GAIN        6

/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
 * 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
 *     back silent or garbled in stereo mode. */
#define USE_BOTH_MICS   1

/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
 * uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS  25

/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
 * shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
    TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);

BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);

/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr;      // WAV_HEADER_LEN + up to REC_BUF_BYTES

/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;

/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool  ok_sd = false;
float g_mic_peak_pct = 0;          // last recording's raw peak, % of full scale
char  g_stt_err[64] = "";          // last STT failure cause, verbatim
char  g_llm_err[64] = "";          // last DeepSeek failure cause, verbatim

/* --- layout --------------------------------------------------------------- */
#define CHAT_H     200              // chat area: y 0..199
#define BAR_Y      202              // button bar below it
#define BTN_SPEAK_X   4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W  58
#define WIFI_X      236             // signal indicator, right end of the bar
#define BTN_H        36

/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
 * This is what prevents new text printing over old - the whole area is
 * rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int     chat_count = 0;

/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
  rgb.setPixelColor(0, rgb.Color(r, g, b));
  rgb.show();
}
#define LED_IDLE()    led(0, 0, 0)
#define LED_LISTEN()  led(0, 60, 255)     // blue   - recording
#define LED_THINK()   led(255, 120, 0)    // amber  - waiting on the cloud
#define LED_SPEAK()   led(0, 255, 40)     // green  - talking
#define LED_ERROR()   led(255, 0, 0)      // red


/* ===========================================================================
 *  Touch
 * =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
  TOUCHINFO ti;
  if (!bbct.getSamples(&ti)) return false;
  if (ti.count < 1) return false;
  *x = ti.y[0];
  *y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
  return true;
}

bool speakButtonHeld() {
  uint16_t x, y;
  if (!getTouch(&x, &y)) return false;
  return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}


/* ===========================================================================
 *  Chat UI  —  word-wrapped bubbles, user right/blue, assistant left/grey
 * =========================================================================== */
#define CHAT_CHARS 42                          // chars per line at textsize 1

static int bubbleLines(const char *t) {
  int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
  return l < 1 ? 1 : l;
}

void drawOneBubble(const ChatMsg &m, int y) {
  int len = strlen(m.text);
  int lines = bubbleLines(m.text);
  int h = lines * 10 + 8;
  uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
  int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
  if (w < 30) w = 30;
  int x = m.from_user ? (316 - w) : 4;

  gfx->fillRoundRect(x, y, w, h, 5, bg);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  for (int i = 0; i < lines; i++) {
    char line[CHAT_CHARS + 1] = {0};
    strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
    gfx->setCursor(x + 5, y + 5 + i * 10);
    gfx->print(line);
  }
}

/* Rebuild the whole chat area from history: newest message anchored at the
 * bottom, older ones stacked upward until the area is full. */
void redrawChat() {
  gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
  int shown = min(chat_count, CHAT_HISTORY);
  int y = CHAT_H - 2;
  for (int i = 0; i < shown; i++) {
    ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
    int h = bubbleLines(m.text) * 10 + 8;
    y -= h;
    if (y < 0) break;                          // area full - older ones drop off
    drawOneBubble(m, y);
    y -= 4;
  }
}

void clearChat() {
  chat_count = 0;
  redrawChat();
}

void chatBubble(const char *text, bool from_user) {
  ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
  strncpy(m.text, text, sizeof(m.text) - 1);
  m.text[sizeof(m.text) - 1] = 0;
  m.from_user = from_user;
  chat_count++;
  redrawChat();
}

/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
  gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
  bool up = (WiFi.status() == WL_CONNECTED);
  long rssi = up ? WiFi.RSSI() : -100;
  // -55 dBm or better = full bars; each 10 dB drops one
  int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;

  for (int b = 0; b < 4; b++) {
    int bh = 6 + b * 6;                       // heights 6,12,18,24
    uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
    gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
  }
  gfx->setTextSize(1);
  gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
  if (up) {
    gfx->setTextColor(CYAN);
    gfx->printf("%lddBm", rssi);
  } else {
    gfx->setTextColor(RED);
    gfx->print("DOWN");
  }
  gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
  gfx->setTextColor(gfx->color565(120, 120, 120));
  gfx->print("WiFi");
}

void drawBar(const char *label, uint16_t colour) {
  gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
  gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
  gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
  gfx->setTextSize(2);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
  gfx->print(label);

  // CLEAR wipes the chat history
  gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
  gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
  gfx->print("CLEAR");

  drawWifi();
}


/* ===========================================================================
 *  I2S  —  microphones on port 0, speaker on port 1. Separate hardware
 *  ports, so recording and playback can never fight over a bus.
 * =========================================================================== */
void micInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    /* BOTH channels - this is the two-microphone fix. The vendor examples
     * use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
     * falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
    .channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = false,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_MIC_SCK,
    .ws_io_num = I2S_MIC_WS,
    .data_out_num = I2S_PIN_NO_CHANGE,
    .data_in_num = I2S_MIC_SD
  };
  i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_MIC_PORT, &pins);
}

void spkInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
    .sample_rate = SAMPLE_RATE,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,   // mono - the amp downmixes anyway
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = true,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_SPK_BCLK,
    .ws_io_num = I2S_SPK_LRC,
    .data_out_num = I2S_SPK_DOUT,
    .data_in_num = I2S_PIN_NO_CHANGE
  };
  i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_SPK_PORT, &pins);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
}


/* ===========================================================================
 *  Recording  —  runs while the SPEAK button is held (up to RECORD_MAX_S).
 *  Reads stereo pairs, averages L+R into one mono stream, applies a little
 *  software gain, and fills wav_buf after the 44-byte header slot.
 *  Returns the number of audio bytes recorded.
 * =========================================================================== */
size_t recordWhileHeld() {
  int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
  size_t   mono_samples = 0;
  const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;

  int16_t chunk[512];
  uint32_t last_touch_ok = millis();
  int32_t  peak = 0;                        // loudest raw sample - mic health check

  i2s_zero_dma_buffer(I2S_MIC_PORT);

  while (mono_samples < max_samples) {
    size_t got = 0;
    i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);

#if USE_BOTH_MICS
    size_t n = got / 4;                     // 4 bytes = one L+R pair
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
    size_t n = got / 2;                     // 2 bytes = one mono sample
    for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
      int32_t raw = chunk[i];
#endif
      if (abs(raw) > peak) peak = abs(raw);
      int32_t mixed = raw * MIC_GAIN;
      if (mixed >  32767) mixed =  32767;
      if (mixed < -32768) mixed = -32768;
      mono[mono_samples++] = (int16_t)mixed;
    }

    /* The GT911 is polled between I2S reads. A 250 ms grace period stops a
     * momentary missed touch sample from cutting the recording short. */
    if (speakButtonHeld()) last_touch_ok = millis();
    else if (millis() - last_touch_ok > 250) break;

    // live progress on the button
    static uint32_t last_draw = 0;
    if (millis() - last_draw > 200) {
      last_draw = millis();
      gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
                    (int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
    }
  }

  /* Mic health line: peak as % of full scale BEFORE gain.
   *   0%       = the mic is not being read at all (config/pin problem)
   *   under 3% = too quiet - speak closer or raise MIC_GAIN
   *   3-40%    = healthy speech level
   */
  g_mic_peak_pct = peak * 100.0 / 32768.0;
  Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
                g_mic_peak_pct, MIC_GAIN,
                peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");

  return mono_samples * 2;
}

/* Dump the exact WAV we are about to POST onto the SD card, so it can be
 * played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
  if (!ok_sd) return;
  SD.remove("/stt_debug.wav");
  File f = SD.open("/stt_debug.wav", FILE_WRITE);
  if (!f) return;
  f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
  f.close();
  Serial.println("debug copy saved to SD as /stt_debug.wav");
}

/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
  uint32_t file_len = data_bytes + 36;
  uint32_t byte_rate = SAMPLE_RATE * 2;
  memcpy(h, "RIFF", 4);           memcpy(h + 4, &file_len, 4);
  memcpy(h + 8, "WAVEfmt ", 8);
  uint32_t fmt_len = 16;          memcpy(h + 16, &fmt_len, 4);
  uint16_t fmt = 1, ch = 1;       memcpy(h + 20, &fmt, 2);   memcpy(h + 22, &ch, 2);
  uint32_t rate = SAMPLE_RATE;    memcpy(h + 24, &rate, 4);  memcpy(h + 28, &byte_rate, 4);
  uint16_t align = 2, bits = 16;  memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
  memcpy(h + 36, "data", 4);      memcpy(h + 40, &data_bytes, 4);
}


/* ===========================================================================
 *  HTTP response reader  —  shared by all three cloud calls.
 *  Returns the status code and fills body_out. Handles chunked transfer
 *  encoding PROPERLY: the chunk-size markers must be stripped, or they end
 *  up embedded inside the JSON body and the parse fails on long replies.
 * =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
 * on timeout / closed-and-empty. Every read below goes through this, because
 * a reasoning model can think for many seconds between the response headers
 * and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
  uint32_t t0 = millis();
  while (!c.available()) {
    if (!c.connected()) return false;
    if (millis() - t0 > ms) return false;
    delay(10);
  }
  return true;
}

static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
  body_out = "";

  if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);

  bool chunked = false;
  while (waitData(client, idle_ms)) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;         // blank line = end of headers
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
  }

  if (chunked) {
    int blanks = 0;
    while (true) {
      /* The chunk-size line may not arrive for a long time while the model
       * reasons. Waiting here - instead of letting read() time out - is the
       * whole fix: a timed-out read looks exactly like "0" (final chunk),
       * which silently truncated the body to nothing. */
      if (!waitData(client, idle_ms)) {
        Serial.println("HTTP: timed out waiting for the next chunk");
        break;
      }
      String szline = client.readStringUntil('\n');
      szline.trim();
      if (szline.length() == 0) {                    // stray blank line
        if (++blanks > 4) break;
        continue;
      }
      blanks = 0;

      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;                            // genuine final chunk

      long got = 0;
      while (got < sz) {
        if (!waitData(client, idle_ms)) break;
        while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
      }
      if (waitData(client, 3000)) client.readStringUntil('\n');   // CRLF after chunk
      if (got < sz) { Serial.println("HTTP: short chunk"); break; }
    }
  } else {
    while (waitData(client, idle_ms))
      while (client.available()) body_out += (char)client.read();
  }
  return code;
}


/* ===========================================================================
 *  CLOUD CALL 1  —  Azure speech-to-text
 *  One POST, one header, plain WAV body. This is why Azure does the ears.
 * =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
  /* HTTPClient's one-shot POST fails on bodies this large (it attempts one
   * giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
   * function speaks HTTP directly and streams the WAV up in 4 KB chunks -
   * reliable, and if it ever stalls we know the exact byte it stopped at. */
  WiFiClientSecure client;
  client.setInsecure();                    // no cert bundle on-device; see notes
  client.setTimeout(15);                   // seconds, for reads

  writeWavHeader(wav_buf, audio_bytes);
  dumpWavToSD(audio_bytes);                // PC-playable copy of what we send
  size_t total = WAV_HEADER_LEN + audio_bytes;

  if (!client.connect(AZURE_STT_HOST, 443)) {
    snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
    Serial.println("STT: TLS connect failed");
    return false;
  }

  /* Two valid host forms use DIFFERENT URL paths - detect which one is in
   * secrets.h:  <resource>.cognitiveservices.azure.com -> /stt/speech/...
   *             <region>.stt.speech.microsoft.com      -> /speech/...     */
  bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
  String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
               "/speech/recognition/conversation/cognitiveservices/v1"
               "?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
               "Host: " AZURE_STT_HOST "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
               "Accept: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(total) + "\r\n\r\n";
  client.print(req);

  /* body, 4 KB at a time */
  size_t sent = 0;
  while (sent < total) {
    size_t n = min((size_t)4096, total - sent);
    size_t w = client.write(wav_buf + sent, n);
    if (w == 0) {
      delay(50);                           // brief stall - retry once
      w = client.write(wav_buf + sent, n);
      if (w == 0) {
        snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
                 (unsigned)(sent / 1024));
        Serial.printf("STT: upload stalled at %u/%u bytes\n",
                      (unsigned)sent, (unsigned)total);
        client.stop();
        return false;
      }
    }
    sent += w;
    yield();
  }
  Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);

  /* read the reply with proper de-chunking */
  String resp;
  int code = readHttpResponse(client, resp, 10000);
  client.stop();

  if (code != 200) {
    snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
    Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
    Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
    return false;
  }

  const char *status = doc["RecognitionStatus"];
  if (!status || strcmp(status, "Success") != 0) {
    /* The status names the exact failure:
     *   InitialSilenceTimeout = Azure heard silence (mic level too low)
     *   NoMatch               = heard sound but no recognisable words
     *   BabbleTimeout         = heard only noise                       */
    snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
    Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
    return false;
  }
  text_out = doc["DisplayText"].as<String>();
  if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
  return text_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 2  —  DeepSeek chat completion
 *  OpenAI-compatible format. Model name is deepseek-v4-flash - the old
 *  deepseek-chat name is dead, see secrets.h.
 * =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
  /* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
   * bodies (long answers + the model's hidden reasoning), which shows up as
   * "bad JSON reply". Reading until the server closes the connection is
   * reliable regardless of reply length. */
  JsonDocument req;
  req["model"] = DEEPSEEK_MODEL;
  req["max_tokens"] = LLM_MAX_TOKENS;
  JsonArray msgs = req["messages"].to<JsonArray>();
  JsonObject sys = msgs.add<JsonObject>();
  sys["role"] = "system";  sys["content"] = SYSTEM_PROMPT;
  JsonObject usr = msgs.add<JsonObject>();
  usr["role"] = "user";    usr["content"] = question;

  String body;
  serializeJson(req, body);

  WiFiClientSecure client;
  client.setInsecure();
  /* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
   * ~2 s, but anything that needs actual working-out (an Ohm's law problem,
   * say) can think for 10-30 s before sending a single byte. */
  client.setTimeout(60);

  if (!client.connect(DEEPSEEK_HOST, 443)) {
    snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
    Serial.println("LLM: TLS connect failed");
    return false;
  }

  client.print(String("POST /chat/completions HTTP/1.1\r\n"
               "Host: " DEEPSEEK_HOST "\r\n"
               "Authorization: Bearer " DEEPSEEK_KEY "\r\n"
               "Content-Type: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: ") + String(body.length()) + "\r\n\r\n");
  client.print(body);

  /* read the reply with proper de-chunking; generous window - long
   * questions make the model think for a while before it responds */
  String resp;
  int code = readHttpResponse(client, resp, 60000);
  client.stop();

  if (code != 200) {
    snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
    Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  JsonDocument doc;
  DeserializationError err = deserializeJson(doc, resp);
  if (err) {
    snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
    Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
    return false;
  }

  const char *content = doc["choices"][0]["message"]["content"];
  const char *finish  = doc["choices"][0]["finish_reason"];
  if (!content || !content[0]) {
    /* v4-flash is a reasoning model: if finish_reason is "length", the whole
     * token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
    if (finish && strcmp(finish, "length") == 0)
      snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
    else
      snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
    Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
    return false;
  }
  answer_out = String(content);
  answer_out.trim();
  if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
  return answer_out.length() > 0;
}


/* ===========================================================================
 *  CLOUD CALL 3  —  Azure text-to-speech, streamed straight to the speaker
 *  We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
 *  the I2S peripheral eats. Skip the 44-byte header, forward the rest.
 *  No MP3 decoder, no audio library, no buffering the whole reply.
 * =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
  size_t got = 0;
  uint32_t t0 = millis();
  while (got < n && millis() - t0 < 10000) {
    int r = c.read(dst + got, n - got);
    if (r > 0) { got += r; t0 = millis(); }
    else if (!c.connected() && !c.available()) break;
    else delay(2);
  }
  return got;
}

bool azureTTSSpeak(const String &text) {
  // Escape the XML special characters for the SSML body
  String safe = text;
  safe.replace("&", "&amp;");
  safe.replace("<", "&lt;");
  safe.replace(">", "&gt;");

  String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
                "<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";

  /* Manual HTTP like the other two cloud calls - and for a hard reason:
   * Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
   * raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
   * into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
   * Here we parse the framing properly and keep only clean audio bytes. */
  WiFiClientSecure client;
  client.setInsecure();
  client.setTimeout(20);

  const char *host = AZURE_REGION ".tts.speech.microsoft.com";
  if (!client.connect(host, 443)) {
    Serial.println("TTS: TLS connect failed");
    return false;
  }

  client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
               "Host: ") + host + "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: application/ssml+xml\r\n"
               "X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
               "User-Agent: MaTouchRobojax\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(ssml.length()) + "\r\n\r\n");
  client.print(ssml);

  /* status + headers; note whether the body is chunked */
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
  bool chunked = false;
  long content_len = -1;
  while (client.connected() || client.available()) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
    if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
  }
  if (code != 200) {
    Serial.printf("TTS HTTP %d\n", code);
    client.stop();
    return false;
  }

  const size_t AUDIO_CAP = 1200 * 1024;    // ~37 s of speech
  uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
  if (!audio) { client.stop(); return false; }
  size_t alen = 0;

  if (chunked) {
    /* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
    while (true) {
      String szline = client.readStringUntil('\n');
      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;
      if (alen + sz > AUDIO_CAP) break;
      size_t got = readExact(client, audio + alen, sz);
      alen += got;
      client.readStringUntil('\n');        // trailing CRLF after each chunk
      if (got < (size_t)sz) break;
    }
  } else if (content_len > 0) {
    alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
  } else {
    /* no framing info: read until the server closes */
    uint32_t idle = millis();
    while ((client.connected() || client.available()) && millis() - idle < 5000) {
      int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
      if (r > 0) { alen += r; idle = millis(); }
      else delay(5);
    }
  }
  client.stop();
  Serial.printf("TTS: %u KB clean audio (%s), playing\n",
                (unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");

  bool ok = (alen > WAV_HEADER_LEN);
  if (ok) {
    /* NOW the audio actually starts - this is the honest moment to go green */
    LED_SPEAK();
    drawBar("SPEAKING...", gfx->color565(0, 130, 40));

    /* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
    static const uint8_t lead_in[640] = {0};             // 20 ms of silence
    size_t w = 0;
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
    i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
  }
  free(audio);

  // let the DMA buffers drain so the last word is not cut off
  delay(150);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
  return ok;
}


/* ===========================================================================
 *  SETUP
 * =========================================================================== */
void setup() {
  Serial.begin(115200);
  delay(400);
  Serial.println("\n=== 04 Voice Assistant  |  Robojax.com ===");
  Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");

  pinMode(TFT_BLK, OUTPUT);
  digitalWrite(TFT_BLK, LOW);
  pinMode(SD_CS, OUTPUT);
  digitalWrite(SD_CS, HIGH);

  // one shared SPI bus for TFT + SD (started before either device)
  SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);

  gfx->begin();
  gfx->fillScreen(BLACK);
  digitalWrite(TFT_BLK, HIGH);

  // SD is optional here - it only stores the /stt_debug.wav diagnostic copy
  ok_sd = SD.begin(SD_CS, SPI, 20000000);
  digitalWrite(SD_CS, HIGH);
  Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
                       : "SD not found - debug WAV dump disabled (not fatal)");

  bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
  delay(50);

  rgb.begin();
  rgb.setBrightness(LED_BRIGHTNESS);
  LED_IDLE();

  /* One recording buffer for the whole session, in PSRAM. This is the 8 MB
   * that makes the board worth buying. */
  wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
  if (!wav_buf) {
    gfx->setTextColor(RED);
    gfx->setTextSize(2);
    gfx->setCursor(10, 100);
    gfx->print("PSRAM alloc failed!");
    gfx->setTextSize(1);
    gfx->setCursor(10, 130);
    gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
    while (1) delay(1000);
  }

  micInit();
  spkInit();

  gfx->setTextSize(1);
  gfx->setTextColor(YELLOW);
  gfx->setCursor(4, 4);
  gfx->printf("Connecting to %s ...", WIFI_SSID);
  Serial.printf("Connecting to %s ", WIFI_SSID);

  WiFi.mode(WIFI_STA);
  WiFi.begin(WIFI_SSID, WIFI_PASS);
  uint32_t t0 = millis();
  while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
    delay(300);
    Serial.print(".");
  }
  Serial.println();

  clearChat();
  if (WiFi.status() == WL_CONNECTED) {
    Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
    chatBubble("Hold SPEAK and ask me anything.", false);
  } else {
    chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
    LED_ERROR();
  }
  drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}


/* ===========================================================================
 *  LOOP  —  one full conversation turn per button press
 * =========================================================================== */
void loop() {
  /* CLEAR button: edge-detected so one tap wipes once. Reading the panel
   * twice per loop (here and in speakButtonHeld) is fine - the GT911 just
   * reports its current state. */
  static bool tap_latch = false;
  static uint8_t tap_release = 0;
  if (state == ST_IDLE) {
    uint16_t cx, cy;
    if (getTouch(&cx, &cy)) {
      tap_release = 0;
      if (!tap_latch) {
        tap_latch = true;
        if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
          clearChat();
          chatBubble("Hold SPEAK and ask me anything.", false);
        }
      }
    } else if (tap_latch && ++tap_release >= 4) {
      tap_latch = false;
      tap_release = 0;
    }

    /* live WiFi signal indicator, refreshed every 2 s while idle */
    static uint32_t last_wifi = 0;
    if (millis() - last_wifi > 2000) {
      last_wifi = millis();
      drawWifi();
    }
  }

  if (state == ST_IDLE && speakButtonHeld()) {

    /* ---- record ---- */
    state = ST_RECORDING;
    LED_LISTEN();
    drawBar("LISTENING...", gfx->color565(0, 60, 200));
    uint32_t t_rec = millis();
    size_t audio_bytes = recordWhileHeld();
    t_rec = millis() - t_rec;
    Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);

    if (audio_bytes < SAMPLE_RATE / 2) {   // under a quarter second - a tap
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }

    /* ---- speech to text ---- */
    state = ST_STT;
    LED_THINK();
    drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
    uint32_t t_stt = millis();
    String question;
    if (!azureSTT(audio_bytes, question)) {
      /* Show the REAL cause on screen - no serial monitor needed. */
      char diag[96];
      snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
               g_stt_err, g_mic_peak_pct,
               ok_sd ? " | saved /stt_debug.wav" : "");
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_stt = millis() - t_stt;
    chatBubble(question.c_str(), true);
    Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());

    /* ---- think ---- */
    state = ST_LLM;
    drawBar("THINKING...", gfx->color565(150, 90, 0));
    uint32_t t_llm = millis();
    String answer;
    if (!deepseekChat(question, answer)) {
      char diag[96];
      snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
      chatBubble(diag, false);
      drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
      LED_IDLE();
      state = ST_IDLE;
      return;
    }
    t_llm = millis() - t_llm;
    chatBubble(answer.c_str(), false);
    Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());

    /* ---- speak ----
     * Still amber here: the voice has to be synthesised and downloaded first
     * (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
     * at the exact moment audio starts coming out of the speaker. */
    state = ST_TTS;
    LED_THINK();
    drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
    uint32_t t_tts = millis();
    bool spoke = azureTTSSpeak(answer);
    t_tts = millis() - t_tts;

    /* Timing summary on serial - this feeds the "honest numbers" segment. */
    Serial.printf("TIMINGS  rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
                  audio_bytes / 32000.0, (unsigned long)t_stt,
                  (unsigned long)t_llm, (unsigned long)t_tts,
                  spoke ? "" : "  (TTS FAILED)");

    drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
    LED_IDLE();
    state = ST_IDLE;
  }

  delay(20);
}

آپ کو جن چیزوں کی ضرورت ہو سکتی ہے

وسائل اور حوالہ جات

فائلیں📁

ضروری فائل (.h)

  • secrets.h
    file for Makerfabs MaTouch AI ESP32S3 2.8" TFT Camera module
    secrets.h 0.01 MB

دیگر فائلیں

  • pins.h
    pins file for MaTouch AI ESP32S3 2.8" camera LCD touch screen.
    pins.h 0.01 MB

نقشہ

  • MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematic
    The latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
    MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF 0.15 MB