This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
The latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
เมกเกอร์แฟบส์ MaTouch ESP32-S3 2.8 นิ้ว กล้อง สร้างผู้ช่วยเสียง AI บน ESP32-S3 (Azure + DeepSeek)
กดปุ่มค้างไว้ ถามคำถาม แล้วบอร์ดจะตอบออกมาเป็นเสียง
ผู้ช่วยเสียงแบบครบวงจรบนบอร์ดขนาดเล็กเพียงหนึ่งเดียว กดปุ่ม SPEAK ค้างไว้แล้วถามคำถาม บอร์ดจะบันทึกเสียงคุณด้วยไมโครโฟนทั้งสองตัว ส่งเสียงไปยัง Microsoft Azure เพื่อแปลงเป็นข้อความ ส่งข้อความนั้นไปยัง DeepSeek เพื่อคิดคำตอบ ส่งคำตอบกลับไปยัง Azure เพื่อแปลงเป็นเสียงพูด และเล่นผ่านลำโพงในตัว บทสนทนาทั้งหมดจะปรากฏบนหน้าจอเป็นฟองแชท
ผู้ช่วยเสียง AI แบบพูดคุยที่ทำงานบนบอร์ด MaTouch AI ESP32-S3
สิ่งที่เกิดขึ้นตั้งแต่วินาทีที่คุณกดปุ่ม
นี่คือเส้นทางทั้งหมดของคำถามหนึ่งข้อ ทีละขั้นตอน คุ้มค่าที่จะอ่านสักครั้ง เพราะทุกสิ่งที่คุณเห็นบนหน้าจอและบน LED จะสอดคล้องกับขั้นตอนใดขั้นตอนหนึ่งเหล่านี้
คุณกดและกดค้างปุ่ม SPEAK เป็นแบบกดค้างเพื่อพูด ไม่ใช่แตะเพื่อพูด: การบันทึกจะทำงานตราบเท่าที่นิ้วของคุณกดอยู่ สูงสุดหกวินาที ไฟ LED สถานะจะเปลี่ยนเป็นสีน้ำเงิน และปุ่มจะแสดงคำว่า LISTENING
ไมโครโฟนทั้งสองตัวบันทึกเสียงคุณ บอร์ดสุ่มตัวอย่าง 16,000 ครั้งต่อวินาทีจากคู่สเตอริโอ หาค่าเฉลี่ยของสองช่องสัญญาณเป็นช่องเดียว ปรับเกนเล็กน้อย และเก็บผลลัพธ์ไว้ใน PSRAM แถบความคืบหน้าจะค่อยๆ เคลื่อนผ่านปุ่มขณะที่คุณพูด เสียงพูดสองวินาทีมีขนาดประมาณ 64 KB
คุณปล่อยปุ่ม การบันทึกหยุดลง บอร์ดเขียนส่วนหัว WAV ขนาด 44 ไบต์ไว้ที่ด้านหน้าของไฟล์เสียง - ป้ายกำกับเล็กๆ นั้นคือสิ่งที่เปลี่ยนตัวอย่างเสียงดิบให้เป็นไฟล์ที่ Azure จะยอมรับ
เสียงถูกส่งไปยัง Azure Speech-to-Text ไฟล์ถูกอัปโหลดเป็นชิ้นส่วนขนาด 4 KB ผ่านการเชื่อมต่อที่ปลอดภัย และกลับมาเป็นข้อความหนึ่งบรรทัด เสียง 64 KB ของคุณกลายเป็นข้อความประมาณ 25 ไบต์ LED เปลี่ยนเป็นสีเหลืองอำพัน
คำถามของคุณปรากฏบนหน้าจอ เป็นฟองแชทสีน้ำเงิน เพื่อให้คุณเห็นได้ชัดเจนว่าระบบได้ยินอะไร - ซึ่งมีประโยชน์ เพราะคำที่ฟังผิดอธิบายคำตอบแปลกๆ ส่วนใหญ่ได้
ข้อความถูกส่งไปยัง DeepSeek บอร์ดส่งคำถามของคุณพร้อมคำสั่งประจำให้จำกัดคำตอบไว้ที่สองประโยคสั้นๆ โมเดลจะคิด - คิดจริงๆ เพราะเป็นโมเดลแบบใช้เหตุผล - และส่งคำตอบกลับมา
คำตอบปรากฏบนหน้าจอ เป็นฟองสีเทา คุณสามารถอ่านได้ก่อนที่จะได้ยิน
คำตอบถูกส่งกลับไปยัง Azure เพื่อแปลงเป็นเสียงพูด บอร์ดขอเสียง PCM ดิบที่ 16 kHz ซึ่งเป็นรูปแบบที่แอมพลิฟายเออร์ต้องการพอดี ดังนั้นจึงไม่มีตัวถอดรหัส MP3 ในโปรเจกต์นี้เลย ปุ่มตอนนี้แสดงคำว่า GETTING VOICE และไฟ LED ยังคงเป็นสีเหลืองอำพัน เพราะยังไม่มีเสียงใดๆ ให้ได้ยิน
คลิปเสียงทั้งหมดถูกดาวน์โหลดลงใน PSRAM ก่อนที่จะเล่นตัวอย่างเสียงแม้แต่ตัวอย่างเดียว สิ่งนี้สำคัญ - ดูหมายเหตุด้านล่าง
การเล่นเสียง ทันทีที่เสียงถูกส่งไปยังลำโพง LED จะเปลี่ยนเป็นสีเขียว และปุ่มจะแสดงคำว่า SPEAKING คุณจะได้ยินคำตอบ
เหตุใดจึงดาวน์โหลดเสียงก่อนแทนที่จะเล่นทันทีที่มาถึง การสตรีมจากเครือข่ายตรงเข้าลำโพงจะให้เสียงเหมือนเสียงเคาะ บัฟเฟอร์ของลำโพงเก็บเสียงได้เพียงประมาณหนึ่งในสิบของวินาที และทุกครั้งที่การถ่ายโอน WiFi หยุดชะงักนานกว่านั้น บัฟเฟอร์จะว่างเปล่า ทำให้เกิดเสียงเคาะที่ได้ยิน การดาวน์โหลดคำตอบทั้งหมดลงใน PSRAM ก่อนต้องใช้เวลารอเพิ่มอีกประมาณหนึ่งวินาที แต่ช่วยขจัดช่องว่างทั้งหมด นั่นคือเหตุผลที่จอแสดงผลบอกว่า GETTING VOICE ก่อนที่จะบอกว่า SPEAKING - ทั้งสองเป็นขั้นตอนที่แตกต่างกันจริงๆ
แต่ละขั้นตอนใช้เวลานานเท่าใด
วัดจากฮาร์ดแวร์จริง สำหรับคำถามง่ายๆ:
ขั้นตอน | เวลาโดยทั่วไป |
|---|---|
การบันทึกเสียง | ตราบเท่าที่คุณกดปุ่มค้างไว้ |
การแปลงเสียงเป็นข้อความ (Azure) | ประมาณ 1.8 วินาที |
การคิด (DeepSeek) | ประมาณ 1.8 วินาทีสำหรับคำถามง่ายๆ นานกว่านั้นมากสำหรับคำถามที่ต้องคำนวณจริง |
การดึงเสียง (Azure) | ประมาณ 7 วินาที - เป็นส่วนที่ใหญ่ที่สุด |
รวมทั้งหมด ตั้งแต่ปล่อยปุ่มจนถึงเสียงแรก | ประมาณ 11 วินาที |
ทุกการสนทนาจะพิมพ์เวลาของตัวเองลงในซีเรียลมอนิเตอร์ เพื่อให้คุณวัดค่าได้ด้วยตัวเองแทนที่จะเชื่อตัวเลขเหล่านี้ หากต้องการให้เร็วขึ้น การเปลี่ยนแปลงที่มีประสิทธิภาพที่สุดคือขอคำตอบที่สั้นลงใน SYSTEM_PROMPT - ข้อความที่ต้องพูดน้อยลงหมายถึงเสียงที่ต้องสังเคราะห์และดาวน์โหลดน้อยลง
ตัวบอร์ดเองไม่เคยเข้าใจอะไรเลย มันเป็นเพียงผู้ส่งสารที่มีหูที่ดีและเสียงที่ดี - ความฉลาดถูกเช่าเป็นรายวินาที
เหตุใดจึงใช้บริการทั้งสามนี้
Azure จัดการเสียงพูดเข้าและออก การแปลงข้อความเป็นเสียงสามารถส่งคืนเสียง PCM ดิบที่ 16 kHz ซึ่งเป็นสิ่งที่ชิปลำโพงต้องการพอดี ดังนั้นจึงไม่มีตัวถอดรหัส MP3 ในโปรเจกต์นี้เลย การแปลงเสียงเป็นข้อความรับไฟล์ WAV ธรรมดาผ่านการ POST ธรรมดา
DeepSeek เป็นสมองของการสนทนา มันเร็วและมีค่าใช้จ่ายเพียงเศษเสี้ยวของเซนต์ต่อคำตอบ
OpenAI ไม่ได้ใช้ในโปรเจกต์นี้ - ดูโปรเจกต์ 05 ซึ่งใช้สำหรับงานด้านภาพ
ชื่อโมเดล DeepSeek เปลี่ยนไป ชื่อเก่า deepseek-chat และ deepseek-reasoner ถูกยกเลิกในเดือนกรกฎาคม 2026 บทช่วยสอนส่วนใหญ่ในอินเทอร์เน็ตยังคงใช้ชื่อเหล่านี้และจะส่งคืนข้อผิดพลาด ชื่อปัจจุบันคือ deepseek-v4-flash และ deepseek-v4-pro โปรเจกต์นี้ใช้ v4-flash
กับดักของโมเดลแบบใช้เหตุผล
DeepSeek v4-flash คิดก่อนที่จะตอบ และการคิดนั้นนับรวมอยู่ในโควต้าโทเค็นของคุณ หากตั้งค่า LLM_MAX_TOKENS ต่ำเกินไป งบประมาณทั้งหมดจะหมดไปกับการใช้เหตุผล คำตอบจะกลับมาว่างเปล่า และบอร์ดจะไม่พูดอะไรเลย ด้วยเหตุนี้จึงตั้งไว้ที่ 400 ที่นี่ คำถามยากๆ ก็ใช้เวลานานกว่าเช่นกัน - คำตอบข้อเท็จจริงง่ายๆ ใช้เวลาประมาณสองวินาที แต่คำถามที่ต้องคำนวณจริงอาจใช้เวลานานกว่านั้นมาก
การอ่านไฟแสดงสถานะ
สี | ความหมาย |
|---|---|
สีน้ำเงิน | กำลังฟังคุณ |
สีเหลืองอำพัน | คลาวด์กำลังคิด หรือกำลังดึงเสียง |
สีเขียว | กำลังพูด - จะเปลี่ยนเป็นสีเขียวในจังหวะที่เสียงเริ่มออกพอดี |
สีแดง | มีบางอย่างล้มเหลว - ตรวจสอบซีเรียลมอนิเตอร์ |
ตัวควบคุมบนหน้าจอ
แชทจะเลื่อนเหมือนการสนทนาทางโทรศัพท์ ข้อความเก่าจะเลื่อนขึ้นและออกไป CLEAR จะล้างทั้งหมด มาตรวัดสัญญาณ WiFi พร้อมค่า dBm จริงอยู่ที่มุม ซึ่งมีประโยชน์เมื่อคุณสงสัยว่าการตอบช้าเกิดจากเครือข่ายหรือบริการ
เกี่ยวกับบอร์ด MaTouch AI ESP32-S3 2.8"
ทุกโปรเจกต์ในหน้านี้ทำงานบน MaTouch AI ESP32-S3 2.8" TFT ST7789V จาก Makerfabs เป็นบอร์ดแบบครบวงจร: หน้าจอสัมผัสสี กล้อง 3 เมกะพิกเซล ไมโครโฟนสองตัว และแอมป์ลำโพงจริง ทั้งหมดขับเคลื่อนด้วย ESP32-S3 ที่มี PSRAM 8 MB การผสมผสานนี้เองที่ทำให้โปรเจกต์ AI เหล่านี้เป็นไปได้บนบอร์ดเดียวโดยไม่ต้องต่ออุปกรณ์อื่นเพิ่ม
PSRAM 8 MB มีความสำคัญมากกว่าตัวเลขอื่นๆ ทั้งหมดในหน้านี้ มันคือสิ่งที่ทำให้บอร์ดสามารถเก็บเฟรมกล้อง เสียงที่บันทึกไว้ไม่กี่วินาที หรือรูปภาพที่เข้ารหัส base64 ในหน่วยความจำพร้อมกันได้ - ซึ่งไม่มีสิ่งใดที่พอดีกับ RAM ปกติของ ESP32
เอกสารจากผู้ผลิต: หน้า wiki ของ Makerfabs
ข้อมูลจำเพาะหลัก
โปรเซสเซอร์: ESP32-S3, dual core 240 MHz, WiFi 2.4 GHz + Bluetooth 5.0
หน่วยความจำ: แฟลช 16 MB, PSRAM 8 MB (จำเป็นสำหรับเกือบทุกโปรเจกต์ในหน้านี้)
จอแสดงผล: IPS 2.8", 320×240, ไดรเวอร์ ST7789V, SPI
ระบบสัมผัส: GT911 แบบ capacitive รองรับ 5 นิ้วพร้อมกัน
กล้อง: OV3660, 3 เมกะพิกเซล, สูงสุด 2048×1536
ไมโครโฟน: ไมค์ดิจิทัล I2S INMP441 สองตัว (คู่สเตอริโอแท้)
ลำโพง: แอมป์คลาส-D MAX98357A, 3.2 W ที่ 4 Ω
ที่เก็บข้อมูล: ช่องการ์ด microSD (โหมด SPI)
พลังงาน: USB-C, ขั้วต่อแบตเตอรี่ JST, เครื่องชาร์จ TP4056, สวิตช์เปิดปิด
อื่นๆ บนบอร์ด: LED RGB WS2812B, นาฬิกาเวลาจริง PCF8563T ที่ใช้แบตเตอรี่สำรอง และเกจวัดแบตเตอรี่ MAX17048 ที่ ไม่ได้ ระบุในข้อมูลจำเพาะอย่างเป็นทางการ
พอร์ต USB-C สองพอร์ตไม่เหมือนกัน ลำโพงของบอร์ดใช้พินสัญญาณร่วม (IO19 และ IO20) กับพอร์ต USB เนทีฟ เนื่องจากพินเหล่านั้นคือสายข้อมูล USB ที่ถูกเดินสายตายตัวของ ESP32-S3 อัปโหลดและจ่ายไฟผ่าน พอร์ต USB-C CH340K เสมอ (พอร์ตที่อยู่ข้างปุ่ม RESET) และตั้งค่า USB CDC On Boot เป็น Disabled หากใช้พอร์ตผิด เสียงจะทำงานผิดปกติหรือการอัปโหลดจะล้มเหลว
การตั้งค่า Arduino IDE
การตั้งค่าเหล่านี้สำคัญมาก ปัญหาส่วนใหญ่ที่ผู้คนรายงานเกี่ยวกับบอร์ดนี้เกิดจากการตั้งค่าอย่างใดอย่างหนึ่งผิด และการตั้งค่าเหล่านี้จะรีเซ็ตเมื่อคุณเปลี่ยนเวอร์ชันคอร์ ดังนั้นควรตรวจสอบอีกครั้งหลังการเปลี่ยนแปลงใดๆ
การตั้งค่า | ค่า |
|---|---|
บอร์ด | ESP32S3 Dev Module |
เวอร์ชันคอร์ ESP32 | 2.0.17 |
PSRAM | OPI PSRAM |
ขนาดแฟลช | 16MB (128Mb) |
รูปแบบพาร์ติชัน | 16M Flash (3MB APP/9.9MB FATFS) |
USB CDC On Boot | Disabled |
ความเร็วอัปโหลด | 921600 |
ล้างแฟลชทั้งหมดก่อนอัปโหลด | Disabled |
พอร์ต | พอร์ต USB-C CH340K |
ใช้ ESP32 core 2.0.17 ไม่ใช่ 3.x Espressif นำโมเดลตรวจจับใบหน้าบนอุปกรณ์ออกใน core 3 ดังนั้นโปรเจกต์เกี่ยวกับใบหน้าจะคอมไพล์ไม่ได้ที่นั่น การล็อกไว้ที่ 2.0.17 ทำให้ทุกโปรเจกต์ในหน้านี้ทำงานได้ด้วยการกำหนดค่าเดียว ใน Boards Manager เมนูเลือกเวอร์ชันให้คุณสลับไปมาได้ตามต้องการ
ใช้ GFX Library for Arduino เวอร์ชัน 1.5.6 ไม่ใช่ 1.6.x เวอร์ชัน 1.6 สร้างขึ้นสำหรับ ESP32 core 3 และอาจค้างตอนเริ่มต้นบน core 2.0.17 หากหน้าจอยังคงดำหลังอัปโหลด นี่คือสิ่งแรกที่ควรตรวจสอบ
ไลบรารีที่จำเป็น
ติดตั้งผ่าน Tools → Manage Libraries ใน Arduino IDE หมายเลขเวอร์ชันมีความสำคัญ - กรุณาใช้เวอร์ชันที่ระบุไว้
ไลบรารี | เวอร์ชัน | ผู้เขียน |
|---|---|---|
GFX Library for Arduino | 1.5.6 | moononournation |
bb_captouch | 1.3.1 | Larry Bank |
ArduinoJson | 7.x | Benoit Blanchon |
Adafruit NeoPixel | เวอร์ชันล่าสุดใดก็ได้ | Adafruit |
การตั้งค่า secrets.h
รายละเอียด WiFi และคีย์ API ใดๆ ของคุณจะอยู่ใน secrets.h ซึ่งรวมอยู่ในไฟล์ดาวน์โหลดพร้อมค่าตัวอย่าง เปิดแท็บนั้นใน Arduino IDE และแทนที่ด้วยค่าของคุณเอง
WiFi ต้องเป็น 2.4 GHz ESP32-S3 ไม่สามารถมองเห็นเครือข่าย 5 GHz ได้เลย หากเราเตอร์ของคุณรวมทั้งสองย่านความถี่ไว้ภายใต้ชื่อเดียว (Asus เรียกสิ่งนี้ว่า Smart Connect) ให้ปิดฟีเจอร์นั้นหรือตั้งชื่อเฉพาะให้กับย่าน 2.4 GHz แล้วใช้ชื่อนั้นใน secrets.h
การรับคีย์ API ของคุณ
โปรเจกต์นี้สื่อสารกับบริการ AI บนคลาวด์ ดังนั้นคุณต้องมีคีย์ของตัวเอง หากคุณไม่เคยทำสิ่งนี้มาก่อน ไม่ต้องกังวล - มันเป็นแนวคิดเดียวกับรหัสผ่านที่ระบุบัญชีของคุณให้กับบริการ ใช้เวลาไม่กี่นาที เพียงครั้งเดียว
คีย์ไม่ใช่การสมัครสมาชิกเว็บไซต์ การจ่ายเงินสำหรับ ChatGPT Plus เป็นต้น ไม่ได้ ให้คีย์ API แก่คุณ - ทั้งสองเป็นผลิตภัณฑ์แยกกันที่มีการเรียกเก็บเงินแยกกัน คุณต้องมีบัญชีบนแพลตฟอร์มนักพัฒนา ตามที่อธิบายด้านล่าง
Microsoft Azure Speech - สำหรับการฟังและการพูด
Azure แปลงคำพูดของคุณเป็นข้อความและแปลงคำตอบกลับเป็นเสียง ระดับฟรีเพียงพอสำหรับทุกอย่างในหน้านี้
ไปที่ portal.azure.com และลงชื่อเข้าด้วยบัญชี Microsoft (บัญชีฟรีก็ใช้ได้)
หากคุณไม่เคยใช้ Azure มาก่อน คุณจะเห็นหน้าจอ Welcome to Azure ที่เสนอสามตัวเลือก เลือก Start with an Azure free trial - คุณต้องมีการสมัครสมาชิกก่อนที่ Azure จะให้คุณสร้างสิ่งใดได้ (นักเรียนควรเลือก Azure for Students แทน: ผลลัพธ์เดียวกัน ไม่ต้องใช้บัตร) ไม่ต้องสนใจ Manage Microsoft Entra ID ซึ่งเป็นคนละสิ่งโดยสิ้นเชิง
คลิก Create a resource ค้นหา Speech และเลือก Speech service ที่เผยแพร่โดย Microsoft
กรอกแบบฟอร์ม: กลุ่มทรัพยากรใดก็ได้ ชื่อใดก็ได้ และเลือก Region ที่ใกล้คุณ - จดภูมิภาคนั้นตามที่ปรากฏทุกตัวอักษร เช่น
eastusสำหรับ Pricing tier เลือก F0 (Free) ซึ่งอนุญาตให้ใช้การแปลงคำพูดเป็นข้อความประมาณห้าชั่วโมงและการแปลงข้อความเป็นคำพูดครึ่งล้านตัวอักษรต่อเดือน
คลิก Review + create จากนั้น Create รอประมาณหนึ่งนาที จากนั้นคลิก Go to resource
ในเมนูด้านซ้าย เปิด Keys and Endpoint คัดลอก KEY 1 และ Location/Region
ใส่ค่าเหล่านั้นลงใน secrets.h เป็น AZURE_SPEECH_KEY และ AZURE_REGION สำหรับ AZURE_STT_HOST ให้ใช้ <region>.stt.speech.microsoft.com - ดังนั้นหากภูมิภาคคือ eastus ก็จะเป็น eastus.stt.speech.microsoft.com
เกี่ยวกับบัตรเครดิต การทดลองใช้ฟรีของ Azure ขอบัตรเพื่อยืนยันตัวตนของคุณ มันจะไม่เรียกเก็บเงินจากคุณ คุณได้รับเครดิต $200 เป็นเวลา 30 วัน และหลังจากนั้นบัญชีจะเปลี่ยนเป็น Pay-As-You-Go - แต่ ระดับ F0 Speech ยังคงฟรี เดือนแล้วเดือนเล่า และทุกอย่างในโปรเจกต์เหล่านี้อยู่ในขอบเขตนั้นอย่างสบายๆ หากคุณไม่ต้องการให้บัตรเลยและเป็นนักเรียน ตัวเลือก Azure for Students จะให้เครดิตโดยไม่ต้องใช้บัตร
ต้องเป็นทรัพยากร "Speech service" คีย์จากทรัพยากร Translator, Language หรือ Cognitive Services ทั่วไปมีลักษณะเหมือนกันและใช้ได้อย่างสมบูรณ์ - แต่ทุกคำขอ speech จะส่งคืนข้อผิดพลาด 401 สิ่งนี้ทำให้เราติดปัญหาในระหว่างการทดสอบและเสียเวลาไปหนึ่งชั่วโมง หาก speech ล้มเหลวด้วยรหัส 401 ในขณะที่คีย์ดูถูกต้อง ให้ตรวจสอบว่าคุณสร้างทรัพยากรประเภทใด
DeepSeek - ส่วนที่ใช้คิด
DeepSeek คือโมเดลภาษาที่ตอบคำถามของคุณจริงๆ มันมีราคาถูก - เครดิตไม่กี่ดอลลาร์ครอบคลุมการตอบกลับหลายพันครั้ง
ไปที่ platform.deepseek.com และสร้างบัญชี
เปิด API keys ในเมนูและคลิก Create new API key
คัดลอกทันที คีย์จะแสดงเพียงครั้งเดียวและจะไม่แสดงอีก - หากทำหาย ให้ลบคีย์นั้นและสร้างใหม่
เพิ่มเครดิตจำนวนเล็กน้อยภายใต้ Top up ไม่มีระดับฟรี แต่การเติมเงินขั้นต่ำใช้งานได้นานมากที่อัตราการใช้งานนี้
ใส่คีย์ลงใน secrets.h เป็น DEEPSEEK_KEY คีย์เริ่มต้นด้วย sk-
ชื่อโมเดลเปลี่ยนในเดือนกรกฎาคม 2026 โมเดลเก่า deepseek-chat และ deepseek-reasoner ถูกยกเลิกแล้ว ดังนั้นบทช่วยสอนส่วนใหญ่ที่คุณพบทางออนไลน์จะล้มเหลวด้วยข้อผิดพลาด 400 ให้ใช้ deepseek-v4-flash ซึ่งเป็นสิ่งที่โปรเจกต์เหล่านี้ตั้งค่าไว้แล้ว
ค่าใช้จ่ายในการรัน
น้อยมาก แต่ไม่ฟรี และคุณควรทราบคร่าวๆ ว่าคุณใช้จ่ายเท่าไรก่อนที่จะปล่อยให้โปรเจกต์ทำงานต่อไป
บริการ | ค่าใช้จ่ายโดยประมาณ |
|---|---|
Azure Speech | ระดับฟรีครอบคลุมการฟังประมาณ 5 ชั่วโมงและการพูด 0.5 ล้านตัวอักษรต่อเดือน |
DeepSeek | เศษเสี้ยวของเซนต์ต่อคำตอบ - การตอบกลับหลายพันครั้งในราคาไม่กี่ดอลลาร์ |
OpenAI vision | ประมาณหนึ่งหรือสองเซนต์ต่อรูปภาพ ขึ้นอยู่กับโมเดล |
ราคาเปลี่ยนแปลงได้ ดังนั้นให้ถือว่าเป็นแนวทางมากกว่าข้อเสนอราคา ทุกบริการเหล่านี้มีหน้าอัตราการใช้งานที่คุณสามารถดูยอดใช้จ่ายได้ และทั้งหมดอนุญาตให้คุณตั้งขีดจำกัดการใช้จ่าย - ซึ่งคุ้มค่าที่จะทำตั้งแต่วันแรก
เก็บคีย์ของคุณไว้เป็นความลับ ใครก็ตามที่มีคีย์เหล่านี้สามารถใช้เงินของคุณได้ อย่าใส่ไว้ในวิดีโอ ภาพหน้าจอ โพสต์ในฟอรัม หรือที่เก็บโค้ดสาธารณะ หากคีย์ถูกเปิดเผยไม่ว่าด้วยวิธีใด ให้ลบทิ้งบนเว็บไซต์ของผู้ให้บริการและสร้างคีย์ใหม่ - ใช้เวลาเพียงไม่กี่วินาที และนี่คือทางแก้อันเดียวที่แท้จริง
การแก้ไขปัญหา
อาการ | สาเหตุและวิธีแก้ไข |
|---|---|
หน้าจอค้างเป็นสีดำ | เวอร์ชันไลบรารี GFX ผิด (ใช้ 1.5.6) หรือการตั้งค่าบอร์ดผิด |
|
|
ไม่มีการอัปโหลด / ไม่มีพอร์ต COM | พอร์ต USB-C ผิด หรือไม่ได้ติดตั้งไดรเวอร์ CH340 |
กล้องล้มเหลวและไม่ฟื้นคืน | สายรีเซ็ตของกล้องเชื่อมต่อกับปุ่ม RESET ของบอร์ด ดังนั้นซอฟต์แวร์จึงไม่สามารถรีสตาร์ทได้ กด RESET หากยังล้มเหลว ให้เสียบสายแพของกล้องใหม่ |
ดาวน์โหลดโค้ด
สเก็ตช์ Arduino ที่สมบูรณ์สำหรับโปรเจกต์นี้ พร้อมด้วย pins.h และไฟล์อื่นๆ ที่จำเป็นทั้งหมด ให้ดาวน์โหลดฟรี
แตกไฟล์ เปิดไฟล์ .ino ใน Arduino IDE ตรวจสอบการตั้งค่าด้านบน แล้วอัปโหลดผ่านพอร์ต USB-C CH340K
This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
/*
* ===========================================================================
* 04_Voice_Assistant — MaTouch AI ESP32-S3 2.8" TFT ST7789V
* ===========================================================================
*
----------
* ROBOJAX.COM - MaTouch AI ESP32-S3 2.8" project series
*
* WATCH THE VIDEO
* https://youtu.be/6AL3g3tC_Hk
*
* WRITTEN TUTORIALS - every project, with photos and full explanation
* Camera and touchscreen.... https://robojax.com/RTJ849
* Offline face recognition.. https://robojax.com/RTJ850
* AI voice assistant........ https://robojax.com/RTJ851
* AI vision................. https://robojax.com/RTJ852
*
* GET THE BOARD - SAVE $5 with coupon code: Robojax_Makerfab
* https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
* (enter the code at checkout)
*
* All of this code is free. If it helped you, a subscribe on YouTube is
* the best way to support more of it.
*
* ---------------------------------------------------------------------------
*
* A complete voice assistant on a $40 board:
*
* hold SPEAK -> both INMP441 microphones record you
* -> Azure Speech turns the audio into text
* -> DeepSeek v4-flash thinks of an answer
* -> Azure Speech turns the answer into audio
* -> the MAX98357 speaker says it out loud
*
* and the whole conversation is drawn as chat bubbles on the touchscreen.
*
* WHY THIS COMBINATION OF SERVICES (each is used where it is best):
* - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
* transcription endpoint wants multipart/form-data - miserable on an MCU.
* - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
* which streams straight into the I2S speaker with NO MP3 decoder at all.
* - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
* model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
* Most tutorials online still use them and are broken. See secrets.h.
*
* A detail the vendor examples get wrong: this board has TWO microphones on
* one I2S bus (left + right), but every Makerfabs demo records left-only and
* throws one away. This sketch records both and averages them.
*
* ---------------------------------------------------------------------------
* *** WHICH USB PORT - THIS MATTERS ***
* The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
* through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
* If you use the wrong port the audio will be garbage or uploads will fail.
* ---------------------------------------------------------------------------
*
* FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
*
* BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
* or compile errors. These reset when you switch cores - recheck them!)
*
* Board : ESP32S3 Dev Module
* ESP32 core : 2.0.17
* PSRAM : OPI PSRAM <-- required, audio buffer lives there
* Flash Size : 16MB (128Mb)
* Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
* USB CDC On Boot : Disabled <-- required, see USB note above
* Upload Speed : 921600
* Port : the CH340K USB-C port (the one near RESET)
*
* LIBRARIES
* GFX Library for Arduino v1.5.6 (NOT 1.6.x - that pairs with core 3)
* bb_captouch v1.3.1
* ArduinoJson v7.x
* Adafruit NeoPixel any recent
*
* ---------------------------------------------------------------------------
* FUNCTIONS IN THIS SKETCH
* led(r,g,b) + LED_* macros RGB status colours (blue/amber/green/red)
* getTouch(&x,&y) read the touch panel, mapped to screen coordinates
* speakButtonHeld() true while a finger is on the SPEAK button
* bubbleLines(t) how many lines a message wraps to
* drawOneBubble(m,y) draw a single chat bubble
* redrawChat() rebuild the chat area from history, newest at bottom
* clearChat() wipe the chat history (CLEAR button)
* chatBubble(t,user) add a message to history and redraw
* drawWifi() WiFi signal bars + dBm readout
* drawBar(label,col) bottom bar: SPEAK button + CLEAR + WiFi meter
* micInit() I2S input - BOTH INMP441 mics, stereo
* spkInit() I2S output - MAX98357 speaker
* recordWhileHeld() record while SPEAK held, downmix stereo->mono
* writeWavHeader(...) prepend the 44-byte RIFF/WAVE header
* dumpWavToSD(...) save the exact upload to SD (/stt_debug.wav)
* readHttpResponse() read an HTTPS reply, de-chunking it properly
* azureSTT(...) chunked upload of the WAV -> recognised text
* deepseekChat(...) question -> deepseek-v4-flash -> answer text
* azureTTSSpeak(text) answer -> Azure voice -> PSRAM -> speaker
* setup() / loop() boot + WiFi / one conversation turn per press
*
* Robojax.com
* ===========================================================================
*/
#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"
/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE 16000
#define RECORD_MAX_S 6 // hard cap on one question
#define WAV_HEADER_LEN 44
#define REC_BUF_BYTES (SAMPLE_RATE * RECORD_MAX_S * 2) // 16-bit mono
/* Software gain applied to the recording. The INMP441 capture is quiet at
* 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
* speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
* lower it. */
#define MIC_GAIN 6
/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
* 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
* back silent or garbled in stereo mode. */
#define USE_BOTH_MICS 1
/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
* uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS 25
/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
* shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);
BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);
/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr; // WAV_HEADER_LEN + up to REC_BUF_BYTES
/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;
/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool ok_sd = false;
float g_mic_peak_pct = 0; // last recording's raw peak, % of full scale
char g_stt_err[64] = ""; // last STT failure cause, verbatim
char g_llm_err[64] = ""; // last DeepSeek failure cause, verbatim
/* --- layout --------------------------------------------------------------- */
#define CHAT_H 200 // chat area: y 0..199
#define BAR_Y 202 // button bar below it
#define BTN_SPEAK_X 4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W 58
#define WIFI_X 236 // signal indicator, right end of the bar
#define BTN_H 36
/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
* This is what prevents new text printing over old - the whole area is
* rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int chat_count = 0;
/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
rgb.setPixelColor(0, rgb.Color(r, g, b));
rgb.show();
}
#define LED_IDLE() led(0, 0, 0)
#define LED_LISTEN() led(0, 60, 255) // blue - recording
#define LED_THINK() led(255, 120, 0) // amber - waiting on the cloud
#define LED_SPEAK() led(0, 255, 40) // green - talking
#define LED_ERROR() led(255, 0, 0) // red
/* ===========================================================================
* Touch
* =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
TOUCHINFO ti;
if (!bbct.getSamples(&ti)) return false;
if (ti.count < 1) return false;
*x = ti.y[0];
*y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
return true;
}
bool speakButtonHeld() {
uint16_t x, y;
if (!getTouch(&x, &y)) return false;
return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}
/* ===========================================================================
* Chat UI — word-wrapped bubbles, user right/blue, assistant left/grey
* =========================================================================== */
#define CHAT_CHARS 42 // chars per line at textsize 1
static int bubbleLines(const char *t) {
int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
return l < 1 ? 1 : l;
}
void drawOneBubble(const ChatMsg &m, int y) {
int len = strlen(m.text);
int lines = bubbleLines(m.text);
int h = lines * 10 + 8;
uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
if (w < 30) w = 30;
int x = m.from_user ? (316 - w) : 4;
gfx->fillRoundRect(x, y, w, h, 5, bg);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
for (int i = 0; i < lines; i++) {
char line[CHAT_CHARS + 1] = {0};
strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
gfx->setCursor(x + 5, y + 5 + i * 10);
gfx->print(line);
}
}
/* Rebuild the whole chat area from history: newest message anchored at the
* bottom, older ones stacked upward until the area is full. */
void redrawChat() {
gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
int shown = min(chat_count, CHAT_HISTORY);
int y = CHAT_H - 2;
for (int i = 0; i < shown; i++) {
ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
int h = bubbleLines(m.text) * 10 + 8;
y -= h;
if (y < 0) break; // area full - older ones drop off
drawOneBubble(m, y);
y -= 4;
}
}
void clearChat() {
chat_count = 0;
redrawChat();
}
void chatBubble(const char *text, bool from_user) {
ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
strncpy(m.text, text, sizeof(m.text) - 1);
m.text[sizeof(m.text) - 1] = 0;
m.from_user = from_user;
chat_count++;
redrawChat();
}
/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
bool up = (WiFi.status() == WL_CONNECTED);
long rssi = up ? WiFi.RSSI() : -100;
// -55 dBm or better = full bars; each 10 dB drops one
int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;
for (int b = 0; b < 4; b++) {
int bh = 6 + b * 6; // heights 6,12,18,24
uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
}
gfx->setTextSize(1);
gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
if (up) {
gfx->setTextColor(CYAN);
gfx->printf("%lddBm", rssi);
} else {
gfx->setTextColor(RED);
gfx->print("DOWN");
}
gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
gfx->setTextColor(gfx->color565(120, 120, 120));
gfx->print("WiFi");
}
void drawBar(const char *label, uint16_t colour) {
gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
gfx->setTextSize(2);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
gfx->print(label);
// CLEAR wipes the chat history
gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
gfx->print("CLEAR");
drawWifi();
}
/* ===========================================================================
* I2S — microphones on port 0, speaker on port 1. Separate hardware
* ports, so recording and playback can never fight over a bus.
* =========================================================================== */
void micInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
/* BOTH channels - this is the two-microphone fix. The vendor examples
* use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
* falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
.channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = false,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_MIC_SCK,
.ws_io_num = I2S_MIC_WS,
.data_out_num = I2S_PIN_NO_CHANGE,
.data_in_num = I2S_MIC_SD
};
i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_MIC_PORT, &pins);
}
void spkInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT, // mono - the amp downmixes anyway
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = true,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_SPK_BCLK,
.ws_io_num = I2S_SPK_LRC,
.data_out_num = I2S_SPK_DOUT,
.data_in_num = I2S_PIN_NO_CHANGE
};
i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_SPK_PORT, &pins);
i2s_zero_dma_buffer(I2S_SPK_PORT);
}
/* ===========================================================================
* Recording — runs while the SPEAK button is held (up to RECORD_MAX_S).
* Reads stereo pairs, averages L+R into one mono stream, applies a little
* software gain, and fills wav_buf after the 44-byte header slot.
* Returns the number of audio bytes recorded.
* =========================================================================== */
size_t recordWhileHeld() {
int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
size_t mono_samples = 0;
const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;
int16_t chunk[512];
uint32_t last_touch_ok = millis();
int32_t peak = 0; // loudest raw sample - mic health check
i2s_zero_dma_buffer(I2S_MIC_PORT);
while (mono_samples < max_samples) {
size_t got = 0;
i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);
#if USE_BOTH_MICS
size_t n = got / 4; // 4 bytes = one L+R pair
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
size_t n = got / 2; // 2 bytes = one mono sample
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = chunk[i];
#endif
if (abs(raw) > peak) peak = abs(raw);
int32_t mixed = raw * MIC_GAIN;
if (mixed > 32767) mixed = 32767;
if (mixed < -32768) mixed = -32768;
mono[mono_samples++] = (int16_t)mixed;
}
/* The GT911 is polled between I2S reads. A 250 ms grace period stops a
* momentary missed touch sample from cutting the recording short. */
if (speakButtonHeld()) last_touch_ok = millis();
else if (millis() - last_touch_ok > 250) break;
// live progress on the button
static uint32_t last_draw = 0;
if (millis() - last_draw > 200) {
last_draw = millis();
gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
(int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
}
}
/* Mic health line: peak as % of full scale BEFORE gain.
* 0% = the mic is not being read at all (config/pin problem)
* under 3% = too quiet - speak closer or raise MIC_GAIN
* 3-40% = healthy speech level
*/
g_mic_peak_pct = peak * 100.0 / 32768.0;
Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
g_mic_peak_pct, MIC_GAIN,
peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");
return mono_samples * 2;
}
/* Dump the exact WAV we are about to POST onto the SD card, so it can be
* played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
if (!ok_sd) return;
SD.remove("/stt_debug.wav");
File f = SD.open("/stt_debug.wav", FILE_WRITE);
if (!f) return;
f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
f.close();
Serial.println("debug copy saved to SD as /stt_debug.wav");
}
/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
uint32_t file_len = data_bytes + 36;
uint32_t byte_rate = SAMPLE_RATE * 2;
memcpy(h, "RIFF", 4); memcpy(h + 4, &file_len, 4);
memcpy(h + 8, "WAVEfmt ", 8);
uint32_t fmt_len = 16; memcpy(h + 16, &fmt_len, 4);
uint16_t fmt = 1, ch = 1; memcpy(h + 20, &fmt, 2); memcpy(h + 22, &ch, 2);
uint32_t rate = SAMPLE_RATE; memcpy(h + 24, &rate, 4); memcpy(h + 28, &byte_rate, 4);
uint16_t align = 2, bits = 16; memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
memcpy(h + 36, "data", 4); memcpy(h + 40, &data_bytes, 4);
}
/* ===========================================================================
* HTTP response reader — shared by all three cloud calls.
* Returns the status code and fills body_out. Handles chunked transfer
* encoding PROPERLY: the chunk-size markers must be stripped, or they end
* up embedded inside the JSON body and the parse fails on long replies.
* =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
* on timeout / closed-and-empty. Every read below goes through this, because
* a reasoning model can think for many seconds between the response headers
* and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
uint32_t t0 = millis();
while (!c.available()) {
if (!c.connected()) return false;
if (millis() - t0 > ms) return false;
delay(10);
}
return true;
}
static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
body_out = "";
if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
while (waitData(client, idle_ms)) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break; // blank line = end of headers
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
}
if (chunked) {
int blanks = 0;
while (true) {
/* The chunk-size line may not arrive for a long time while the model
* reasons. Waiting here - instead of letting read() time out - is the
* whole fix: a timed-out read looks exactly like "0" (final chunk),
* which silently truncated the body to nothing. */
if (!waitData(client, idle_ms)) {
Serial.println("HTTP: timed out waiting for the next chunk");
break;
}
String szline = client.readStringUntil('\n');
szline.trim();
if (szline.length() == 0) { // stray blank line
if (++blanks > 4) break;
continue;
}
blanks = 0;
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break; // genuine final chunk
long got = 0;
while (got < sz) {
if (!waitData(client, idle_ms)) break;
while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
}
if (waitData(client, 3000)) client.readStringUntil('\n'); // CRLF after chunk
if (got < sz) { Serial.println("HTTP: short chunk"); break; }
}
} else {
while (waitData(client, idle_ms))
while (client.available()) body_out += (char)client.read();
}
return code;
}
/* ===========================================================================
* CLOUD CALL 1 — Azure speech-to-text
* One POST, one header, plain WAV body. This is why Azure does the ears.
* =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
/* HTTPClient's one-shot POST fails on bodies this large (it attempts one
* giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
* function speaks HTTP directly and streams the WAV up in 4 KB chunks -
* reliable, and if it ever stalls we know the exact byte it stopped at. */
WiFiClientSecure client;
client.setInsecure(); // no cert bundle on-device; see notes
client.setTimeout(15); // seconds, for reads
writeWavHeader(wav_buf, audio_bytes);
dumpWavToSD(audio_bytes); // PC-playable copy of what we send
size_t total = WAV_HEADER_LEN + audio_bytes;
if (!client.connect(AZURE_STT_HOST, 443)) {
snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
Serial.println("STT: TLS connect failed");
return false;
}
/* Two valid host forms use DIFFERENT URL paths - detect which one is in
* secrets.h: <resource>.cognitiveservices.azure.com -> /stt/speech/...
* <region>.stt.speech.microsoft.com -> /speech/... */
bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
"/speech/recognition/conversation/cognitiveservices/v1"
"?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
"Host: " AZURE_STT_HOST "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
"Accept: application/json\r\n"
"Connection: close\r\n"
"Content-Length: " + String(total) + "\r\n\r\n";
client.print(req);
/* body, 4 KB at a time */
size_t sent = 0;
while (sent < total) {
size_t n = min((size_t)4096, total - sent);
size_t w = client.write(wav_buf + sent, n);
if (w == 0) {
delay(50); // brief stall - retry once
w = client.write(wav_buf + sent, n);
if (w == 0) {
snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
(unsigned)(sent / 1024));
Serial.printf("STT: upload stalled at %u/%u bytes\n",
(unsigned)sent, (unsigned)total);
client.stop();
return false;
}
}
sent += w;
yield();
}
Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);
/* read the reply with proper de-chunking */
String resp;
int code = readHttpResponse(client, resp, 10000);
client.stop();
if (code != 200) {
snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
return false;
}
const char *status = doc["RecognitionStatus"];
if (!status || strcmp(status, "Success") != 0) {
/* The status names the exact failure:
* InitialSilenceTimeout = Azure heard silence (mic level too low)
* NoMatch = heard sound but no recognisable words
* BabbleTimeout = heard only noise */
snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
return false;
}
text_out = doc["DisplayText"].as<String>();
if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
return text_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 2 — DeepSeek chat completion
* OpenAI-compatible format. Model name is deepseek-v4-flash - the old
* deepseek-chat name is dead, see secrets.h.
* =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
/* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
* bodies (long answers + the model's hidden reasoning), which shows up as
* "bad JSON reply". Reading until the server closes the connection is
* reliable regardless of reply length. */
JsonDocument req;
req["model"] = DEEPSEEK_MODEL;
req["max_tokens"] = LLM_MAX_TOKENS;
JsonArray msgs = req["messages"].to<JsonArray>();
JsonObject sys = msgs.add<JsonObject>();
sys["role"] = "system"; sys["content"] = SYSTEM_PROMPT;
JsonObject usr = msgs.add<JsonObject>();
usr["role"] = "user"; usr["content"] = question;
String body;
serializeJson(req, body);
WiFiClientSecure client;
client.setInsecure();
/* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
* ~2 s, but anything that needs actual working-out (an Ohm's law problem,
* say) can think for 10-30 s before sending a single byte. */
client.setTimeout(60);
if (!client.connect(DEEPSEEK_HOST, 443)) {
snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
Serial.println("LLM: TLS connect failed");
return false;
}
client.print(String("POST /chat/completions HTTP/1.1\r\n"
"Host: " DEEPSEEK_HOST "\r\n"
"Authorization: Bearer " DEEPSEEK_KEY "\r\n"
"Content-Type: application/json\r\n"
"Connection: close\r\n"
"Content-Length: ") + String(body.length()) + "\r\n\r\n");
client.print(body);
/* read the reply with proper de-chunking; generous window - long
* questions make the model think for a while before it responds */
String resp;
int code = readHttpResponse(client, resp, 60000);
client.stop();
if (code != 200) {
snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
return false;
}
const char *content = doc["choices"][0]["message"]["content"];
const char *finish = doc["choices"][0]["finish_reason"];
if (!content || !content[0]) {
/* v4-flash is a reasoning model: if finish_reason is "length", the whole
* token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
if (finish && strcmp(finish, "length") == 0)
snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
else
snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
return false;
}
answer_out = String(content);
answer_out.trim();
if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
return answer_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 3 — Azure text-to-speech, streamed straight to the speaker
* We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
* the I2S peripheral eats. Skip the 44-byte header, forward the rest.
* No MP3 decoder, no audio library, no buffering the whole reply.
* =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
size_t got = 0;
uint32_t t0 = millis();
while (got < n && millis() - t0 < 10000) {
int r = c.read(dst + got, n - got);
if (r > 0) { got += r; t0 = millis(); }
else if (!c.connected() && !c.available()) break;
else delay(2);
}
return got;
}
bool azureTTSSpeak(const String &text) {
// Escape the XML special characters for the SSML body
String safe = text;
safe.replace("&", "&");
safe.replace("<", "<");
safe.replace(">", ">");
String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
"<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";
/* Manual HTTP like the other two cloud calls - and for a hard reason:
* Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
* raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
* into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
* Here we parse the framing properly and keep only clean audio bytes. */
WiFiClientSecure client;
client.setInsecure();
client.setTimeout(20);
const char *host = AZURE_REGION ".tts.speech.microsoft.com";
if (!client.connect(host, 443)) {
Serial.println("TTS: TLS connect failed");
return false;
}
client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
"Host: ") + host + "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: application/ssml+xml\r\n"
"X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
"User-Agent: MaTouchRobojax\r\n"
"Connection: close\r\n"
"Content-Length: " + String(ssml.length()) + "\r\n\r\n");
client.print(ssml);
/* status + headers; note whether the body is chunked */
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
long content_len = -1;
while (client.connected() || client.available()) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break;
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
}
if (code != 200) {
Serial.printf("TTS HTTP %d\n", code);
client.stop();
return false;
}
const size_t AUDIO_CAP = 1200 * 1024; // ~37 s of speech
uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
if (!audio) { client.stop(); return false; }
size_t alen = 0;
if (chunked) {
/* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
while (true) {
String szline = client.readStringUntil('\n');
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break;
if (alen + sz > AUDIO_CAP) break;
size_t got = readExact(client, audio + alen, sz);
alen += got;
client.readStringUntil('\n'); // trailing CRLF after each chunk
if (got < (size_t)sz) break;
}
} else if (content_len > 0) {
alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
} else {
/* no framing info: read until the server closes */
uint32_t idle = millis();
while ((client.connected() || client.available()) && millis() - idle < 5000) {
int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
if (r > 0) { alen += r; idle = millis(); }
else delay(5);
}
}
client.stop();
Serial.printf("TTS: %u KB clean audio (%s), playing\n",
(unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");
bool ok = (alen > WAV_HEADER_LEN);
if (ok) {
/* NOW the audio actually starts - this is the honest moment to go green */
LED_SPEAK();
drawBar("SPEAKING...", gfx->color565(0, 130, 40));
/* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
static const uint8_t lead_in[640] = {0}; // 20 ms of silence
size_t w = 0;
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
}
free(audio);
// let the DMA buffers drain so the last word is not cut off
delay(150);
i2s_zero_dma_buffer(I2S_SPK_PORT);
return ok;
}
/* ===========================================================================
* SETUP
* =========================================================================== */
void setup() {
Serial.begin(115200);
delay(400);
Serial.println("\n=== 04 Voice Assistant | Robojax.com ===");
Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");
pinMode(TFT_BLK, OUTPUT);
digitalWrite(TFT_BLK, LOW);
pinMode(SD_CS, OUTPUT);
digitalWrite(SD_CS, HIGH);
// one shared SPI bus for TFT + SD (started before either device)
SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);
gfx->begin();
gfx->fillScreen(BLACK);
digitalWrite(TFT_BLK, HIGH);
// SD is optional here - it only stores the /stt_debug.wav diagnostic copy
ok_sd = SD.begin(SD_CS, SPI, 20000000);
digitalWrite(SD_CS, HIGH);
Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
: "SD not found - debug WAV dump disabled (not fatal)");
bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
delay(50);
rgb.begin();
rgb.setBrightness(LED_BRIGHTNESS);
LED_IDLE();
/* One recording buffer for the whole session, in PSRAM. This is the 8 MB
* that makes the board worth buying. */
wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
if (!wav_buf) {
gfx->setTextColor(RED);
gfx->setTextSize(2);
gfx->setCursor(10, 100);
gfx->print("PSRAM alloc failed!");
gfx->setTextSize(1);
gfx->setCursor(10, 130);
gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
while (1) delay(1000);
}
micInit();
spkInit();
gfx->setTextSize(1);
gfx->setTextColor(YELLOW);
gfx->setCursor(4, 4);
gfx->printf("Connecting to %s ...", WIFI_SSID);
Serial.printf("Connecting to %s ", WIFI_SSID);
WiFi.mode(WIFI_STA);
WiFi.begin(WIFI_SSID, WIFI_PASS);
uint32_t t0 = millis();
while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
delay(300);
Serial.print(".");
}
Serial.println();
clearChat();
if (WiFi.status() == WL_CONNECTED) {
Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
chatBubble("Hold SPEAK and ask me anything.", false);
} else {
chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
LED_ERROR();
}
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}
/* ===========================================================================
* LOOP — one full conversation turn per button press
* =========================================================================== */
void loop() {
/* CLEAR button: edge-detected so one tap wipes once. Reading the panel
* twice per loop (here and in speakButtonHeld) is fine - the GT911 just
* reports its current state. */
static bool tap_latch = false;
static uint8_t tap_release = 0;
if (state == ST_IDLE) {
uint16_t cx, cy;
if (getTouch(&cx, &cy)) {
tap_release = 0;
if (!tap_latch) {
tap_latch = true;
if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
clearChat();
chatBubble("Hold SPEAK and ask me anything.", false);
}
}
} else if (tap_latch && ++tap_release >= 4) {
tap_latch = false;
tap_release = 0;
}
/* live WiFi signal indicator, refreshed every 2 s while idle */
static uint32_t last_wifi = 0;
if (millis() - last_wifi > 2000) {
last_wifi = millis();
drawWifi();
}
}
if (state == ST_IDLE && speakButtonHeld()) {
/* ---- record ---- */
state = ST_RECORDING;
LED_LISTEN();
drawBar("LISTENING...", gfx->color565(0, 60, 200));
uint32_t t_rec = millis();
size_t audio_bytes = recordWhileHeld();
t_rec = millis() - t_rec;
Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);
if (audio_bytes < SAMPLE_RATE / 2) { // under a quarter second - a tap
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
/* ---- speech to text ---- */
state = ST_STT;
LED_THINK();
drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
uint32_t t_stt = millis();
String question;
if (!azureSTT(audio_bytes, question)) {
/* Show the REAL cause on screen - no serial monitor needed. */
char diag[96];
snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
g_stt_err, g_mic_peak_pct,
ok_sd ? " | saved /stt_debug.wav" : "");
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_stt = millis() - t_stt;
chatBubble(question.c_str(), true);
Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());
/* ---- think ---- */
state = ST_LLM;
drawBar("THINKING...", gfx->color565(150, 90, 0));
uint32_t t_llm = millis();
String answer;
if (!deepseekChat(question, answer)) {
char diag[96];
snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_llm = millis() - t_llm;
chatBubble(answer.c_str(), false);
Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());
/* ---- speak ----
* Still amber here: the voice has to be synthesised and downloaded first
* (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
* at the exact moment audio starts coming out of the speaker. */
state = ST_TTS;
LED_THINK();
drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
uint32_t t_tts = millis();
bool spoke = azureTTSSpeak(answer);
t_tts = millis() - t_tts;
/* Timing summary on serial - this feeds the "honest numbers" segment. */
Serial.printf("TIMINGS rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
audio_bytes / 32000.0, (unsigned long)t_stt,
(unsigned long)t_llm, (unsigned long)t_tts,
spoke ? "" : " (TTS FAILED)");
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
}
delay(20);
}
สิ่งที่คุณอาจต้องการ
-
อื่น ๆProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
แหล่งข้อมูลและเอกสารอ้างอิง
-
เอกสารMakerfabs MaTouch ESP32-S3 2.8" Camera and Touchscreen: user's manualwiki.makerfabs.com
-
เอกสาร
-
เอกสารProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
-
ดาวน์โหลดArduino GFX Library on Githubgithub.com
ไฟล์📁
ไฟล์ที่ต้องการ (.h)
ไฟล์อื่น ๆ
แผนภาพวงจร
-
MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematicThe latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF0.15 MB