Ten poradnik jest częścią: Makerfabs MaTouch AI ESP32S3 2.8" Camera
Najnowsza płyta MaTouch AI integruje wejście audio I2S/głośnik I2S/ kamerę 3 miliony pikseli OV3660/ wyświetlacz o rozdzielczości 320*240, z mocnym procesorem ESP32S3 i możliwością WiFi, co czyni tę płytę dobrym narzędziem/platformą do rozwoju AI z ESP32.
Makerfabs MaTouch ESP32-S3 2.8" kamera Zbuduj asystenta głosowego AI na ESP32-S3 (Azure + DeepSeek)
Przytrzymaj przycisk, zadaj pytanie, a płytka odpowie na głos
Kompletny asystent głosowy na jednej małej płytce. Przytrzymaj przycisk SPEAK i zadaj pytanie. Płytka nagrywa Cię obydwoma mikrofonami, wysyła dźwięk do Microsoft Azure, gdzie jest zamieniany na tekst, wysyła ten tekst do DeepSeek do przemyślenia, odsyła odpowiedź do Azure w celu zamiany na mowę i odtwarza ją przez własny głośnik. Cała rozmowa pojawia się na ekranie w postaci dymków czatu.
Mówiący asystent głosowy AI działający na płycie MaTouch AI ESP32-S3
Co się dzieje od momentu naciśnięcia przycisku
Oto cała podróż jednego pytania, krok po kroku. Warto przeczytać to raz, ponieważ wszystko, co widzisz na ekranie i na diodzie LED, odpowiada jednemu z tych etapów.
Naciskasz i PRZYTRZYMUJESZ przycisk SPEAK. To tryb mówienia przez przytrzymanie, a nie dotknięcie: nagrywanie trwa dokładnie tak długo, jak długo trzymasz palec, maksymalnie do sześciu sekund. Dioda statusu zmienia kolor na niebieski, a przycisk pokazuje LISTENING.
Oba mikrofony Cię nagrywają. Płytka próbkuje 16 000 razy na sekundę z pary stereo, uśrednia oba kanały do jednego, dodaje odrobinę wzmocnienia i zapisuje wynik w PSRAM. Pasek postępu przesuwa się po przycisku, gdy mówisz. Dwie sekundy mowy to około 64 KB.
Zwalniasz przycisk. Nagrywanie się zatrzymuje. Płytka zapisuje 44-bajtowy nagłówek WAV na początku dźwięku – ta drobna etykieta to wszystko, co zamienia surowe próbki w plik, który Azure zaakceptuje.
Dźwięk trafia do Azure Speech-to-Text. Jest przesyłany w kawałkach po 4 KB przez bezpieczne połączenie i wraca jako pojedyncza linia tekstu. Twoje 64 KB dźwięku zamienia się w około 25 bajtów tekstu. Dioda LED zmienia kolor na bursztynowy.
Twoje pytanie pojawia się na ekranie jako niebieski dymek czatu, więc możesz zobaczyć dokładnie to, co usłyszało – co jest przydatne, ponieważ błędnie usłyszane słowa wyjaśniają większość dziwnych odpowiedzi.
Tekst trafia do DeepSeek. Płytka wysyła Twoje pytanie wraz ze stałą instrukcją, aby odpowiedzi były ograniczone do dwóch krótkich zdań. Model myśli – naprawdę myśli, to model rozumujący – i zwraca odpowiedź.
Odpowiedź pojawia się na ekranie jako szary dymek. Możesz ją przeczytać, zanim ją usłyszysz.
Odpowiedź wraca do Azure, aby zostać wypowiedziana. Płytka żąda surowego dźwięku PCM 16 kHz, który jest dokładnie formatem oczekiwanym przez jej wzmacniacz, więc w tym projekcie nie ma żadnego dekodera MP3. Przycisk pokazuje teraz GETTING VOICE, a dioda LED pozostaje bursztynowa, ponieważ nic jeszcze nie jest słyszalne.
Cały klip jest pobierany do PSRAM, zanim zostanie odtworzona choćby jedna próbka. To ważne – zobacz notkę poniżej.
Odtwarzanie. W momencie, gdy dźwięk trafia do głośnika, dioda LED zmienia kolor na zielony, a przycisk pokazuje SPEAKING. Słyszysz odpowiedź.
Dlaczego dźwięk jest najpierw pobierany, a nie odtwarzany w miarę docierania. Strumieniowanie bezpośrednio z sieci do głośnika brzmi jak stukanie. Bufor głośnika mieści tylko około jednej dziesiątej sekundy, a każda przerwa w transferze WiFi dłuższa niż ta opróżnia go, powodując słyszalne stuknięcie. Pobranie całej odpowiedzi do PSRAM kosztuje około sekundy dodatkowego czekania i eliminuje wszystkie przerwy. To również dlatego wyświetlacz pokazuje GETTING VOICE, zanim pokaże SPEAKING – te dwa etapy są uczciwie różne.
Ile czasu zajmuje każdy etap
Zmierzono na prawdziwym sprzęcie, dla prostego pytania:
Etap | Typowy czas |
|---|---|
Nagrywanie | tak długo, jak trzymasz przycisk |
Zamiana mowy na tekst (Azure) | około 1,8 sekundy |
Myślenie (DeepSeek) | około 1,8 sekundy dla prostego pytania, znacznie dłużej dla wymagającego prawdziwego rozumowania |
Pobieranie głosu (Azure) | około 7 sekund – największa pojedyncza część |
Łącznie, od zwolnienia do pierwszego dźwięku | mniej więcej 11 sekund |
Każda wymiana wypisuje własne czasy na monitorze szeregowym, więc możesz zmierzyć własne, zamiast ufać tym. Jeśli chcesz przyspieszyć, najskuteczniejszą zmianą jest proszenie o krótsze odpowiedzi w SYSTEM_PROMPT – mniej tekstu do wypowiedzenia oznacza mniej dźwięku do syntezy i pobrania.
Sama płytka nigdy niczego nie rozumie. Jest posłańcem z dobrymi uszami i dobrym głosem – inteligencja jest wynajmowana za sekundę.
Dlaczego te trzy usługi
Azure obsługuje mowę na wejściu i wyjściu. Jego zamiana tekstu na mowę może zwracać surowe PCM 16 kHz, co jest dokładnie tym, czego chce układ głośnika, więc w tym projekcie nie ma dekodera MP3. Jego zamiana mowy na tekst przyjmuje zwykły WAV w zwykłym żądaniu POST.
DeepSeek to mózg konwersacji. Jest szybki i kosztuje ułamek centa za odpowiedź.
OpenAI nie jest tutaj używane – zobacz projekt 05, gdzie wykonuje pracę związaną z widzeniem.
Zmieniły się nazwy modeli DeepSeek. Stare nazwy deepseek-chat i deepseek-reasoner zostały wycofane w lipcu 2026. Większość poradników online nadal ich używa i zwróci błąd. Obecne nazwy to deepseek-v4-flash i deepseek-v4-pro. Ten projekt używa v4-flash.
Pułapka modelu rozumującego
DeepSeek v4-flash myśli, zanim odpowie, a to myślenie wlicza się do limitu tokenów. Jeśli ustawisz LLM_MAX_TOKENS zbyt nisko, cały budżet zostanie zużyty na rozumowanie, odpowiedź wróci pusta, a tablica nic nie powie. Dlatego tutaj ustawiono go na 400. Trudne pytania również zajmują więcej czasu – prosta odpowiedź na fakt pojawia się w około dwie sekundy, a pytanie wymagające faktycznego rozwiązania może zająć znacznie więcej.
Odczytywanie wskaźnika statusu
Kolor | Znaczenie |
|---|---|
Niebieski | słucha Ciebie |
Bursztynowy | chmura myśli lub głos jest pobierany |
Zielony | mówienie – zmienia się na zielony dokładnie w momencie rozpoczęcia dźwięku |
Czerwony | coś się nie powiodło – sprawdź monitor szeregowy |
Kontrolki na ekranie
Czat przewija się jak rozmowa telefoniczna, a najstarsze wiadomości przesuwają się w górę i znikają. CLEAR go czyści. W rogu znajduje się wskaźnik siły sygnału WiFi z rzeczywistym odczytem w dBm, co jest przydatne, gdy zastanawiasz się, czy powolna odpowiedź to wina sieci, czy usługi.
O płycie MaTouch AI ESP32-S3 2.8"
Każdy projekt na tej stronie działa na płycie MaTouch AI ESP32-S3 2.8" TFT ST7789V firmy Makerfabs. To płyta typu „wszystko w jednym”: kolorowy ekran dotykowy, aparat 3 megapikseli, dwa mikrofony i prawdziwy wzmacniacz głośnikowy, wszystko sterowane przez ESP32-S3 z 8 MB PSRAM. To właśnie ta kombinacja umożliwia realizację tych projektów AI na jednej płycie bez żadnych dodatkowych elementów.
8 MB PSRAM ma większe znaczenie niż jakakolwiek inna liczba tutaj. To dzięki niemu płyta może jednocześnie przechowywać w pamięci klatkę z kamery, kilka sekund nagranego dźwięku lub zdjęcie zakodowane w base64 – żadne z tych rzeczy nie mieści się w normalnej pamięci RAM ESP32.
Dokumentacja producenta: strona wiki Makerfabs.
Kluczowe specyfikacje
Procesor: ESP32-S3, dwurdzeniowy 240 MHz, WiFi 2.4 GHz + Bluetooth 5.0
Pamięć: 16 MB flash, 8 MB PSRAM (wymagane przez prawie każdy projekt tutaj)
Wyświetlacz: 2.8" IPS, 320×240, sterownik ST7789V, SPI
Dotyk: pojemnościowy GT911, śledzi 5 palców jednocześnie
Aparat: OV3660, 3 megapiksele, do 2048×1536
Mikrofony: dwa cyfrowe mikrofony I2S INMP441 (prawdziwa para stereo)
Głośnik: wzmacniacz klasy D MAX98357A, 3.2 W przy 4 Ω
Pamięć masowa: gniazdo kart microSD (tryb SPI)
Zasilanie: USB-C, złącze baterii JST, ładowarka TP4056, wyłącznik zasilania
Dodatkowo na płycie: dioda RGB WS2812B, zegar czasu rzeczywistego z baterią PCF8563T oraz miernik stanu baterii MAX17048, który nie jest wymieniony w oficjalnych specyfikacjach
Dwa porty USB-C nie są takie same. Głośnik płyty współdzieli piny sygnałowe (IO19 i IO20) z natywnym portem USB, ponieważ te piny to sprzętowe linie danych USB ESP32-S3. Zawsze wgrywaj i zasilaj przez port USB-C CH340K (ten obok przycisku RESET) i ustaw USB CDC On Boot na Disabled. Użycie niewłaściwego portu spowoduje problemy z dźwiękiem lub nieudane wgrywanie.
Ustawienia Arduino IDE
Te ustawienia mają znaczenie. Większość problemów zgłaszanych przez użytkowników tej płyty wynika z jednego z nich, a one resetują się po zmianie wersji rdzenia, więc sprawdź je ponownie po każdej zmianie.
Ustawienie | Wartość |
|---|---|
Płyta | ESP32S3 Dev Module |
Wersja rdzenia ESP32 | 2.0.17 |
PSRAM | OPI PSRAM |
Rozmiar flash | 16MB (128Mb) |
Schemat partycji | 16M Flash (3MB APP/9.9MB FATFS) |
USB CDC On Boot | Disabled |
Prędkość wgrywania | 921600 |
Erase All Flash Before Upload | Disabled |
Port | port USB-C CH340K |
Użyj rdzenia ESP32 2.0.17, a nie 3.x. Espressif usunął modele wykrywania twarzy na urządzeniu w rdzeniu 3, więc projekty związane z twarzą nie skompilują się tam. Przypięcie wersji 2.0.17 sprawia, że każdy projekt na tej stronie działa z jedną konfiguracją. W Menedżerze płyt rozwijane menu wersji pozwala przełączać się w dowolnym momencie.
Użyj GFX Library for Arduino w wersji 1.5.6, a nie 1.6.x. Wydania 1.6 są zbudowane dla rdzenia ESP32 3 i mogą zawieszać się przy starcie na rdzeniu 2.0.17. Jeśli ekran pozostaje czarny po wgraniu, to pierwsza rzecz do sprawdzenia.
Wymagane biblioteki
Zainstaluj je przez Tools → Manage Libraries w Arduino IDE. Numery wersji mają znaczenie – użyj podanych.
Biblioteka | Wersja | Autor |
|---|---|---|
GFX Library for Arduino | 1.5.6 | moononournation |
bb_captouch | 1.3.1 | Larry Bank |
ArduinoJson | 7.x | Benoit Blanchon |
Adafruit NeoPixel | dowolna najnowsza | Adafruit |
Konfigurowanie pliku secrets.h
Twoje dane WiFi oraz wszelkie klucze API znajdują się w secrets.h, który jest dołączony do pobranego pliku z przykładowymi wartościami. Otwórz tę kartę w Arduino IDE i zastąp je własnymi.
WiFi musi działać w paśmie 2,4 GHz. ESP32-S3 w ogóle nie widzi sieci 5 GHz. Jeśli Twój router łączy oba pasma pod jedną nazwą (Asus nazywa to Smart Connect), wyłącz tę funkcję lub nadaj pasmu 2,4 GHz osobną nazwę i użyj jej w secrets.h.
Uzyskiwanie kluczy API
Ten projekt komunikuje się z chmurową usługą AI, więc potrzebujesz własnego klucza. Jeśli nigdy tego nie robiłeś, nie martw się – to działa na tej samej zasadzie co hasło identyfikujące Twoje konto w usłudze. Zajmie to kilka minut, tylko raz.
Klucz to nie subskrypcja strony internetowej. Na przykład płacenie za ChatGPT Plus nie daje Ci klucza API – to dwa osobne produkty z osobnymi rozliczeniami. Potrzebujesz konta na platformie deweloperskiej, co opisano poniżej.
Microsoft Azure Speech – do słuchania i mówienia
Azure zamienia Twoją mowę na tekst, a odpowiedź z powrotem na głos. Darmowy poziom jest wystarczający do wszystkiego na tej stronie.
Przejdź do portal.azure.com i zaloguj się na konto Microsoft (darmowe jest w porządku).
Jeśli nigdy nie korzystałeś z Azure, zobaczysz ekran Witamy w Azure z trzema opcjami. Wybierz Start z bezpłatną wersją próbną Azure – potrzebujesz subskrypcji, zanim Azure pozwoli Ci cokolwiek utworzyć. (Studenci powinni zamiast tego wybrać Azure dla studentów: ten sam efekt, bez wymogu podania karty.) Zignoruj Zarządzaj Microsoft Entra ID, które dotyczy czegoś zupełnie innego.
Kliknij Utwórz zasób, wyszukaj Mowa i wybierz Usługa mowy opublikowaną przez Microsoft.
Wypełnij formularz: dowolna grupa zasobów, dowolna nazwa i wybierz Region blisko Ciebie – zapisz ten region dokładnie tak, jak się wyświetla, na przykład
eastus.W polu Warstwa cenowa wybierz F0 (bezpłatna). Pozwala to na około pięć godzin zamiany mowy na tekst i pół miliona znaków zamiany tekstu na mowę miesięcznie.
Kliknij Przegląd + utwórz, a następnie Utwórz. Poczekaj około minutę, potem kliknij Przejdź do zasobu.
W menu po lewej stronie otwórz Klucze i punkt końcowy. Skopiuj KLUCZ 1 oraz Lokalizację/Region.
Umieść je w secrets.h jako AZURE_SPEECH_KEY i AZURE_REGION. Dla AZURE_STT_HOST użyj <region>.stt.speech.microsoft.com – więc dla regionu eastus będzie to eastus.stt.speech.microsoft.com.
O karcie kredytowej. Bezpłatna wersja próbna Azure prosi o kartę w celu weryfikacji tożsamości. Nie pobiera opłat. Otrzymujesz 200 USD kredytu na 30 dni, a po tym czasie konto przechodzi na model Pay-As-You-Go – ale warstwa F0 Speech pozostaje bezpłatna miesiąc po miesiącu, a wszystkie projekty na tej stronie mieszczą się w niej bez problemu. Jeśli wolisz w ogóle nie podawać karty i jesteś studentem, opcja Azure dla studentów daje Ci kredyt bez niej.
To musi być zasób „Usługa mowy”. Klucz z zasobu Translator, Language lub ogólnego Cognitive Services wygląda identycznie i jest w pełni poprawny – ale każde żądanie mowy zwraca błąd 401. To nas złapało podczas testów i kosztowało godzinę. Jeśli mowa zawodzi z błędem 401, a klucz wygląda dobrze, sprawdź, jaki rodzaj zasobu utworzyłeś.
DeepSeek – część myśląca
DeepSeek to model językowy, który faktycznie odpowiada na Twoje pytanie. Jest niedrogi – kilka dolarów kredytu wystarcza na tysiące odpowiedzi.
Przejdź do platform.deepseek.com i utwórz konto.
Otwórz Klucze API w menu i kliknij Utwórz nowy klucz API.
Skopiuj go natychmiast. Jest wyświetlany tylko raz i nigdy więcej – jeśli go zgubisz, usuń ten klucz i utwórz nowy.
Dodaj niewielką kwotę kredytu w sekcji Doładuj. Nie ma darmowego poziomu, ale najmniejsze doładowanie wystarcza na bardzo długo przy takim użyciu.
Umieść klucz w secrets.h jako DEEPSEEK_KEY. Zaczyna się od sk-.
Nazwy modeli zmieniły się w lipcu 2026. Stare deepseek-chat i deepseek-reasoner zostały wycofane, więc większość poradników online zakończy się błędem 400. Użyj deepseek-v4-flash, który jest już ustawiony w tych projektach.
Ile to kosztuje w działaniu
Bardzo niewiele, ale nie jest darmowe i warto wiedzieć z grubsza, ile wydajesz, zanim zostawisz projekt działający.
Usługa | Przybliżony koszt |
|---|---|
Azure Speech | darmowy poziom obejmuje około 5 godzin słuchania i 0,5 M znaków mówienia miesięcznie |
DeepSeek | ułamek centa za odpowiedź – tysiące odpowiedzi za kilka dolarów |
OpenAI vision | mniej więcej cent lub dwa za zdjęcie, w zależności od modelu |
Ceny się zmieniają, więc traktuj to jako wskazówkę, a nie wycenę. Każda z tych usług ma stronę z użyciem, na której możesz śledzić wydatki, a wszystkie pozwalają ustawić limit wydatków – warto to zrobić pierwszego dnia.
Trzymaj swoje klucze prywatne. Każdy, kto je posiada, może wydać Twoje pieniądze. Nie umieszczaj ich w filmie, zrzucie ekranu, poście na forum ani w publicznym repozytorium kodu. Jeśli klucz kiedykolwiek zostanie ujawniony, usuń go na stronie dostawcy i utwórz nowy - to zajmuje sekundy i jest jedynym prawdziwym rozwiązaniem.
Rozwiązywanie problemów
Objaw | Przyczyna i rozwiązanie |
|---|---|
Ekran pozostaje czarny | Nieprawidłowa wersja biblioteki GFX (użyj 1.5.6) lub nieprawidłowe ustawienia płyty. |
|
|
Nic się nie wgrywa / brak portu COM | Nieprawidłowy port USB-C lub sterownik CH340 nie jest zainstalowany. |
Kamera zawodzi i nigdy nie wraca do działania | Linia resetu kamery jest podłączona do przycisku RESET na płycie, więc oprogramowanie nie może jej zrestartować. Naciśnij RESET. Jeśli nadal nie działa, ponownie podłącz taśmę kamery. |
Pobierz kod
Kompletny szkic Arduino dla tego projektu, wraz z pins.h i wszystkim innym, czego potrzebuje, jest dostępny do bezpłatnego pobrania.
Rozpakuj go, otwórz plik .ino w Arduino IDE, sprawdź powyższe ustawienia i wgraj przez port USB-C CH340K.
Ten poradnik jest częścią: Makerfabs MaTouch AI ESP32S3 2.8" Camera
/*
* ===========================================================================
* 04_Voice_Assistant — MaTouch AI ESP32-S3 2.8" TFT ST7789V
* ===========================================================================
*
----------
* ROBOJAX.COM - MaTouch AI ESP32-S3 2.8" project series
*
* WATCH THE VIDEO
* https://youtu.be/6AL3g3tC_Hk
*
* WRITTEN TUTORIALS - every project, with photos and full explanation
* Camera and touchscreen.... https://robojax.com/RTJ849
* Offline face recognition.. https://robojax.com/RTJ850
* AI voice assistant........ https://robojax.com/RTJ851
* AI vision................. https://robojax.com/RTJ852
*
* GET THE BOARD - SAVE $5 with coupon code: Robojax_Makerfab
* https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
* (enter the code at checkout)
*
* All of this code is free. If it helped you, a subscribe on YouTube is
* the best way to support more of it.
*
* ---------------------------------------------------------------------------
*
* A complete voice assistant on a $40 board:
*
* hold SPEAK -> both INMP441 microphones record you
* -> Azure Speech turns the audio into text
* -> DeepSeek v4-flash thinks of an answer
* -> Azure Speech turns the answer into audio
* -> the MAX98357 speaker says it out loud
*
* and the whole conversation is drawn as chat bubbles on the touchscreen.
*
* WHY THIS COMBINATION OF SERVICES (each is used where it is best):
* - Azure STT accepts a plain WAV in a plain POST with one header. OpenAI's
* transcription endpoint wants multipart/form-data - miserable on an MCU.
* - Azure TTS can return RAW 16 kHz PCM ("riff-16khz-16bit-mono-pcm"),
* which streams straight into the I2S speaker with NO MP3 decoder at all.
* - DeepSeek v4-flash is fast and nearly free per reply. NOTE: the old
* model names deepseek-chat / deepseek-reasoner were RETIRED in July 2026.
* Most tutorials online still use them and are broken. See secrets.h.
*
* A detail the vendor examples get wrong: this board has TWO microphones on
* one I2S bus (left + right), but every Makerfabs demo records left-only and
* throws one away. This sketch records both and averages them.
*
* ---------------------------------------------------------------------------
* *** WHICH USB PORT - THIS MATTERS ***
* The speaker shares IO19/IO20 with the NATIVE USB port. Upload and power
* through the CH340K UART USB-C port, and set USB CDC On Boot = Disabled.
* If you use the wrong port the audio will be garbage or uploads will fail.
* ---------------------------------------------------------------------------
*
* FILL IN secrets.h BEFORE FLASHING (WiFi + all three API keys).
*
* BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
* or compile errors. These reset when you switch cores - recheck them!)
*
* Board : ESP32S3 Dev Module
* ESP32 core : 2.0.17
* PSRAM : OPI PSRAM <-- required, audio buffer lives there
* Flash Size : 16MB (128Mb)
* Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
* USB CDC On Boot : Disabled <-- required, see USB note above
* Upload Speed : 921600
* Port : the CH340K USB-C port (the one near RESET)
*
* LIBRARIES
* GFX Library for Arduino v1.5.6 (NOT 1.6.x - that pairs with core 3)
* bb_captouch v1.3.1
* ArduinoJson v7.x
* Adafruit NeoPixel any recent
*
* ---------------------------------------------------------------------------
* FUNCTIONS IN THIS SKETCH
* led(r,g,b) + LED_* macros RGB status colours (blue/amber/green/red)
* getTouch(&x,&y) read the touch panel, mapped to screen coordinates
* speakButtonHeld() true while a finger is on the SPEAK button
* bubbleLines(t) how many lines a message wraps to
* drawOneBubble(m,y) draw a single chat bubble
* redrawChat() rebuild the chat area from history, newest at bottom
* clearChat() wipe the chat history (CLEAR button)
* chatBubble(t,user) add a message to history and redraw
* drawWifi() WiFi signal bars + dBm readout
* drawBar(label,col) bottom bar: SPEAK button + CLEAR + WiFi meter
* micInit() I2S input - BOTH INMP441 mics, stereo
* spkInit() I2S output - MAX98357 speaker
* recordWhileHeld() record while SPEAK held, downmix stereo->mono
* writeWavHeader(...) prepend the 44-byte RIFF/WAVE header
* dumpWavToSD(...) save the exact upload to SD (/stt_debug.wav)
* readHttpResponse() read an HTTPS reply, de-chunking it properly
* azureSTT(...) chunked upload of the WAV -> recognised text
* deepseekChat(...) question -> deepseek-v4-flash -> answer text
* azureTTSSpeak(text) answer -> Azure voice -> PSRAM -> speaker
* setup() / loop() boot + WiFi / one conversation turn per press
*
* Robojax.com
* ===========================================================================
*/
#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include <SPI.h>
#include <SD.h>
#include "driver/i2s.h"
#include "pins.h"
#include "secrets.h"
/* --- audio geometry ------------------------------------------------------- */
#define SAMPLE_RATE 16000
#define RECORD_MAX_S 6 // hard cap on one question
#define WAV_HEADER_LEN 44
#define REC_BUF_BYTES (SAMPLE_RATE * RECORD_MAX_S * 2) // 16-bit mono
/* Software gain applied to the recording. The INMP441 capture is quiet at
* 16-bit depth; if the serial monitor reports "mic peak" under ~10% while you
* speak normally, raise this (6 -> 10 -> 16). If it reports clipping (100%),
* lower it. */
#define MIC_GAIN 6
/* 1 = read BOTH microphones (stereo bus) and average them - better SNR.
* 0 = vendor-style single left mic. Use 0 as a fallback if recordings come
* back silent or garbled in stereo mode. */
#define USE_BOTH_MICS 1
/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
* uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS 25
/* HWSPI (not ESP32SPI): the debug WAV dump writes to the SD card, which
* shares these pins - both must go through the same SPI driver. */
Arduino_HWSPI *bus = new Arduino_HWSPI(
TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, &SPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);
BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);
/* --- big buffers live in PSRAM, allocated once at boot -------------------- */
uint8_t *wav_buf = nullptr; // WAV_HEADER_LEN + up to REC_BUF_BYTES
/* --- state ---------------------------------------------------------------- */
enum State { ST_IDLE, ST_RECORDING, ST_STT, ST_LLM, ST_TTS, ST_ERROR };
State state = ST_IDLE;
/* --- diagnostics: shown ON SCREEN so the serial monitor is optional -------- */
bool ok_sd = false;
float g_mic_peak_pct = 0; // last recording's raw peak, % of full scale
char g_stt_err[64] = ""; // last STT failure cause, verbatim
char g_llm_err[64] = ""; // last DeepSeek failure cause, verbatim
/* --- layout --------------------------------------------------------------- */
#define CHAT_H 200 // chat area: y 0..199
#define BAR_Y 202 // button bar below it
#define BTN_SPEAK_X 4
#define BTN_SPEAK_W 160
#define BTN_CLEAR_X 170
#define BTN_CLEAR_W 58
#define WIFI_X 236 // signal indicator, right end of the bar
#define BTN_H 36
/* --- chat history: last 8 messages, redrawn newest-at-bottom like a phone.
* This is what prevents new text printing over old - the whole area is
* rebuilt from history on every message, older lines scroll up and out. */
#define CHAT_HISTORY 8
struct ChatMsg { char text[160]; bool from_user; };
ChatMsg chat_hist[CHAT_HISTORY];
int chat_count = 0;
/* --- LED status colours: visible from across the room --------------------- */
void led(uint8_t r, uint8_t g, uint8_t b) {
rgb.setPixelColor(0, rgb.Color(r, g, b));
rgb.show();
}
#define LED_IDLE() led(0, 0, 0)
#define LED_LISTEN() led(0, 60, 255) // blue - recording
#define LED_THINK() led(255, 120, 0) // amber - waiting on the cloud
#define LED_SPEAK() led(0, 255, 40) // green - talking
#define LED_ERROR() led(255, 0, 0) // red
/* ===========================================================================
* Touch
* =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
TOUCHINFO ti;
if (!bbct.getSamples(&ti)) return false;
if (ti.count < 1) return false;
*x = ti.y[0];
*y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
return true;
}
bool speakButtonHeld() {
uint16_t x, y;
if (!getTouch(&x, &y)) return false;
return (x >= BTN_SPEAK_X && x < BTN_SPEAK_X + BTN_SPEAK_W && y >= BAR_Y);
}
/* ===========================================================================
* Chat UI — word-wrapped bubbles, user right/blue, assistant left/grey
* =========================================================================== */
#define CHAT_CHARS 42 // chars per line at textsize 1
static int bubbleLines(const char *t) {
int l = ((int)strlen(t) + CHAT_CHARS - 1) / CHAT_CHARS;
return l < 1 ? 1 : l;
}
void drawOneBubble(const ChatMsg &m, int y) {
int len = strlen(m.text);
int lines = bubbleLines(m.text);
int h = lines * 10 + 8;
uint16_t bg = m.from_user ? gfx->color565(0, 70, 140) : gfx->color565(50, 50, 55);
int w = (len > CHAT_CHARS ? CHAT_CHARS : len) * 6 + 10;
if (w < 30) w = 30;
int x = m.from_user ? (316 - w) : 4;
gfx->fillRoundRect(x, y, w, h, 5, bg);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
for (int i = 0; i < lines; i++) {
char line[CHAT_CHARS + 1] = {0};
strncpy(line, m.text + i * CHAT_CHARS, CHAT_CHARS);
gfx->setCursor(x + 5, y + 5 + i * 10);
gfx->print(line);
}
}
/* Rebuild the whole chat area from history: newest message anchored at the
* bottom, older ones stacked upward until the area is full. */
void redrawChat() {
gfx->fillRect(0, 0, 320, CHAT_H, BLACK);
int shown = min(chat_count, CHAT_HISTORY);
int y = CHAT_H - 2;
for (int i = 0; i < shown; i++) {
ChatMsg &m = chat_hist[(chat_count - 1 - i) % CHAT_HISTORY];
int h = bubbleLines(m.text) * 10 + 8;
y -= h;
if (y < 0) break; // area full - older ones drop off
drawOneBubble(m, y);
y -= 4;
}
}
void clearChat() {
chat_count = 0;
redrawChat();
}
void chatBubble(const char *text, bool from_user) {
ChatMsg &m = chat_hist[chat_count % CHAT_HISTORY];
strncpy(m.text, text, sizeof(m.text) - 1);
m.text[sizeof(m.text) - 1] = 0;
m.from_user = from_user;
chat_count++;
redrawChat();
}
/* WiFi bars + dBm, right end of the button bar. Refreshed from the loop. */
void drawWifi() {
gfx->fillRect(WIFI_X, BAR_Y, 320 - WIFI_X, BTN_H, BLACK);
bool up = (WiFi.status() == WL_CONNECTED);
long rssi = up ? WiFi.RSSI() : -100;
// -55 dBm or better = full bars; each 10 dB drops one
int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;
for (int b = 0; b < 4; b++) {
int bh = 6 + b * 6; // heights 6,12,18,24
uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
gfx->fillRect(WIFI_X + 2 + b * 8, BAR_Y + 28 - bh, 6, bh, col);
}
gfx->setTextSize(1);
gfx->setCursor(WIFI_X + 38, BAR_Y + 6);
if (up) {
gfx->setTextColor(CYAN);
gfx->printf("%lddBm", rssi);
} else {
gfx->setTextColor(RED);
gfx->print("DOWN");
}
gfx->setCursor(WIFI_X + 38, BAR_Y + 18);
gfx->setTextColor(gfx->color565(120, 120, 120));
gfx->print("WiFi");
}
void drawBar(const char *label, uint16_t colour) {
gfx->fillRect(0, BAR_Y, 320, 240 - BAR_Y, BLACK);
gfx->fillRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, colour);
gfx->drawRoundRect(BTN_SPEAK_X, BAR_Y, BTN_SPEAK_W, BTN_H, 6, WHITE);
gfx->setTextSize(2);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_SPEAK_X + 10, BAR_Y + 10);
gfx->print(label);
// CLEAR wipes the chat history
gfx->fillRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, gfx->color565(110, 35, 35));
gfx->drawRoundRect(BTN_CLEAR_X, BAR_Y, BTN_CLEAR_W, BTN_H, 6, WHITE);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_CLEAR_X + 14, BAR_Y + 15);
gfx->print("CLEAR");
drawWifi();
}
/* ===========================================================================
* I2S — microphones on port 0, speaker on port 1. Separate hardware
* ports, so recording and playback can never fight over a bus.
* =========================================================================== */
void micInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
/* BOTH channels - this is the two-microphone fix. The vendor examples
* use ONLY_LEFT here and waste the second microphone. USE_BOTH_MICS 0
* falls back to the vendor-proven single-mic configuration. */
#if USE_BOTH_MICS
.channel_format = I2S_CHANNEL_FMT_RIGHT_LEFT,
#else
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
#endif
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = false,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_MIC_SCK,
.ws_io_num = I2S_MIC_WS,
.data_out_num = I2S_PIN_NO_CHANGE,
.data_in_num = I2S_MIC_SD
};
i2s_driver_install(I2S_MIC_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_MIC_PORT, &pins);
}
void spkInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT, // mono - the amp downmixes anyway
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = true,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_SPK_BCLK,
.ws_io_num = I2S_SPK_LRC,
.data_out_num = I2S_SPK_DOUT,
.data_in_num = I2S_PIN_NO_CHANGE
};
i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_SPK_PORT, &pins);
i2s_zero_dma_buffer(I2S_SPK_PORT);
}
/* ===========================================================================
* Recording — runs while the SPEAK button is held (up to RECORD_MAX_S).
* Reads stereo pairs, averages L+R into one mono stream, applies a little
* software gain, and fills wav_buf after the 44-byte header slot.
* Returns the number of audio bytes recorded.
* =========================================================================== */
size_t recordWhileHeld() {
int16_t *mono = (int16_t *)(wav_buf + WAV_HEADER_LEN);
size_t mono_samples = 0;
const size_t max_samples = SAMPLE_RATE * RECORD_MAX_S;
int16_t chunk[512];
uint32_t last_touch_ok = millis();
int32_t peak = 0; // loudest raw sample - mic health check
i2s_zero_dma_buffer(I2S_MIC_PORT);
while (mono_samples < max_samples) {
size_t got = 0;
i2s_read(I2S_MIC_PORT, chunk, sizeof(chunk), &got, 80 / portTICK_PERIOD_MS);
#if USE_BOTH_MICS
size_t n = got / 4; // 4 bytes = one L+R pair
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = ((int32_t)chunk[i * 2] + (int32_t)chunk[i * 2 + 1]) / 2;
#else
size_t n = got / 2; // 2 bytes = one mono sample
for (size_t i = 0; i < n && mono_samples < max_samples; i++) {
int32_t raw = chunk[i];
#endif
if (abs(raw) > peak) peak = abs(raw);
int32_t mixed = raw * MIC_GAIN;
if (mixed > 32767) mixed = 32767;
if (mixed < -32768) mixed = -32768;
mono[mono_samples++] = (int16_t)mixed;
}
/* The GT911 is polled between I2S reads. A 250 ms grace period stops a
* momentary missed touch sample from cutting the recording short. */
if (speakButtonHeld()) last_touch_ok = millis();
else if (millis() - last_touch_ok > 250) break;
// live progress on the button
static uint32_t last_draw = 0;
if (millis() - last_draw > 200) {
last_draw = millis();
gfx->fillRect(BTN_SPEAK_X + 2, BAR_Y + BTN_H - 6,
(int)((BTN_SPEAK_W - 4) * mono_samples / max_samples), 4, WHITE);
}
}
/* Mic health line: peak as % of full scale BEFORE gain.
* 0% = the mic is not being read at all (config/pin problem)
* under 3% = too quiet - speak closer or raise MIC_GAIN
* 3-40% = healthy speech level
*/
g_mic_peak_pct = peak * 100.0 / 32768.0;
Serial.printf("mic peak: %.1f%% of full scale (gain x%d applied%s)\n",
g_mic_peak_pct, MIC_GAIN,
peak == 0 ? " - MIC IS SILENT, check USE_BOTH_MICS" : "");
return mono_samples * 2;
}
/* Dump the exact WAV we are about to POST onto the SD card, so it can be
* played on a PC - you hear exactly what Azure hears. Overwritten each time. */
void dumpWavToSD(size_t audio_bytes) {
if (!ok_sd) return;
SD.remove("/stt_debug.wav");
File f = SD.open("/stt_debug.wav", FILE_WRITE);
if (!f) return;
f.write(wav_buf, WAV_HEADER_LEN + audio_bytes);
f.close();
Serial.println("debug copy saved to SD as /stt_debug.wav");
}
/* Standard 44-byte RIFF/WAVE header for 16 kHz 16-bit mono PCM. */
void writeWavHeader(uint8_t *h, uint32_t data_bytes) {
uint32_t file_len = data_bytes + 36;
uint32_t byte_rate = SAMPLE_RATE * 2;
memcpy(h, "RIFF", 4); memcpy(h + 4, &file_len, 4);
memcpy(h + 8, "WAVEfmt ", 8);
uint32_t fmt_len = 16; memcpy(h + 16, &fmt_len, 4);
uint16_t fmt = 1, ch = 1; memcpy(h + 20, &fmt, 2); memcpy(h + 22, &ch, 2);
uint32_t rate = SAMPLE_RATE; memcpy(h + 24, &rate, 4); memcpy(h + 28, &byte_rate, 4);
uint16_t align = 2, bits = 16; memcpy(h + 32, &align, 2); memcpy(h + 34, &bits, 2);
memcpy(h + 36, "data", 4); memcpy(h + 40, &data_bytes, 4);
}
/* ===========================================================================
* HTTP response reader — shared by all three cloud calls.
* Returns the status code and fills body_out. Handles chunked transfer
* encoding PROPERLY: the chunk-size markers must be stripped, or they end
* up embedded inside the JSON body and the parse fails on long replies.
* =========================================================================== */
/* Block until the connection has data (or the budget runs out). Returns false
* on timeout / closed-and-empty. Every read below goes through this, because
* a reasoning model can think for many seconds between the response headers
* and the first byte of the body - and a bare read() would just time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
uint32_t t0 = millis();
while (!c.available()) {
if (!c.connected()) return false;
if (millis() - t0 > ms) return false;
delay(10);
}
return true;
}
static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
body_out = "";
if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
while (waitData(client, idle_ms)) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break; // blank line = end of headers
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
}
if (chunked) {
int blanks = 0;
while (true) {
/* The chunk-size line may not arrive for a long time while the model
* reasons. Waiting here - instead of letting read() time out - is the
* whole fix: a timed-out read looks exactly like "0" (final chunk),
* which silently truncated the body to nothing. */
if (!waitData(client, idle_ms)) {
Serial.println("HTTP: timed out waiting for the next chunk");
break;
}
String szline = client.readStringUntil('\n');
szline.trim();
if (szline.length() == 0) { // stray blank line
if (++blanks > 4) break;
continue;
}
blanks = 0;
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break; // genuine final chunk
long got = 0;
while (got < sz) {
if (!waitData(client, idle_ms)) break;
while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
}
if (waitData(client, 3000)) client.readStringUntil('\n'); // CRLF after chunk
if (got < sz) { Serial.println("HTTP: short chunk"); break; }
}
} else {
while (waitData(client, idle_ms))
while (client.available()) body_out += (char)client.read();
}
return code;
}
/* ===========================================================================
* CLOUD CALL 1 — Azure speech-to-text
* One POST, one header, plain WAV body. This is why Azure does the ears.
* =========================================================================== */
bool azureSTT(size_t audio_bytes, String &text_out) {
/* HTTPClient's one-shot POST fails on bodies this large (it attempts one
* giant TLS write and dies with error -3 SEND_PAYLOAD_FAILED). So this
* function speaks HTTP directly and streams the WAV up in 4 KB chunks -
* reliable, and if it ever stalls we know the exact byte it stopped at. */
WiFiClientSecure client;
client.setInsecure(); // no cert bundle on-device; see notes
client.setTimeout(15); // seconds, for reads
writeWavHeader(wav_buf, audio_bytes);
dumpWavToSD(audio_bytes); // PC-playable copy of what we send
size_t total = WAV_HEADER_LEN + audio_bytes;
if (!client.connect(AZURE_STT_HOST, 443)) {
snprintf(g_stt_err, sizeof(g_stt_err), "TLS connect failed");
Serial.println("STT: TLS connect failed");
return false;
}
/* Two valid host forms use DIFFERENT URL paths - detect which one is in
* secrets.h: <resource>.cognitiveservices.azure.com -> /stt/speech/...
* <region>.stt.speech.microsoft.com -> /speech/... */
bool custom_subdomain = (strstr(AZURE_STT_HOST, ".cognitiveservices.azure.com") != NULL);
String req = String("POST ") + (custom_subdomain ? "/stt" : "") +
"/speech/recognition/conversation/cognitiveservices/v1"
"?language=" AZURE_STT_LANG "&format=simple HTTP/1.1\r\n"
"Host: " AZURE_STT_HOST "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: audio/wav; codecs=audio/pcm; samplerate=16000\r\n"
"Accept: application/json\r\n"
"Connection: close\r\n"
"Content-Length: " + String(total) + "\r\n\r\n";
client.print(req);
/* body, 4 KB at a time */
size_t sent = 0;
while (sent < total) {
size_t n = min((size_t)4096, total - sent);
size_t w = client.write(wav_buf + sent, n);
if (w == 0) {
delay(50); // brief stall - retry once
w = client.write(wav_buf + sent, n);
if (w == 0) {
snprintf(g_stt_err, sizeof(g_stt_err), "upload stalled at %uKB",
(unsigned)(sent / 1024));
Serial.printf("STT: upload stalled at %u/%u bytes\n",
(unsigned)sent, (unsigned)total);
client.stop();
return false;
}
}
sent += w;
yield();
}
Serial.printf("STT: uploaded %u bytes\n", (unsigned)sent);
/* read the reply with proper de-chunking */
String resp;
int code = readHttpResponse(client, resp, 10000);
client.stop();
if (code != 200) {
snprintf(g_stt_err, sizeof(g_stt_err), "HTTP %d", code);
Serial.printf("STT HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_stt_err, sizeof(g_stt_err), "bad JSON reply");
Serial.printf("STT parse error, raw response: %s\n", resp.c_str());
return false;
}
const char *status = doc["RecognitionStatus"];
if (!status || strcmp(status, "Success") != 0) {
/* The status names the exact failure:
* InitialSilenceTimeout = Azure heard silence (mic level too low)
* NoMatch = heard sound but no recognisable words
* BabbleTimeout = heard only noise */
snprintf(g_stt_err, sizeof(g_stt_err), "%s", status ? status : "no status");
Serial.printf("STT status: %s\nraw: %s\n", status ? status : "null", resp.c_str());
return false;
}
text_out = doc["DisplayText"].as<String>();
if (text_out.length() == 0) snprintf(g_stt_err, sizeof(g_stt_err), "empty text");
return text_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 2 — DeepSeek chat completion
* OpenAI-compatible format. Model name is deepseek-v4-flash - the old
* deepseek-chat name is dead, see secrets.h.
* =========================================================================== */
bool deepseekChat(const String &question, String &answer_out) {
/* Manual HTTP, same as azureSTT: HTTPClient truncates larger TLS response
* bodies (long answers + the model's hidden reasoning), which shows up as
* "bad JSON reply". Reading until the server closes the connection is
* reliable regardless of reply length. */
JsonDocument req;
req["model"] = DEEPSEEK_MODEL;
req["max_tokens"] = LLM_MAX_TOKENS;
JsonArray msgs = req["messages"].to<JsonArray>();
JsonObject sys = msgs.add<JsonObject>();
sys["role"] = "system"; sys["content"] = SYSTEM_PROMPT;
JsonObject usr = msgs.add<JsonObject>();
usr["role"] = "user"; usr["content"] = question;
String body;
serializeJson(req, body);
WiFiClientSecure client;
client.setInsecure();
/* 60 s: deepseek-v4-flash is a REASONING model. Easy questions answer in
* ~2 s, but anything that needs actual working-out (an Ohm's law problem,
* say) can think for 10-30 s before sending a single byte. */
client.setTimeout(60);
if (!client.connect(DEEPSEEK_HOST, 443)) {
snprintf(g_llm_err, sizeof(g_llm_err), "TLS connect failed");
Serial.println("LLM: TLS connect failed");
return false;
}
client.print(String("POST /chat/completions HTTP/1.1\r\n"
"Host: " DEEPSEEK_HOST "\r\n"
"Authorization: Bearer " DEEPSEEK_KEY "\r\n"
"Content-Type: application/json\r\n"
"Connection: close\r\n"
"Content-Length: ") + String(body.length()) + "\r\n\r\n");
client.print(body);
/* read the reply with proper de-chunking; generous window - long
* questions make the model think for a while before it responds */
String resp;
int code = readHttpResponse(client, resp, 60000);
client.stop();
if (code != 200) {
snprintf(g_llm_err, sizeof(g_llm_err), "HTTP %d", code);
Serial.printf("LLM HTTP %d: %s\n", code, resp.c_str());
return false;
}
JsonDocument doc;
DeserializationError err = deserializeJson(doc, resp);
if (err) {
snprintf(g_llm_err, sizeof(g_llm_err), "bad JSON reply");
Serial.printf("LLM parse error, raw: %s\n", resp.c_str());
return false;
}
const char *content = doc["choices"][0]["message"]["content"];
const char *finish = doc["choices"][0]["finish_reason"];
if (!content || !content[0]) {
/* v4-flash is a reasoning model: if finish_reason is "length", the whole
* token budget went to internal reasoning - raise LLM_MAX_TOKENS. */
if (finish && strcmp(finish, "length") == 0)
snprintf(g_llm_err, sizeof(g_llm_err), "empty - raise LLM_MAX_TOKENS");
else
snprintf(g_llm_err, sizeof(g_llm_err), "empty content");
Serial.printf("LLM empty content, raw: %s\n", resp.c_str());
return false;
}
answer_out = String(content);
answer_out.trim();
if (answer_out.length() == 0) snprintf(g_llm_err, sizeof(g_llm_err), "blank answer");
return answer_out.length() > 0;
}
/* ===========================================================================
* CLOUD CALL 3 — Azure text-to-speech, streamed straight to the speaker
* We ask for riff-16khz-16bit-mono-pcm: a WAV whose payload is exactly what
* the I2S peripheral eats. Skip the 44-byte header, forward the rest.
* No MP3 decoder, no audio library, no buffering the whole reply.
* =========================================================================== */
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
size_t got = 0;
uint32_t t0 = millis();
while (got < n && millis() - t0 < 10000) {
int r = c.read(dst + got, n - got);
if (r > 0) { got += r; t0 = millis(); }
else if (!c.connected() && !c.available()) break;
else delay(2);
}
return got;
}
bool azureTTSSpeak(const String &text) {
// Escape the XML special characters for the SSML body
String safe = text;
safe.replace("&", "&");
safe.replace("<", "<");
safe.replace(">", ">");
String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
"<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";
/* Manual HTTP like the other two cloud calls - and for a hard reason:
* Azure sends this audio with CHUNKED transfer encoding, and HTTPClient's
* raw stream hands over the chunk framing (ASCII "2000\r\n" lines) mixed
* into the PCM. Played as sound, every chunk boundary is an audible KNOCK.
* Here we parse the framing properly and keep only clean audio bytes. */
WiFiClientSecure client;
client.setInsecure();
client.setTimeout(20);
const char *host = AZURE_REGION ".tts.speech.microsoft.com";
if (!client.connect(host, 443)) {
Serial.println("TTS: TLS connect failed");
return false;
}
client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
"Host: ") + host + "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: application/ssml+xml\r\n"
"X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
"User-Agent: MaTouchRobojax\r\n"
"Connection: close\r\n"
"Content-Length: " + String(ssml.length()) + "\r\n\r\n");
client.print(ssml);
/* status + headers; note whether the body is chunked */
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
long content_len = -1;
while (client.connected() || client.available()) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break;
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
}
if (code != 200) {
Serial.printf("TTS HTTP %d\n", code);
client.stop();
return false;
}
const size_t AUDIO_CAP = 1200 * 1024; // ~37 s of speech
uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
if (!audio) { client.stop(); return false; }
size_t alen = 0;
if (chunked) {
/* chunked: <hex size>\r\n <bytes> \r\n ... 0\r\n\r\n */
while (true) {
String szline = client.readStringUntil('\n');
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break;
if (alen + sz > AUDIO_CAP) break;
size_t got = readExact(client, audio + alen, sz);
alen += got;
client.readStringUntil('\n'); // trailing CRLF after each chunk
if (got < (size_t)sz) break;
}
} else if (content_len > 0) {
alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
} else {
/* no framing info: read until the server closes */
uint32_t idle = millis();
while ((client.connected() || client.available()) && millis() - idle < 5000) {
int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
if (r > 0) { alen += r; idle = millis(); }
else delay(5);
}
}
client.stop();
Serial.printf("TTS: %u KB clean audio (%s), playing\n",
(unsigned)(alen / 1024), chunked ? "de-chunked" : "plain");
bool ok = (alen > WAV_HEADER_LEN);
if (ok) {
/* NOW the audio actually starts - this is the honest moment to go green */
LED_SPEAK();
drawBar("SPEAKING...", gfx->color565(0, 130, 40));
/* skip the RIFF header; a silence pre-roll softens the amp wake-up pop */
static const uint8_t lead_in[640] = {0}; // 20 ms of silence
size_t w = 0;
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, audio + WAV_HEADER_LEN, alen - WAV_HEADER_LEN, &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
}
free(audio);
// let the DMA buffers drain so the last word is not cut off
delay(150);
i2s_zero_dma_buffer(I2S_SPK_PORT);
return ok;
}
/* ===========================================================================
* SETUP
* =========================================================================== */
void setup() {
Serial.begin(115200);
delay(400);
Serial.println("\n=== 04 Voice Assistant | Robojax.com ===");
Serial.println("Mics -> Azure STT -> DeepSeek -> Azure TTS -> speaker");
pinMode(TFT_BLK, OUTPUT);
digitalWrite(TFT_BLK, LOW);
pinMode(SD_CS, OUTPUT);
digitalWrite(SD_CS, HIGH);
// one shared SPI bus for TFT + SD (started before either device)
SPI.begin(TFT_SCLK, TFT_MISO, TFT_MOSI);
gfx->begin();
gfx->fillScreen(BLACK);
digitalWrite(TFT_BLK, HIGH);
// SD is optional here - it only stores the /stt_debug.wav diagnostic copy
ok_sd = SD.begin(SD_CS, SPI, 20000000);
digitalWrite(SD_CS, HIGH);
Serial.println(ok_sd ? "SD ok - will save /stt_debug.wav after each recording"
: "SD not found - debug WAV dump disabled (not fatal)");
bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
delay(50);
rgb.begin();
rgb.setBrightness(LED_BRIGHTNESS);
LED_IDLE();
/* One recording buffer for the whole session, in PSRAM. This is the 8 MB
* that makes the board worth buying. */
wav_buf = (uint8_t *)ps_malloc(WAV_HEADER_LEN + REC_BUF_BYTES);
if (!wav_buf) {
gfx->setTextColor(RED);
gfx->setTextSize(2);
gfx->setCursor(10, 100);
gfx->print("PSRAM alloc failed!");
gfx->setTextSize(1);
gfx->setCursor(10, 130);
gfx->print("Tools > PSRAM > OPI PSRAM must be set.");
while (1) delay(1000);
}
micInit();
spkInit();
gfx->setTextSize(1);
gfx->setTextColor(YELLOW);
gfx->setCursor(4, 4);
gfx->printf("Connecting to %s ...", WIFI_SSID);
Serial.printf("Connecting to %s ", WIFI_SSID);
WiFi.mode(WIFI_STA);
WiFi.begin(WIFI_SSID, WIFI_PASS);
uint32_t t0 = millis();
while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) {
delay(300);
Serial.print(".");
}
Serial.println();
clearChat();
if (WiFi.status() == WL_CONNECTED) {
Serial.printf("Connected, IP %s\n", WiFi.localIP().toString().c_str());
chatBubble("Hold SPEAK and ask me anything.", false);
} else {
chatBubble("WiFi failed. Remember: the ESP32 is 2.4GHz only. Check secrets.h, then press RESET.", false);
LED_ERROR();
}
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
}
/* ===========================================================================
* LOOP — one full conversation turn per button press
* =========================================================================== */
void loop() {
/* CLEAR button: edge-detected so one tap wipes once. Reading the panel
* twice per loop (here and in speakButtonHeld) is fine - the GT911 just
* reports its current state. */
static bool tap_latch = false;
static uint8_t tap_release = 0;
if (state == ST_IDLE) {
uint16_t cx, cy;
if (getTouch(&cx, &cy)) {
tap_release = 0;
if (!tap_latch) {
tap_latch = true;
if (cx >= BTN_CLEAR_X && cx < BTN_CLEAR_X + BTN_CLEAR_W && cy >= BAR_Y) {
clearChat();
chatBubble("Hold SPEAK and ask me anything.", false);
}
}
} else if (tap_latch && ++tap_release >= 4) {
tap_latch = false;
tap_release = 0;
}
/* live WiFi signal indicator, refreshed every 2 s while idle */
static uint32_t last_wifi = 0;
if (millis() - last_wifi > 2000) {
last_wifi = millis();
drawWifi();
}
}
if (state == ST_IDLE && speakButtonHeld()) {
/* ---- record ---- */
state = ST_RECORDING;
LED_LISTEN();
drawBar("LISTENING...", gfx->color565(0, 60, 200));
uint32_t t_rec = millis();
size_t audio_bytes = recordWhileHeld();
t_rec = millis() - t_rec;
Serial.printf("Recorded %u bytes (%.1f s)\n", (unsigned)audio_bytes, audio_bytes / 32000.0);
if (audio_bytes < SAMPLE_RATE / 2) { // under a quarter second - a tap
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
/* ---- speech to text ---- */
state = ST_STT;
LED_THINK();
drawBar("HEARD YOU...", gfx->color565(150, 90, 0));
uint32_t t_stt = millis();
String question;
if (!azureSTT(audio_bytes, question)) {
/* Show the REAL cause on screen - no serial monitor needed. */
char diag[96];
snprintf(diag, sizeof(diag), "STT failed: %s | mic peak %.1f%%%s",
g_stt_err, g_mic_peak_pct,
ok_sd ? " | saved /stt_debug.wav" : "");
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_stt = millis() - t_stt;
chatBubble(question.c_str(), true);
Serial.printf("STT (%lu ms): %s\n", (unsigned long)t_stt, question.c_str());
/* ---- think ---- */
state = ST_LLM;
drawBar("THINKING...", gfx->color565(150, 90, 0));
uint32_t t_llm = millis();
String answer;
if (!deepseekChat(question, answer)) {
char diag[96];
snprintf(diag, sizeof(diag), "DeepSeek failed: %s", g_llm_err);
chatBubble(diag, false);
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
return;
}
t_llm = millis() - t_llm;
chatBubble(answer.c_str(), false);
Serial.printf("LLM (%lu ms): %s\n", (unsigned long)t_llm, answer.c_str());
/* ---- speak ----
* Still amber here: the voice has to be synthesised and downloaded first
* (a few seconds). azureTTSSpeak() itself flips the bar and LED to green
* at the exact moment audio starts coming out of the speaker. */
state = ST_TTS;
LED_THINK();
drawBar("GETTING VOICE", gfx->color565(150, 90, 0));
uint32_t t_tts = millis();
bool spoke = azureTTSSpeak(answer);
t_tts = millis() - t_tts;
/* Timing summary on serial - this feeds the "honest numbers" segment. */
Serial.printf("TIMINGS rec %.1fs | stt %lums | llm %lums | tts %lums%s\n",
audio_bytes / 32000.0, (unsigned long)t_stt,
(unsigned long)t_llm, (unsigned long)t_tts,
spoke ? "" : " (TTS FAILED)");
drawBar("HOLD+TALK", gfx->color565(0, 90, 160));
LED_IDLE();
state = ST_IDLE;
}
delay(20);
}
Rzeczy, których możesz potrzebować
-
InnyProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
Zasoby i odniesienia
-
DokumentacjaMakerfabs MaTouch ESP32-S3 2.8" Camera and Touchscreen: User's Manualwiki.makerfabs.com
-
Dokumentacja
-
DokumentacjaProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
-
PobierzArduino GFX Library on GitHubgithub.com
Pliki📁
Wymagany plik (.h)
Inne Pliki
Schemat
-
MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematicThe latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF0.15 MB