Search Code

Makerfabs MaTouch ESP32-S3 2.8" cameră ESP32-S3 AI Vision: Îndreaptă camera și întreabă ce vede

Makerfabs MaTouch ESP32-S3 2.8" cameră ESP32-S3 AI Vision: Îndreaptă camera și întreabă ce vede

Placa face o fotografie, o IA descrie ceea ce vede și rostește răspunsul cu voce tare

Îndreaptă camera spre ceva, apasă un buton, iar o IA îți spune ce vede - afișat pe ecran și rostit prin difuzor. Două moduri: denumește obiectele din câmpul vizual sau descrie persoana din fața camerei.

AI Vision - Point It At Something running on the MaTouch AI ESP32-S3 board AI Vision - Point It At Something rulând pe placa MaTouch AI ESP32-S3

Cum ajunge imaginea acolo

Camera produce un JPEG de 800×600 de aproximativ 20-40 KB. Placa îl codifică în base64 în PSRAM, ceea ce îl face cu aproximativ o treime mai mare ca text, și îl încorporează în cerere ca URI de date. Cei 8 MB de PSRAM sunt cei care fac acest lucru confortabil.

Descriere, niciodată identificare

Modul PERSON descrie ceea ce poate vedea - starea de spirit, ochelarii, ceea ce pare să facă cineva. Nu îți va spune cine este cineva, iar acest lucru este deliberat. Identificarea persoanelor după față este împotriva politicilor de utilizare ale OpenAI, iar serviciul echivalent de identificare facială al Microsoft este blocat în spatele unui proces de aprobare la care dezvoltatorii obișnuiți nu se pot înscrie pur și simplu. Orice tutorial care îți promite identificare facială în cloud descrie ceva care nu este disponibil în general.

Dacă vrei o placă care recunoaște persoane specifice, folosește proiectul 03b - o face offline și nu trimite niciodată fața nimănui nicăieri.

Camera trebuie să se oprească pentru rețea

Viewfinder-ul live este oprit în timp ce se pune întrebarea. Acest lucru nu este cosmetic: driverul camerei care rulează deține memoria internă de care are nevoie o conexiune securizată, iar cu previzualizarea activă conexiunea pur și simplu eșuează. Oprirea camerei o eliberează. Înseamnă, de asemenea, că răspunsul este lizibil în loc să fie repictat de peste 25 de ori pe secundă.

Răspunsul rămâne până când îl închizi

Răspunsul rămâne pe ecran până când apeși CLEAR - nu există cronometru. În timp ce este afișat, un buton REPLAY redă răspunsul rostit din memorie, fără un al doilea apel API și fără costuri suplimentare.

Despre placa MaTouch AI ESP32-S3 2.8"

Fiecare proiect de pe această pagină rulează pe MaTouch AI ESP32-S3 2.8" TFT ST7789V de la Makerfabs. Este o placă all-in-one: un ecran tactil color, o cameră de 3 megapixeli, două microfoane și un amplificator real pentru difuzor, toate conduse de un ESP32-S3 cu 8 MB de PSRAM. Această combinație este cea care face posibile aceste proiecte AI pe o singură placă, fără nimic altceva atașat.

Cei 8 MB de PSRAM contează mai mult decât orice alt număr de aici. Este ceea ce permite plăcii să țină în memorie simultan un cadru al camerei, câteva secunde de audio înregistrat sau o fotografie codificată în base64 - nimic din toate acestea nu încape în RAM-ul normal al ESP32.

Documentația producătorului: pagina wiki Makerfabs.

Specificații cheie

  • Procesor: ESP32-S3, dual core 240 MHz, WiFi 2.4 GHz + Bluetooth 5.0

  • Memorie: 16 MB flash, 8 MB PSRAM (necesară pentru aproape fiecare proiect de aici)

  • Display: 2.8" IPS, 320×240, driver ST7789V, SPI

  • Touch: GT911 capacitiv, urmărește 5 degete simultan

  • Cameră: OV3660, 3 megapixeli, până la 2048×1536

  • Microfoane: două microfoane digitale I2S INMP441 (o pereche stereo autentică)

  • Difuzor: amplificator clasa D MAX98357A, 3.2 W pe 4 Ω

  • Stocare: slot card microSD (mod SPI)

  • Alimentare: USB-C, conector baterie JST, încărcător TP4056, întrerupător de alimentare

  • De asemenea, pe placă: LED RGB WS2812B, ceas de timp real cu baterie PCF8563T și un indicator al nivelului bateriei MAX17048 care nu este listat în specificațiile oficiale

Cele două porturi USB-C nu sunt identice. Difuzorul plăcii împarte pinii de semnal (IO19 și IO20) cu portul USB nativ, deoarece acești pini sunt liniile de date USB hardwired ale ESP32-S3. Încarcă și alimentează întotdeauna prin portul USB-C CH340K (cel de lângă butonul RESET) și setează USB CDC On Boot la Disabled. Folosește portul greșit și audio-ul se va comporta necorespunzător sau încărcările vor eșua.

Setări Arduino IDE

Aceste setări contează. Cele mai multe probleme raportate de oameni cu această placă sunt una dintre acestea fiind greșită și se resetează când schimbi versiunea nucleului, așa că verifică-le din nou după orice modificare.

Setare

Valoare

Placă

ESP32S3 Dev Module

Versiune nucleu ESP32

2.0.17

PSRAM

OPI PSRAM

Dimensiune Flash

16MB (128Mb)

Schema de partiții

16M Flash (3MB APP/9.9MB FATFS)

USB CDC On Boot

Disabled

Viteză încărcare

921600

Șterge tot flash-ul înainte de încărcare

Disabled

Port

portul USB-C CH340K

Folosește nucleul ESP32 2.0.17, nu 3.x. Espressif a eliminat modelele de detectare facială de pe dispozitiv în nucleul 3, astfel încât proiectele faciale nu vor compila acolo. Fixarea versiunii 2.0.17 menține fiecare proiect de pe această pagină funcțional cu o singură configurație. În Boards Manager, meniul derulant de versiuni îți permite să comuți înainte și înapoi oricând dorești.

Folosiți GFX Library for Arduino versiunea 1.5.6, nu 1.6.x. Versiunile 1.6 sunt create pentru ESP32 core 3 și pot rămâne blocate la pornire pe core 2.0.17. Dacă ecranul rămâne negru după încărcare, acesta este primul lucru de verificat.

Biblioteci necesare

Instalați-le prin Tools → Manage Libraries în Arduino IDE. Numerele de versiune contează - vă rugăm să folosiți cele listate.

Bibliotecă

Versiune

Autor

GFX Library for Arduino

1.5.6

moononournation

bb_captouch

1.3.1

Larry Bank

ArduinoJson

7.x

Benoit Blanchon

Adafruit NeoPixel

orice recentă

Adafruit

Configurarea secrets.h

Detaliile WiFi și orice chei API se pun în secrets.h, care este inclus în descărcare cu valori placeholder. Deschideți acea filă în Arduino IDE și înlocuiți-le cu ale dumneavoastră.

WiFi-ul trebuie să fie de 2,4 GHz. ESP32-S3 nu poate vedea deloc o rețea de 5 GHz. Dacă routerul dumneavoastră combină ambele benzi sub un singur nume (Asus numește aceasta Smart Connect), fie dezactivați această funcție, fie dați benzii de 2,4 GHz un nume propriu și folosiți-l în secrets.h.

Obținerea cheilor API

Acest proiect comunică cu un serviciu AI cloud, deci aveți nevoie de propria cheie. Dacă nu ați făcut niciodată acest lucru, nu vă îngrijorați - este aceeași idee ca o parolă care identifică contul dumneavoastră la serviciu. Durează câteva minute, o singură dată.

O cheie nu este un abonament la un site web. De exemplu, plata pentru ChatGPT Plus nu vă oferă o cheie API - cele două sunt produse separate cu facturare separată. Aveți nevoie de un cont pe platforma de dezvoltatori, descrisă mai jos.

OpenAI - ochii

OpenAI oferă modelul de viziune care se uită la imagini. Doar proiectele care trimit o imagine au nevoie de acesta.

  1. Mergeți la platform.openai.com/api-keys și conectați-vă sau creați un cont.

  2. Faceți clic pe Create new secret key, dați-i un nume și creați-o.

  3. Copiați-o imediat - ca la DeepSeek, este afișată doar o dată.

  4. Deschideți Billing și adăugați o mică sumă de credit. API-ul este preplătit și separat de orice abonament ChatGPT pe care îl aveți deja.

Puneți cheia în secrets.h ca OPENAI_KEY.

Microsoft Azure Speech - pentru ascultare și vorbire

Azure transformă vorbirea dumneavoastră în text și răspunsul înapoi în voce. Nivelul gratuit este suficient de generos pentru tot ce este pe această pagină.

  1. Mergeți la portal.azure.com și conectați-vă cu un cont Microsoft (unul gratuit este suficient).

  2. Dacă nu ați folosit niciodată Azure, veți vedea un ecran Welcome to Azure care oferă trei opțiuni. Alegeți Start with an Azure free trial - aveți nevoie de un abonament înainte ca Azure să vă permită să creați ceva. (Studenții ar trebui să aleagă în schimb Azure for Students: același rezultat, fără card necesar.) Ignorați Manage Microsoft Entra ID, care este cu totul altceva.

  3. Faceți clic pe Create a resource, căutați Speech și alegeți Speech service publicat de Microsoft.

  4. Completați formularul: orice grup de resurse, orice nume și alegeți o Regiune aproape de dumneavoastră - notați acea regiune exact cum apare, de exemplu eastus.

  5. Pentru Pricing tier alegeți F0 (Free). Aceasta permite aproximativ cinci ore de vorbire în text și jumătate de milion de caractere de text în vorbire în fiecare lună.

  6. Faceți clic pe Review + create, apoi pe Create. Așteptați aproximativ un minut, apoi faceți clic pe Go to resource.

  7. În meniul din stânga deschideți Keys and Endpoint. Copiați KEY 1 și Location/Region.

Puneți-le în secrets.h ca AZURE_SPEECH_KEY și AZURE_REGION. Pentru AZURE_STT_HOST, folosiți <regiune>.stt.speech.microsoft.com - deci cu regiunea eastus aceasta este eastus.stt.speech.microsoft.com.

Despre cardul de credit. Procesul gratuit de încercare Azure cere un card pentru a vă verifica identitatea. Nu vă taxează. Primiți 200 de dolari credit pentru 30 de zile, iar după aceea contul trece la Pay-As-You-Go - dar nivelul F0 Speech rămâne gratuit, lună de lună, iar tot ce este în aceste proiecte se încadrează confortabil în el. Dacă preferați să nu dați deloc un card și sunteți student, opțiunea Azure for Students vă oferă credit fără unul.

Trebuie să fie o resursă "Speech service". O cheie de la o resursă Translator, Language sau Cognitive Services generală arată identic și este perfect validă - dar fiecare cerere de vorbire returnează eroarea 401. Aceasta ne-a prins în timpul testării și ne-a costat o oră. Dacă vorbirea eșuează cu 401 în timp ce cheia pare corectă, verificați ce tip de resursă ați creat.

Cât costă să ruleze

Foarte puțin, dar nu este gratuit și ar trebui să știți aproximativ cât cheltuiți înainte de a lăsa un proiect să ruleze.

Serviciu

Cost aproximativ

Azure Speech

nivelul gratuit acoperă aproximativ 5 ore de ascultare și 0,5 M de caractere de vorbire pe lună

DeepSeek

o fracțiune de cent pe răspuns - mii de răspunsuri pentru câțiva dolari

OpenAI vision

aproximativ un cent sau doi pe imagine, în funcție de model

Prețurile se schimbă, așa că tratați-le ca pe un ghid, nu ca pe o ofertă. Fiecare dintre aceste servicii are o pagină de utilizare unde puteți urmări ce ați cheltuit, iar toate vă permit să setați o limită de cheltuieli - ceea ce merită făcut din prima zi.

Păstrați-vă cheile private. Oricine le are le poate cheltui banii. Nu le puneți într-un videoclip, o captură de ecran, o postare pe forum sau un depozit de cod public. Dacă o cheie este vreodată expusă, ștergeți-o pe site-ul furnizorului și creați una nouă - durează secunde și este singura soluție reală.

Depanare

Simptom

Cauză și soluție

Ecranul rămâne negru

Versiune greșită a bibliotecii GFX (folosiți 1.5.6) sau setări incorecte ale plăcii.

PSRAM alloc failed sau eroare cameră 0xffffffff

Tools → PSRAM nu este setat pe OPI PSRAM.

Nu se încarcă nimic / niciun port COM

Port USB-C greșit sau driverul CH340 nu este instalat.

Camera eșuează și nu se recuperează

Linia de reset a camerei este legată de butonul RESET al plăcii, astfel încât software-ul nu o poate reporni. Apăsați RESET. Dacă încă eșuează, repoziționați cablul plat al camerei.

Descărcați codul

Schița Arduino completă pentru acest proiect, împreună cu pins.h și tot ce mai este necesar, este disponibilă gratuit pentru descărcare.

Descărcați 05_Vision_AI

Dezarhivați-o, deschideți fișierul .ino în Arduino IDE, verificați setările de mai sus și încărcați prin portul USB-C CH340K.

885-Arduin code for MaTouch AI ESP32S3 2.8in AI Camera: AI Vision
Limba: C++
/*
 * ===========================================================================
 *  05_Vision_AI  —  MaTouch AI ESP32-S3 2.8" TFT ST7789V
 * ===========================================================================
 *
 ----------
 *  ROBOJAX.COM  -  MaTouch AI ESP32-S3 2.8" project series
 *
 *    WATCH THE VIDEO
 *        https://youtu.be/6AL3g3tC_Hk
 *
 *    WRITTEN TUTORIALS - every project, with photos and full explanation
 *        Camera and touchscreen.... https://robojax.com/RTJ849
 *        Offline face recognition.. https://robojax.com/RTJ850
 *        AI voice assistant........ https://robojax.com/RTJ851
 *        AI vision................. https://robojax.com/RTJ852
 *
 *    GET THE BOARD - SAVE $5 with coupon code:  Robojax_Makerfab
 *        https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
 *        (enter the code at checkout)
 *
 *  All of this code is free. If it helped you, a subscribe on YouTube is
 *  the best way to support more of it.
 *  
 *
 *  Watching the video first will save you time - it shows the Arduino IDE
 *  settings and the library versions being set up step by step.
 *  ---------------------------------------------------------------------------
 *
 *  The board SEES. Point the camera at something, touch a button, and an
 *  OpenAI vision model tells you what it is looking at - drawn on the screen
 *  and (optionally) spoken out of the board's own speaker via Azure TTS.
 *
 *  Two modes, two buttons:
 *
 *      OBJECTS   "What do you see?" - names the things in front of the lens.
 *      PERSON    Describes the person in frame: glasses, expression, what
 *                they are doing. DESCRIPTION ONLY - it will not and must not
 *                try to say WHO someone is. Identifying people by face is
 *                against OpenAI's usage policies, and it is the right call:
 *                say this in the video, it is worth 15 honest seconds.
 *
 *  WHY OPENAI FOR THIS DEMO AND NOT DEEPSEEK: DeepSeek's API is text-only.
 *  It cannot accept an image at all. Azure OpenAI could do it, but plain
 *  OpenAI is one endpoint with no deployment setup - simplest to follow.
 *
 *  HOW THE IMAGE TRAVELS: the OV3660 gives us a JPEG directly (800x600,
 *  ~40 KB). We base64-encode it in PSRAM (~55 KB of text) and embed it in
 *  the JSON request as a data: URI. The 8 MB PSRAM makes this trivial.
 *
 *  ---------------------------------------------------------------------------
 *  FILL IN secrets.h BEFORE FLASHING
 *  (needs WIFI_*, OPENAI_*; AZURE_* only if SPEAK_REPLIES is 1).
 *
 *  BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
 *  or compile errors. These reset when you switch cores - recheck them!)
 *
 *      Board            : ESP32S3 Dev Module
 *      ESP32 core       : 2.0.17
 *      PSRAM            : OPI PSRAM        <-- required, image buffers live there
 *      Flash Size       : 16MB (128Mb)
 *      Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
 *      USB CDC On Boot  : Disabled         <-- speaker shares pins with native USB
 *      Upload Speed     : 921600
 *      Port             : the CH340K USB-C port (the one near RESET)
 *
 *  LIBRARIES
 *      GFX Library for Arduino   v1.5.6   (NOT 1.6.x - that pairs with core 3)
 *      bb_captouch               v1.3.1
 *      ArduinoJson               v7.x
 *      Adafruit NeoPixel         any recent
 *
 *  ---------------------------------------------------------------------------
 *  FUNCTIONS IN THIS SKETCH
 *      led(r,g,b)          set the RGB status LED colour
 *      getTouch(&x,&y)     read the touch panel, mapped to screen coordinates
 *      camStart(...)       start the camera in a given format/size
 *      camStartPreview()   start the 240x240 RGB565 live-view camera
 *      captureJpeg(&len)   take one 800x600 JPEG into PSRAM (camera stays OFF
 *                          afterwards - caller restarts the preview)
 *      readHttpResponse()  read an HTTPS reply, de-chunking it properly
 *      askVision(...)      send photo + prompt to OpenAI, return the answer
 *      spkInit()           configure the I2S speaker output
 *      readExact(...)      read exactly N bytes from a TLS connection
 *      speak(text)         Azure TTS -> download voice to PSRAM -> play it
 *      playLastAnswer()    replay the kept voice from PSRAM (REPLAY button)
 *      drawButton(...)     draw one side-column button
 *      drawWifi()          WiFi signal bars + dBm readout
 *      drawButtons()       normal side column (the two ask buttons)
 *      drawClearSide()     answer-mode side column (CLEAR + REPLAY)
 *      showAnswer(text)    word-wrapped answer overlay on the viewfinder
 *      lookAndTell(...)    one full cycle: capture -> ask -> show -> speak
 *      setup() / loop()    boot sequence / viewfinder + touch handling
 *
 *  Robojax.com
 * ===========================================================================
 */

#define SPEAK_REPLIES  1     // 1 = read the answer aloud with Azure TTS, 0 = screen only

/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
 * uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS 15

/* CAMERA ORIENTATION
 * 1 = camera faces the SAME way as the screen (how the board ships) - the
 *     image is mirrored so it looks natural when you point it at yourself.
 * 0 = you folded the ribbon so the lens faces AWAY from the screen
 *     (phone-style, screen to you / camera to the subject).
 * If the picture looks left-right reversed, flip this number. */
#define CAMERA_FACES_USER  1

#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <Wire.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include "mbedtls/base64.h"
#include "esp_camera.h"
#include "pins.h"
#include "secrets.h"

#if SPEAK_REPLIES
#include "driver/i2s.h"
#define WAV_HEADER_LEN 44
#endif

Arduino_ESP32SPI *bus = new Arduino_ESP32SPI(
    TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, HSPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);

BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);

/* --- the two vision prompts ------------------------------------------------
 * Keep answers short: they must fit a 320x240 screen and, if spoken, must not
 * leave the presenter waiting awkwardly on camera. */
const char *PROMPT_OBJECTS =
    "Look at this photo from a small camera. In ONE short sentence of at most "
    "20 words, plain text only, name the main object(s) you see.";

const char *PROMPT_PERSON =
    "Look at this photo. If there is a person, describe them in ONE short "
    "sentence of at most 20 words: mood, glasses or not, what they are doing. "
    "Never guess who they are. If no person is visible, say so. Plain text only.";

/* --- layout: viewfinder left, buttons right ------------------------------- */
#define BTN_X   242
#define BTN_W    78
#define BTN_H    52
#define BTN_OBJ_Y     4
#define BTN_PERSON_Y 62

void led(uint8_t r, uint8_t g, uint8_t b) { rgb.setPixelColor(0, rgb.Color(r, g, b)); rgb.show(); }


/* ===========================================================================
 *  Touch
 * =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
  TOUCHINFO ti;
  if (!bbct.getSamples(&ti)) return false;
  if (ti.count < 1) return false;
  *x = ti.y[0];
  *y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
  return true;
}


/* ===========================================================================
 *  Camera  —  RGB565 for the live view; JPEG capture happens by restarting
 *  the driver, same technique as sketch 02.
 * =========================================================================== */
bool mirrored = CAMERA_FACES_USER;   // see the define at the top of the file

/* The viewfinder is OFF while a cloud call runs and while the answer is on
 * screen. Two reasons, both learned on the bench:
 *  1. RAM: the live camera driver eats the internal memory a TLS handshake
 *     needs - with the preview running, connecting to OpenAI fails (HTTP -1).
 *  2. Readability: the viewfinder repaints 25x/s and would wipe the answer.
 * The answer STAYS on screen until the user taps CLEAR - no timer. */
bool preview_on = false;

/* The last spoken answer is KEPT in PSRAM so the REPLAY button can play it
 * again without another cloud call. Overwritten by the next answer. */
uint8_t *last_audio = nullptr;
size_t   last_audio_len = 0;

bool camStart(pixformat_t fmt, framesize_t size, int fb_count, int quality) {
  camera_config_t c;
  c.ledc_channel = LEDC_CHANNEL_0;
  c.ledc_timer   = LEDC_TIMER_0;
  c.pin_d0 = CAM_PIN_D0;  c.pin_d1 = CAM_PIN_D1;
  c.pin_d2 = CAM_PIN_D2;  c.pin_d3 = CAM_PIN_D3;
  c.pin_d4 = CAM_PIN_D4;  c.pin_d5 = CAM_PIN_D5;
  c.pin_d6 = CAM_PIN_D6;  c.pin_d7 = CAM_PIN_D7;
  c.pin_xclk     = CAM_PIN_XCLK;
  c.pin_pclk     = CAM_PIN_PCLK;
  c.pin_vsync    = CAM_PIN_VSYNC;
  c.pin_href     = CAM_PIN_HREF;
  /* Share the I2C bus Wire already drives (the touch panel lives there too)
   * instead of letting the camera install a second driver on the same pins -
   * that kills touch. Wire.begin() must run before this function. */
  c.pin_sccb_sda = -1;
  c.pin_sccb_scl = -1;
  c.sccb_i2c_port = 0;               // Wire = I2C port 0
  c.pin_pwdn     = CAM_PIN_PWDN;
  c.pin_reset    = CAM_PIN_RESET;
  c.xclk_freq_hz = 20000000;
  c.frame_size   = size;
  c.pixel_format = fmt;
  c.grab_mode    = CAMERA_GRAB_WHEN_EMPTY;
  c.fb_location  = CAMERA_FB_IN_PSRAM;
  c.jpeg_quality = quality;
  c.fb_count     = fb_count;

  if (esp_camera_init(&c) != ESP_OK) return false;

  sensor_t *s = esp_camera_sensor_get();
  if (s) {
    s->set_hmirror(s, mirrored ? 1 : 0);
    s->set_vflip(s,   mirrored ? 1 : 0);
    s->set_brightness(s, 1);
  }
  return true;
}

bool camStartPreview() { return camStart(PIXFORMAT_RGB565, FRAMESIZE_240X240, 2, 12); }

/* Capture one SVGA JPEG into a PSRAM buffer the caller owns. */
uint8_t *captureJpeg(size_t *len_out) {
  esp_camera_deinit();
  delay(120);
  if (!camStart(PIXFORMAT_JPEG, FRAMESIZE_SVGA, 1, 12)) { *len_out = 0; return nullptr; }

  // a few warm-up frames so exposure settles
  for (int i = 0; i < 3; i++) {
    camera_fb_t *w = esp_camera_fb_get();
    if (w) esp_camera_fb_return(w);
    delay(100);
  }

  uint8_t *copy = nullptr;
  *len_out = 0;
  camera_fb_t *fb = esp_camera_fb_get();
  if (fb && fb->len > 0) {
    copy = (uint8_t *)ps_malloc(fb->len);
    if (copy) { memcpy(copy, fb->buf, fb->len); *len_out = fb->len; }
  }
  if (fb) esp_camera_fb_return(fb);

  /* Deliberately leave the camera OFF here. The caller restarts the preview
   * after the cloud call - a running camera driver starves the TLS handshake
   * of internal RAM (that was the "Vision HTTP -1" bug). */
  esp_camera_deinit();
  delay(120);
  return copy;
}


/* ===========================================================================
 *  HTTP response reader  —  shared by the cloud calls below.
 *  Returns the status code and fills body_out. Handles chunked transfer
 *  encoding PROPERLY: the chunk-size markers must be stripped, or they end
 *  up embedded inside the JSON and the parse fails. (OpenAI pretty-prints
 *  its replies so they span several chunks - this bit us on the bench.)
 * =========================================================================== */
/* Block until the connection has data (or the budget runs out). Every read
 * below goes through this, because the server can go quiet for many seconds
 * while it analyses the image - and a bare read() would simply time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
  uint32_t t0 = millis();
  while (!c.available()) {
    if (!c.connected()) return false;
    if (millis() - t0 > ms) return false;
    delay(10);
  }
  return true;
}

static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
  body_out = "";

  if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);

  bool chunked = false;
  while (waitData(client, idle_ms)) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;         // blank line = end of headers
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
  }

  if (chunked) {
    int blanks = 0;
    while (true) {
      /* Waiting here - instead of letting read() time out - is essential: a
       * timed-out read looks exactly like "0" (final chunk), which silently
       * truncates the body to nothing while the server is still thinking. */
      if (!waitData(client, idle_ms)) {
        Serial.println("HTTP: timed out waiting for the next chunk");
        break;
      }
      String szline = client.readStringUntil('\n');
      szline.trim();
      if (szline.length() == 0) {
        if (++blanks > 4) break;
        continue;
      }
      blanks = 0;

      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;                            // genuine final chunk

      long got = 0;
      while (got < sz) {
        if (!waitData(client, idle_ms)) break;
        while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
      }
      if (waitData(client, 3000)) client.readStringUntil('\n');
      if (got < sz) { Serial.println("HTTP: short chunk"); break; }
    }
  } else {
    while (waitData(client, idle_ms))
      while (client.available()) body_out += (char)client.read();
  }
  return code;
}


/* ===========================================================================
 *  OpenAI vision call
 *  The request body is built by hand in PSRAM rather than through ArduinoJson,
 *  because embedding a 55 KB base64 string in a JSON document would mean
 *  holding two copies. Base64 text never needs JSON escaping, so this is safe.
 * =========================================================================== */
bool askVision(const uint8_t *jpg, size_t jpg_len, const char *prompt, String &answer_out) {
  // --- base64 encode the image into PSRAM ---
  size_t b64_cap = ((jpg_len + 2) / 3) * 4 + 16;
  unsigned char *b64 = (unsigned char *)ps_malloc(b64_cap);
  if (!b64) return false;
  size_t b64_len = 0;
  if (mbedtls_base64_encode(b64, b64_cap, &b64_len, jpg, jpg_len) != 0) {
    free(b64);
    return false;
  }

  // --- assemble the JSON request around it ---
  /* Current OpenAI models reject the old "max_tokens" name - it must be
   * "max_completion_tokens" (verified against the live API, July 2026). */
  const char *head_fmt =
      "{\"model\":\"%s\",\"max_completion_tokens\":%d,\"messages\":[{\"role\":\"user\","
      "\"content\":[{\"type\":\"text\",\"text\":\"%s\"},"
      "{\"type\":\"image_url\",\"image_url\":{\"url\":\"data:image/jpeg;base64,";
  const char *tail = "\"}}]}]}";

  char head[640];
  int head_len = snprintf(head, sizeof(head), head_fmt, OPENAI_MODEL, LLM_MAX_TOKENS, prompt);

  size_t body_len = head_len + b64_len + strlen(tail);
  char *body = (char *)ps_malloc(body_len + 1);
  if (!body) { free(b64); return false; }
  memcpy(body, head, head_len);
  memcpy(body + head_len, b64, b64_len);
  strcpy(body + head_len + b64_len, tail);
  free(b64);

  // --- send it: manual HTTP, chunk-wise upload (proven pattern from 04) ---
  WiFiClientSecure client;
  client.setInsecure();
  client.setTimeout(25);

  if (!client.connect(OPENAI_HOST, 443)) {
    Serial.println("Vision: TLS connect failed (is the camera still running?)");
    free(body);
    return false;
  }

  client.print(String("POST /v1/chat/completions HTTP/1.1\r\n"
               "Host: " OPENAI_HOST "\r\n"
               "Authorization: Bearer " OPENAI_KEY "\r\n"
               "Content-Type: application/json\r\n"
               "Connection: close\r\n"
               "Content-Length: ") + String(body_len) + "\r\n\r\n");

  size_t sent = 0;
  while (sent < body_len) {
    size_t n = min((size_t)4096, body_len - sent);
    size_t w = client.write((uint8_t *)body + sent, n);
    if (w == 0) {
      delay(50);
      w = client.write((uint8_t *)body + sent, n);
      if (w == 0) break;
    }
    sent += w;
    yield();
  }
  free(body);
  if (sent < body_len) {
    Serial.printf("Vision: upload stalled at %u/%u bytes\n",
                  (unsigned)sent, (unsigned)body_len);
    client.stop();
    return false;
  }

  /* read the reply with proper de-chunking */
  String resp;
  int code = readHttpResponse(client, resp, 30000);
  client.stop();

  if (code != 200) {
    Serial.printf("Vision HTTP %d: %s\n", code, resp.c_str());
    return false;
  }

  bool ok = false;
  JsonDocument doc;
  if (!deserializeJson(doc, resp)) {
    const char *content = doc["choices"][0]["message"]["content"];
    if (content) {
      answer_out = String(content);
      answer_out.trim();
      ok = answer_out.length() > 0;
    }
  } else {
    Serial.printf("Vision: JSON parse failed, %u bytes received\n", resp.length());
  }
  return ok;
}


/* ===========================================================================
 *  Azure TTS  —  same streaming trick as sketch 04: raw PCM into I2S.
 * =========================================================================== */
#if SPEAK_REPLIES
void spkInit() {
  i2s_config_t cfg = {
    .mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
    .sample_rate = 16000,
    .bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
    .channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
    .communication_format = I2S_COMM_FORMAT_STAND_I2S,
    .intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
    .dma_buf_count = 8,
    .dma_buf_len = 256,
    .use_apll = false,
    .tx_desc_auto_clear = true,
    .fixed_mclk = 0
  };
  i2s_pin_config_t pins = {
    .mck_io_num = I2S_PIN_NO_CHANGE,
    .bck_io_num = I2S_SPK_BCLK,
    .ws_io_num = I2S_SPK_LRC,
    .data_out_num = I2S_SPK_DOUT,
    .data_in_num = I2S_PIN_NO_CHANGE
  };
  i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
  i2s_set_pin(I2S_SPK_PORT, &pins);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
}

/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
  size_t got = 0;
  uint32_t t0 = millis();
  while (got < n && millis() - t0 < 10000) {
    int r = c.read(dst + got, n - got);
    if (r > 0) { got += r; t0 = millis(); }
    else if (!c.connected() && !c.available()) break;
    else delay(2);
  }
  return got;
}

void playLastAnswer();   // defined below; explicit prototype for the IDE

void speak(const String &text) {
  String safe = text;
  safe.replace("&", "&amp;");
  safe.replace("<", "&lt;");
  safe.replace(">", "&gt;");
  String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
                "<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";

  /* Manual HTTP with proper de-chunking: Azure sends this audio chunked, and
   * HTTPClient's raw stream leaks the ASCII chunk-size lines into the PCM -
   * each one plays as an audible KNOCK. */
  WiFiClientSecure client;
  client.setInsecure();
  client.setTimeout(20);

  const char *host = AZURE_REGION ".tts.speech.microsoft.com";
  if (!client.connect(host, 443)) { Serial.println("TTS: TLS connect failed"); return; }

  client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
               "Host: ") + host + "\r\n"
               "Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
               "Content-Type: application/ssml+xml\r\n"
               "X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
               "User-Agent: MaTouchRobojax\r\n"
               "Connection: close\r\n"
               "Content-Length: " + String(ssml.length()) + "\r\n\r\n");
  client.print(ssml);

  String status_line = client.readStringUntil('\n');
  int code = 0;
  sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
  bool chunked = false;
  long content_len = -1;
  while (client.connected() || client.available()) {
    String h = client.readStringUntil('\n');
    if (h == "\r" || h.length() <= 1) break;
    h.toLowerCase();
    if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
    if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
  }
  if (code != 200) { Serial.printf("TTS HTTP %d\n", code); client.stop(); return; }

  const size_t AUDIO_CAP = 1200 * 1024;
  uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
  if (!audio) { client.stop(); return; }
  size_t alen = 0;

  if (chunked) {
    while (true) {
      String szline = client.readStringUntil('\n');
      long sz = strtol(szline.c_str(), NULL, 16);
      if (sz <= 0) break;
      if (alen + sz > AUDIO_CAP) break;
      size_t got = readExact(client, audio + alen, sz);
      alen += got;
      client.readStringUntil('\n');
      if (got < (size_t)sz) break;
    }
  } else if (content_len > 0) {
    alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
  } else {
    uint32_t idle = millis();
    while ((client.connected() || client.available()) && millis() - idle < 5000) {
      int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
      if (r > 0) { alen += r; idle = millis(); }
      else delay(5);
    }
  }
  client.stop();

  if (alen > WAV_HEADER_LEN) {
    /* keep this answer for the REPLAY button (replacing the previous one),
     * then play it */
    if (last_audio) free(last_audio);
    last_audio = audio;
    last_audio_len = alen;
    playLastAnswer();
  } else {
    free(audio);
  }
}

/* Play the kept answer from PSRAM - used right after download AND by REPLAY. */
void playLastAnswer() {
  if (!last_audio || last_audio_len <= WAV_HEADER_LEN) return;
  static const uint8_t lead_in[640] = {0};     // 20 ms silence pre-roll
  size_t w = 0;
  i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
  i2s_write(I2S_SPK_PORT, last_audio + WAV_HEADER_LEN,
            last_audio_len - WAV_HEADER_LEN, &w, portMAX_DELAY);
  i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
  delay(150);
  i2s_zero_dma_buffer(I2S_SPK_PORT);
}
#endif  // SPEAK_REPLIES


/* ===========================================================================
 *  UI
 * =========================================================================== */
void drawButton(int y, const char *l1, const char *l2, uint16_t colour) {
  gfx->fillRoundRect(BTN_X, y, BTN_W, BTN_H, 6, colour);
  gfx->drawRoundRect(BTN_X, y, BTN_W, BTN_H, 6, WHITE);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_X + 8, y + 14);
  gfx->print(l1);
  gfx->setCursor(BTN_X + 8, y + 28);
  gfx->print(l2);
}

/* WiFi bars + dBm at the bottom of the side column, refreshed from the loop */
void drawWifi() {
  gfx->fillRect(BTN_X, 188, BTN_W, 52, BLACK);
  bool up = (WiFi.status() == WL_CONNECTED);
  long rssi = up ? WiFi.RSSI() : -100;
  int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;

  for (int b = 0; b < 4; b++) {
    int bh = 6 + b * 6;
    uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
    gfx->fillRect(BTN_X + 4 + b * 9, 216 - bh, 7, bh, col);
  }
  gfx->setTextSize(1);
  gfx->setCursor(BTN_X + 44, 196);
  if (up) {
    gfx->setTextColor(CYAN);
    gfx->printf("%ld", rssi);
  } else {
    gfx->setTextColor(RED);
    gfx->print("DOWN");
  }
  gfx->setCursor(BTN_X + 44, 208);
  gfx->setTextColor(gfx->color565(120, 120, 120));
  gfx->print("dBm");
  gfx->setCursor(BTN_X + 4, 228);
  gfx->setTextColor(gfx->color565(120, 120, 120));
  gfx->print("WiFi signal");
}

/* normal side column: the two ask buttons */
void drawButtons() {
  gfx->fillRect(BTN_X, 0, BTN_W, 188, BLACK);
  drawButton(BTN_OBJ_Y,    "WHAT DO",  "YOU SEE?", gfx->color565(0, 90, 160));
  drawButton(BTN_PERSON_Y, "DESCRIBE", "PERSON",   gfx->color565(120, 60, 140));

  gfx->setTextColor(CYAN);
  gfx->setCursor(BTN_X + 2, 128);
  gfx->print("OpenAI eyes");
  gfx->setCursor(BTN_X + 2, 140);
  gfx->print("Azure voice");
  gfx->setCursor(BTN_X + 2, 152);
  gfx->print("Robojax.com");

  drawWifi();
}

/* answer-mode side column: CLEAR (back to camera) + REPLAY (say it again).
 * The answer stays on screen until CLEAR is tapped. */
void drawClearSide() {
  gfx->fillRect(BTN_X, 0, BTN_W, 188, BLACK);
  gfx->fillRoundRect(BTN_X, BTN_OBJ_Y, BTN_W, BTN_H, 6, gfx->color565(110, 35, 35));
  gfx->drawRoundRect(BTN_X, BTN_OBJ_Y, BTN_W, BTN_H, 6, WHITE);
  gfx->setTextSize(2);
  gfx->setTextColor(WHITE);
  gfx->setCursor(BTN_X + 10, BTN_OBJ_Y + 14);
  gfx->print("CLEAR");

#if SPEAK_REPLIES
  if (last_audio_len > 0) {
    gfx->fillRoundRect(BTN_X, BTN_PERSON_Y, BTN_W, BTN_H, 6, gfx->color565(0, 110, 60));
    gfx->drawRoundRect(BTN_X, BTN_PERSON_Y, BTN_W, BTN_H, 6, WHITE);
    gfx->setTextSize(1);
    gfx->setTextColor(WHITE);
    gfx->setCursor(BTN_X + 16, BTN_PERSON_Y + 20);
    gfx->print("REPLAY");
  }
#endif

  gfx->setTextSize(1);
  gfx->setTextColor(YELLOW);
  gfx->setCursor(BTN_X + 2, 124);
  gfx->print("CLEAR = camera");
  gfx->setCursor(BTN_X + 2, 136);
  gfx->print("REPLAY = again");

  drawWifi();
}

/* Word-wrapped answer overlay across the bottom of the viewfinder. */
void showAnswer(const String &text) {
  const int max_chars = 38;
  int lines = (text.length() + max_chars - 1) / max_chars;
  if (lines > 6) lines = 6;
  int h = lines * 10 + 8;
  int y = 240 - h;

  gfx->fillRect(0, y, 240, h, gfx->color565(0, 0, 0));
  gfx->drawRect(0, y, 240, h, CYAN);
  gfx->setTextSize(1);
  gfx->setTextColor(WHITE);
  for (int i = 0; i < lines; i++) {
    gfx->setCursor(4, y + 5 + i * 10);
    gfx->print(text.substring(i * max_chars, min((int)text.length(), (i + 1) * max_chars)));
  }
}


/* ===========================================================================
 *  One full "look and tell" cycle
 * =========================================================================== */
void lookAndTell(const char *prompt, const char *label) {
  led(255, 120, 0);                                        // amber: working
  preview_on = false;                                      // camera goes OFF here

  gfx->fillRect(0, 0, 240, 240, BLACK);
  gfx->setTextColor(YELLOW);
  gfx->setTextSize(1);
  gfx->setCursor(30, 110);
  gfx->printf("capturing photo (%s)...", label);

  size_t jpg_len = 0;
  uint32_t t_cap = millis();
  uint8_t *jpg = captureJpeg(&jpg_len);
  t_cap = millis() - t_cap;

  if (!jpg || jpg_len == 0) {
    if (jpg) free(jpg);
    showAnswer("Capture failed - tap CLEAR to retry.");
    led(255, 0, 0);
    drawClearSide();
    return;
  }
  Serial.printf("Captured %u KB in %lu ms\n", (unsigned)(jpg_len / 1024), (unsigned long)t_cap);

  gfx->setCursor(30, 124);
  gfx->printf("asking OpenAI (%uKB)...", (unsigned)(jpg_len / 1024));

  String answer;
  uint32_t t_ai = millis();
  bool ok = askVision(jpg, jpg_len, prompt, answer);
  t_ai = millis() - t_ai;
  free(jpg);

  if (!ok) {
    showAnswer("No answer - see serial monitor for the reason.");
    led(255, 0, 0);
    drawClearSide();
    return;
  }

  Serial.printf("Vision (%lu ms): %s\n", (unsigned long)t_ai, answer.c_str());
  showAnswer(answer);
  drawClearSide();                                         // CLEAR replaces the ask buttons

#if SPEAK_REPLIES
  led(0, 255, 40);                                         // green: speaking
  speak(answer);                                           // answer stays on screen
#endif
  led(0, 0, 0);
  /* the answer now stays until the user taps CLEAR */
}


/* ===========================================================================
 *  SETUP
 * =========================================================================== */
void setup() {
  Serial.begin(115200);
  delay(400);
  Serial.println("\n=== 05 Vision AI  |  Robojax.com ===");

  pinMode(TFT_BLK, OUTPUT);
  digitalWrite(TFT_BLK, LOW);
  pinMode(SD_CS, OUTPUT);
  digitalWrite(SD_CS, HIGH);

  gfx->begin();
  gfx->fillScreen(BLACK);
  digitalWrite(TFT_BLK, HIGH);

  bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
  delay(50);

  // The camera's SCCB shares this bus, so Wire must be up before the camera
  Wire.begin(I2C_SDA, I2C_SCL, 100000);
  delay(20);

  rgb.begin();
  rgb.setBrightness(LED_BRIGHTNESS);
  led(0, 0, 0);

#if SPEAK_REPLIES
  spkInit();
#endif

  gfx->setTextSize(1);
  gfx->setTextColor(YELLOW);
  gfx->setCursor(4, 4);
  gfx->printf("Connecting to %s ...", WIFI_SSID);

  WiFi.mode(WIFI_STA);
  WiFi.begin(WIFI_SSID, WIFI_PASS);
  uint32_t t0 = millis();
  while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) delay(300);

  gfx->setCursor(4, 16);
  if (WiFi.status() == WL_CONNECTED) {
    gfx->setTextColor(GREEN);
    gfx->print("WiFi ok");
  } else {
    gfx->setTextColor(RED);
    gfx->print("WiFi FAILED (2.4GHz only! check secrets.h)");
  }

  gfx->setCursor(4, 28);
  gfx->setTextColor(YELLOW);
  gfx->print("starting camera...");
  if (!camStartPreview()) {
    gfx->fillScreen(RED);
    gfx->setTextColor(WHITE);
    gfx->setTextSize(2);
    gfx->setCursor(20, 100);
    gfx->print("CAMERA FAILED");
    gfx->setTextSize(1);
    gfx->setCursor(20, 130);
    gfx->print("Press RESET (camera reset = board reset)");
    while (1) delay(1000);
  }

  delay(400);
  gfx->fillScreen(BLACK);
  drawButtons();
  preview_on = true;
  Serial.println("Running. Touch a button to have the AI look through the camera.");
}


/* ===========================================================================
 *  LOOP
 * =========================================================================== */
void loop() {
  static uint32_t last_touch = 0;

  if (preview_on) {
    camera_fb_t *fb = esp_camera_fb_get();
    if (fb) {
      gfx->draw16bitBeRGBBitmap(0, 0, (uint16_t *)fb->buf, fb->width, fb->height);
      esp_camera_fb_return(fb);
    }
  } else {
    delay(20);                               // answer on screen, waiting for CLEAR
  }

  /* live WiFi signal, refreshed every 2 s in both modes */
  static uint32_t last_wifi = 0;
  if (millis() - last_wifi > 2000) {
    last_wifi = millis();
    drawWifi();
  }

  /* Edge-detected touch: an action fires only on a NEW finger-down, and the
   * finger must fully lift (4 consecutive empty reads) before anything can
   * fire again. This is what stops one tap on CLEAR from also triggering the
   * capture button that appears in the same spot a moment later. */
  static bool    touch_down = false;
  static uint8_t release_count = 0;

  uint16_t x, y;
  if (getTouch(&x, &y)) {
    release_count = 0;
    if (!touch_down) {
      touch_down = true;                     // new tap - act exactly once

      if (!preview_on) {
        /* answer mode: CLEAR returns to the camera, REPLAY says it again */
        if (x >= BTN_X && y >= BTN_OBJ_Y && y < BTN_OBJ_Y + BTN_H) {
          if (camStartPreview()) {
            preview_on = true;
            drawButtons();
          } else {
            showAnswer("Camera restart failed - press RESET.");
          }
        }
#if SPEAK_REPLIES
        else if (x >= BTN_X && y >= BTN_PERSON_Y && y < BTN_PERSON_Y + BTN_H) {
          led(0, 255, 40);                   // green while speaking
          playLastAnswer();
          led(0, 0, 0);
        }
#endif
      } else if (x >= BTN_X && y >= BTN_OBJ_Y && y < BTN_OBJ_Y + BTN_H) {
        lookAndTell(PROMPT_OBJECTS, "objects");
      } else if (x >= BTN_X && y >= BTN_PERSON_Y && y < BTN_PERSON_Y + BTN_H) {
        lookAndTell(PROMPT_PERSON, "person");
      }
    }
  } else if (touch_down) {
    if (++release_count >= 4) { touch_down = false; release_count = 0; }
  }
  (void)last_touch;
}

Fișiere📁

Fișier necesar (.h)

  • secrets.h
    file for Makerfabs MaTouch AI ESP32S3 2.8" TFT Camera module
    secrets.h 0.01 MB

Alte fișiere

  • pins.h
    pins file for MaTouch AI ESP32S3 2.8" camera LCD touch screen.
    pins.h 0.01 MB

Schematic

  • MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematic
    The latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
    MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF 0.15 MB