This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
The latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
Makerfabs MaTouch ESP32-S3 2.8" camera ESP32-S3 AI Vision: Hướng camera và hỏi nó thấy gì
Boards chụp một bức ảnh, AI mô tả những gì nó nhìn thấy, và nó đọc to câu trả lời
Hướng camera vào một vật gì đó, nhấn một nút, và AI sẽ cho bạn biết nó đang nhìn thấy gì - được in trên màn hình và phát qua loa. Hai chế độ: gọi tên các vật thể trong tầm nhìn, hoặc mô tả người trước camera.
AI Vision - Point It At Something chạy trên bo mạch MaTouch AI ESP32-S3
Cách hình ảnh được truyền đi
Camera tạo ra một tệp JPEG 800×600 khoảng 20-40 KB. Bo mạch mã hóa base64 tệp này trong PSRAM, khiến kích thước tăng khoảng một phần ba dưới dạng văn bản, và nhúng nó vào yêu cầu dưới dạng một data URI. 8 MB PSRAM chính là thứ giúp việc này trở nên dễ dàng.
Mô tả, không bao giờ nhận dạng
Chế độ PERSON mô tả những gì nó có thể thấy - tâm trạng, kính mắt, người đó có vẻ đang làm gì. Nó sẽ không cho bạn biết một người là ai, và điều đó là có chủ đích. Nhận dạng người qua khuôn mặt vi phạm chính sách sử dụng của OpenAI, và dịch vụ nhận dạng khuôn mặt tương đương của Microsoft bị khóa sau một quy trình phê duyệt mà các nhà phát triển thông thường không thể đơn giản đăng ký. Bất kỳ hướng dẫn nào hứa hẹn nhận dạng khuôn mặt qua đám mây đều đang mô tả một thứ không phổ biến.
Nếu bạn muốn một bo mạch nhận dạng những người cụ thể, hãy sử dụng dự án 03b - nó hoạt động ngoại tuyến và không bao giờ gửi khuôn mặt của bất kỳ ai đi đâu cả.
Camera phải dừng để mạng hoạt động
Kính ngắm trực tiếp bị tắt trong khi câu hỏi đang được gửi đi. Điều này không chỉ mang tính thẩm mỹ: trình điều khiển camera đang chạy giữ bộ nhớ nội bộ mà kết nối bảo mật cần, và khi chế độ xem trước đang chạy, kết nối sẽ thất bại. Tắt camera sẽ giải phóng bộ nhớ đó. Điều này cũng có nghĩa là câu trả lời có thể đọc được thay vì bị vẽ lại hơn 25 lần mỗi giây.
Câu trả lời ở lại cho đến khi bạn xóa nó
Câu trả lời vẫn hiển thị trên màn hình cho đến khi bạn chạm vào CLEAR - không có bộ hẹn giờ. Trong khi nó đang hiển thị, một nút REPLAY phát lại câu trả lời bằng giọng nói từ bộ nhớ, không cần gọi API lần thứ hai và không tốn thêm chi phí.
Về bo mạch MaTouch AI ESP32-S3 2.8"
Mọi dự án trên trang này chạy trên MaTouch AI ESP32-S3 2.8" TFT ST7789V từ Makerfabs. Đây là một bo mạch tất cả-trong-một: màn hình cảm ứng màu, camera 3 megapixel, hai micro và một bộ khuếch đại loa thực sự, tất cả được điều khiển bởi ESP32-S3 với 8 MB PSRAM. Sự kết hợp đó chính là điều khiến các dự án AI này khả thi trên một bo mạch duy nhất mà không cần thêm thiết bị nào khác.
8 MB PSRAM quan trọng hơn bất kỳ con số nào khác ở đây. Nó cho phép bo mạch giữ một khung hình camera, vài giây âm thanh đã ghi, hoặc một bức ảnh mã hóa base64 trong bộ nhớ cùng một lúc - không thứ nào trong số đó vừa với RAM thông thường của ESP32.
Tài liệu của nhà sản xuất: Trang wiki Makerfabs.
Thông số kỹ thuật chính
Bộ xử lý: ESP32-S3, lõi kép 240 MHz, WiFi 2.4 GHz + Bluetooth 5.0
Bộ nhớ: 16 MB flash, 8 MB PSRAM (bắt buộc cho gần như mọi dự án ở đây)
Màn hình: 2.8" IPS, 320×240, trình điều khiển ST7789V, SPI
Cảm ứng: GT911 điện dung, theo dõi 5 ngón tay cùng lúc
Camera: OV3660, 3 megapixel, lên đến 2048×1536
Micro: hai micro kỹ thuật số I2S INMP441 (một cặp stereo thực sự)
Loa: bộ khuếch đại class-D MAX98357A, 3.2 W vào 4 Ω
Lưu trữ: khe thẻ microSD (chế độ SPI)
Nguồn: USB-C, đầu nối pin JST, bộ sạc TP4056, công tắc nguồn
Cũng có trên bo: đèn LED RGB WS2812B, đồng hồ thời gian thực PCF8563T chạy bằng pin, và một bộ đo pin MAX17048 không được liệt kê trong thông số kỹ thuật chính thức
Hai cổng USB-C không giống nhau. Loa của bo mạch chia sẻ các chân tín hiệu (IO19 và IO20) với cổng USB gốc, vì các chân đó là các đường dữ liệu USB được nối cứng của ESP32-S3. Luôn tải lên và cấp nguồn qua cổng USB-C CH340K (cổng bên cạnh nút RESET), và đặt USB CDC On Boot thành Disabled. Dùng sai cổng sẽ khiến âm thanh hoạt động sai hoặc quá trình tải lên thất bại.
Cài đặt Arduino IDE
Những cài đặt này rất quan trọng. Hầu hết các vấn đề mà mọi người gặp phải với bo mạch này đều do một trong những cài đặt này sai, và chúng tự đặt lại khi bạn thay đổi phiên bản core, vì vậy hãy kiểm tra lại chúng sau bất kỳ thay đổi nào.
Cài đặt | Giá trị |
|---|---|
Bo mạch | ESP32S3 Dev Module |
Phiên bản core ESP32 | 2.0.17 |
PSRAM | OPI PSRAM |
Kích thước Flash | 16MB (128Mb) |
Sơ đồ phân vùng | 16M Flash (3MB APP/9.9MB FATFS) |
USB CDC On Boot | Disabled |
Tốc độ tải lên | 921600 |
Xóa toàn bộ Flash trước khi tải lên | Disabled |
Cổng | cổng USB-C CH340K |
Sử dụng core ESP32 2.0.17, không phải 3.x. Espressif đã loại bỏ các mô hình nhận dạng khuôn mặt trên thiết bị trong core 3, vì vậy các dự án nhận dạng khuôn mặt sẽ không biên dịch được ở đó. Cố định 2.0.17 giúp mọi dự án trên trang này hoạt động với một cấu hình duy nhất. Trong Boards Manager, menu thả xuống phiên bản cho phép bạn chuyển đổi qua lại bất cứ khi nào bạn muốn.
Sử dụng GFX Library for Arduino phiên bản 1.5.6, không phải 1.6.x. Các bản phát hành 1.6 được xây dựng cho ESP32 core 3 và có thể treo khi khởi động trên core 2.0.17. Nếu màn hình của bạn vẫn đen sau khi tải lên, đây là điều đầu tiên cần kiểm tra.
Thư viện cần thiết
Cài đặt chúng qua Tools → Manage Libraries trong Arduino IDE. Số phiên bản rất quan trọng - vui lòng sử dụng các phiên bản được liệt kê.
Thư viện | Phiên bản | Tác giả |
|---|---|---|
GFX Library for Arduino | 1.5.6 | moononournation |
bb_captouch | 1.3.1 | Larry Bank |
ArduinoJson | 7.x | Benoit Blanchon |
Adafruit NeoPixel | bất kỳ bản mới nào | Adafruit |
Thiết lập secrets.h
Thông tin WiFi của bạn và bất kỳ khóa API nào được đặt trong secrets.h, tệp này được bao gồm trong bản tải xuống với các giá trị mẫu. Mở tab đó trong Arduino IDE và thay thế chúng bằng thông tin của bạn.
WiFi phải là 2.4 GHz. ESP32-S3 không thể nhìn thấy mạng 5 GHz. Nếu bộ định tuyến của bạn kết hợp cả hai băng tần dưới một tên (Asus gọi đây là Smart Connect), hãy tắt tính năng đó hoặc đặt cho băng tần 2.4 GHz một tên riêng và sử dụng tên đó trong secrets.h.
Lấy khóa API của bạn
Dự án này kết nối với dịch vụ AI đám mây, vì vậy bạn cần khóa riêng của mình. Nếu bạn chưa từng làm điều này trước đây, đừng lo lắng - nó giống như mật khẩu xác định tài khoản của bạn với dịch vụ. Chỉ mất vài phút, một lần duy nhất.
Khóa không phải là đăng ký trang web. Ví dụ, trả phí cho ChatGPT Plus không cung cấp cho bạn khóa API - hai sản phẩm này tách biệt với hóa đơn riêng. Bạn cần tài khoản trên nền tảng nhà phát triển, được mô tả bên dưới.
OpenAI - đôi mắt
OpenAI cung cấp mô hình thị giác xem hình ảnh. Chỉ các dự án gửi hình ảnh mới cần cái này.
Truy cập platform.openai.com/api-keys và đăng nhập hoặc tạo tài khoản.
Nhấp vào Create new secret key, đặt tên cho nó và tạo.
Sao chép ngay lập tức - giống như DeepSeek, nó chỉ được hiển thị một lần.
Mở Billing và thêm một khoản tín dụng nhỏ. API trả trước và tách biệt với bất kỳ đăng ký ChatGPT nào bạn có thể đã có.
Đặt khóa vào secrets.h dưới dạng OPENAI_KEY.
Microsoft Azure Speech - để nghe và nói
Azure chuyển giọng nói của bạn thành văn bản và chuyển câu trả lời thành giọng nói. Gói miễn phí đủ hào phóng cho mọi thứ trên trang này.
Truy cập portal.azure.com và đăng nhập bằng tài khoản Microsoft (tài khoản miễn phí là đủ).
Nếu bạn chưa từng sử dụng Azure trước đây, bạn sẽ thấy màn hình Welcome to Azure với ba lựa chọn. Chọn Start with an Azure free trial - bạn cần đăng ký trước khi Azure cho phép tạo bất cứ thứ gì. (Sinh viên nên chọn Azure for Students thay thế: kết quả tương tự, không cần thẻ.) Bỏ qua Manage Microsoft Entra ID, một thứ hoàn toàn khác.
Nhấp vào Create a resource, tìm kiếm Speech, và chọn Speech service do Microsoft phát hành.
Điền vào biểu mẫu: bất kỳ nhóm tài nguyên nào, bất kỳ tên nào, và chọn Region gần bạn - ghi lại khu vực đó chính xác như hiển thị, ví dụ
eastus.Đối với Pricing tier, chọn F0 (Free). Điều này cho phép khoảng năm giờ chuyển giọng nói thành văn bản và nửa triệu ký tự chuyển văn bản thành giọng nói mỗi tháng.
Nhấp vào Review + create, sau đó Create. Đợi khoảng một phút, sau đó nhấp vào Go to resource.
Trong menu bên trái, mở Keys and Endpoint. Sao chép KEY 1 và Location/Region.
Đặt chúng vào secrets.h dưới dạng AZURE_SPEECH_KEY và AZURE_REGION. Đối với AZURE_STT_HOST, sử dụng <region>.stt.speech.microsoft.com - vì vậy với khu vực eastus, đó là eastus.stt.speech.microsoft.com.
Về thẻ tín dụng. Bản dùng thử miễn phí của Azure yêu cầu thẻ để xác minh danh tính của bạn. Nó không tính phí bạn. Bạn nhận được $200 tín dụng trong 30 ngày, và sau đó tài khoản chuyển sang Pay-As-You-Go - nhưng F0 Speech tier vẫn miễn phí, tháng này qua tháng khác, và mọi thứ trong các dự án này nằm gọn trong đó. Nếu bạn không muốn cung cấp thẻ và bạn là sinh viên, tùy chọn Azure for Students cung cấp tín dụng mà không cần thẻ.
Nó phải là tài nguyên "Speech service". Khóa từ tài nguyên Translator, Language hoặc Cognitive Services chung trông giống hệt và hoàn toàn hợp lệ - nhưng mọi yêu cầu giọng nói đều trả về lỗi 401. Điều này đã làm chúng tôi mất một giờ trong quá trình thử nghiệm. Nếu giọng nói thất bại với mã 401 trong khi khóa trông có vẻ đúng, hãy kiểm tra loại tài nguyên bạn đã tạo.
Chi phí vận hành là bao nhiêu
Rất ít, nhưng không miễn phí, và bạn nên biết gần đúng số tiền bạn đang chi trước khi để một dự án chạy.
Dịch vụ | Chi phí ước tính |
|---|---|
Azure Speech | gói miễn phí bao gồm khoảng 5 giờ nghe và 0,5 triệu ký tự nói mỗi tháng |
DeepSeek | một phần nhỏ của xu cho mỗi câu trả lời - hàng nghìn phản hồi chỉ với vài đô la |
OpenAI vision | khoảng một hoặc hai xu cho mỗi hình ảnh, tùy thuộc vào mô hình |
Giá cả thay đổi, vì vậy hãy coi đây là hướng dẫn thay vì báo giá. Mỗi dịch vụ này đều có trang sử dụng để bạn theo dõi số tiền đã chi, và tất cả đều cho phép bạn đặt giới hạn chi tiêu - điều đáng làm ngay từ ngày đầu.
Giữ khóa của bạn ở chế độ riêng tư. Bất kỳ ai có chúng đều có thể tiêu tiền của bạn. Đừng đặt chúng trong video, ảnh chụp màn hình, bài đăng diễn đàn hoặc kho mã nguồn công khai. Nếu khóa bị lộ, hãy xóa nó trên trang web của nhà cung cấp và tạo khóa mới - việc này chỉ mất vài giây và là cách khắc phục thực sự duy nhất.
Xử lý sự cố
Triệu chứng | Nguyên nhân và cách khắc phục |
|---|---|
Màn hình vẫn đen | Sai phiên bản thư viện GFX (sử dụng 1.5.6) hoặc sai cài đặt bo mạch. |
|
|
Không tải lên được / không có cổng COM | Sai cổng USB-C, hoặc trình điều khiển CH340 chưa được cài đặt. |
Camera lỗi và không bao giờ phục hồi | Chân reset của camera được nối với nút RESET của bo mạch, vì vậy phần mềm không thể khởi động lại nó. Nhấn RESET. Nếu vẫn lỗi, hãy cắm lại cáp ribbon của camera. |
Tải mã nguồn
Bản phác thảo Arduino hoàn chỉnh cho dự án này, cùng với pins.h và mọi thứ khác cần thiết, có thể tải xuống miễn phí.
Giải nén, mở tệp .ino trong Arduino IDE, kiểm tra các cài đặt ở trên và tải lên qua cổng CH340K USB-C.
This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
- Makerfabs MaTouch ESP32-S3 2.8" Camera và Màn hình cảm ứng: Ảnh 3MP lưu vào thẻ SD
- Makerfabs MaTouch ESP32-S3 2.8" camera Phát hiện khuôn mặt ngoại tuyến trên ESP32-S3 - Không cần Internet, Không cần khóa API
- Makerfabs MaTouch ESP32-S3 2.8" camera Xây dựng trợ lý giọng nói AI trên ESP32-S3 (Azure + DeepSeek)
/*
* ===========================================================================
* 05_Vision_AI — MaTouch AI ESP32-S3 2.8" TFT ST7789V
* ===========================================================================
*
----------
* ROBOJAX.COM - MaTouch AI ESP32-S3 2.8" project series
*
* WATCH THE VIDEO
* https://youtu.be/6AL3g3tC_Hk
*
* WRITTEN TUTORIALS - every project, with photos and full explanation
* Camera and touchscreen.... https://robojax.com/RTJ849
* Offline face recognition.. https://robojax.com/RTJ850
* AI voice assistant........ https://robojax.com/RTJ851
* AI vision................. https://robojax.com/RTJ852
*
* GET THE BOARD - SAVE $5 with coupon code: Robojax_Makerfab
* https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
* (enter the code at checkout)
*
* All of this code is free. If it helped you, a subscribe on YouTube is
* the best way to support more of it.
*
*
* Watching the video first will save you time - it shows the Arduino IDE
* settings and the library versions being set up step by step.
* ---------------------------------------------------------------------------
*
* The board SEES. Point the camera at something, touch a button, and an
* OpenAI vision model tells you what it is looking at - drawn on the screen
* and (optionally) spoken out of the board's own speaker via Azure TTS.
*
* Two modes, two buttons:
*
* OBJECTS "What do you see?" - names the things in front of the lens.
* PERSON Describes the person in frame: glasses, expression, what
* they are doing. DESCRIPTION ONLY - it will not and must not
* try to say WHO someone is. Identifying people by face is
* against OpenAI's usage policies, and it is the right call:
* say this in the video, it is worth 15 honest seconds.
*
* WHY OPENAI FOR THIS DEMO AND NOT DEEPSEEK: DeepSeek's API is text-only.
* It cannot accept an image at all. Azure OpenAI could do it, but plain
* OpenAI is one endpoint with no deployment setup - simplest to follow.
*
* HOW THE IMAGE TRAVELS: the OV3660 gives us a JPEG directly (800x600,
* ~40 KB). We base64-encode it in PSRAM (~55 KB of text) and embed it in
* the JSON request as a data: URI. The 8 MB PSRAM makes this trivial.
*
* ---------------------------------------------------------------------------
* FILL IN secrets.h BEFORE FLASHING
* (needs WIFI_*, OPENAI_*; AZURE_* only if SPEAK_REPLIES is 1).
*
* BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
* or compile errors. These reset when you switch cores - recheck them!)
*
* Board : ESP32S3 Dev Module
* ESP32 core : 2.0.17
* PSRAM : OPI PSRAM <-- required, image buffers live there
* Flash Size : 16MB (128Mb)
* Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
* USB CDC On Boot : Disabled <-- speaker shares pins with native USB
* Upload Speed : 921600
* Port : the CH340K USB-C port (the one near RESET)
*
* LIBRARIES
* GFX Library for Arduino v1.5.6 (NOT 1.6.x - that pairs with core 3)
* bb_captouch v1.3.1
* ArduinoJson v7.x
* Adafruit NeoPixel any recent
*
* ---------------------------------------------------------------------------
* FUNCTIONS IN THIS SKETCH
* led(r,g,b) set the RGB status LED colour
* getTouch(&x,&y) read the touch panel, mapped to screen coordinates
* camStart(...) start the camera in a given format/size
* camStartPreview() start the 240x240 RGB565 live-view camera
* captureJpeg(&len) take one 800x600 JPEG into PSRAM (camera stays OFF
* afterwards - caller restarts the preview)
* readHttpResponse() read an HTTPS reply, de-chunking it properly
* askVision(...) send photo + prompt to OpenAI, return the answer
* spkInit() configure the I2S speaker output
* readExact(...) read exactly N bytes from a TLS connection
* speak(text) Azure TTS -> download voice to PSRAM -> play it
* playLastAnswer() replay the kept voice from PSRAM (REPLAY button)
* drawButton(...) draw one side-column button
* drawWifi() WiFi signal bars + dBm readout
* drawButtons() normal side column (the two ask buttons)
* drawClearSide() answer-mode side column (CLEAR + REPLAY)
* showAnswer(text) word-wrapped answer overlay on the viewfinder
* lookAndTell(...) one full cycle: capture -> ask -> show -> speak
* setup() / loop() boot sequence / viewfinder + touch handling
*
* Robojax.com
* ===========================================================================
*/
#define SPEAK_REPLIES 1 // 1 = read the answer aloud with Azure TTS, 0 = screen only
/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
* uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS 15
/* CAMERA ORIENTATION
* 1 = camera faces the SAME way as the screen (how the board ships) - the
* image is mirrored so it looks natural when you point it at yourself.
* 0 = you folded the ribbon so the lens faces AWAY from the screen
* (phone-style, screen to you / camera to the subject).
* If the picture looks left-right reversed, flip this number. */
#define CAMERA_FACES_USER 1
#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <Wire.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include "mbedtls/base64.h"
#include "esp_camera.h"
#include "pins.h"
#include "secrets.h"
#if SPEAK_REPLIES
#include "driver/i2s.h"
#define WAV_HEADER_LEN 44
#endif
Arduino_ESP32SPI *bus = new Arduino_ESP32SPI(
TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, HSPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);
BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);
/* --- the two vision prompts ------------------------------------------------
* Keep answers short: they must fit a 320x240 screen and, if spoken, must not
* leave the presenter waiting awkwardly on camera. */
const char *PROMPT_OBJECTS =
"Look at this photo from a small camera. In ONE short sentence of at most "
"20 words, plain text only, name the main object(s) you see.";
const char *PROMPT_PERSON =
"Look at this photo. If there is a person, describe them in ONE short "
"sentence of at most 20 words: mood, glasses or not, what they are doing. "
"Never guess who they are. If no person is visible, say so. Plain text only.";
/* --- layout: viewfinder left, buttons right ------------------------------- */
#define BTN_X 242
#define BTN_W 78
#define BTN_H 52
#define BTN_OBJ_Y 4
#define BTN_PERSON_Y 62
void led(uint8_t r, uint8_t g, uint8_t b) { rgb.setPixelColor(0, rgb.Color(r, g, b)); rgb.show(); }
/* ===========================================================================
* Touch
* =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
TOUCHINFO ti;
if (!bbct.getSamples(&ti)) return false;
if (ti.count < 1) return false;
*x = ti.y[0];
*y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
return true;
}
/* ===========================================================================
* Camera — RGB565 for the live view; JPEG capture happens by restarting
* the driver, same technique as sketch 02.
* =========================================================================== */
bool mirrored = CAMERA_FACES_USER; // see the define at the top of the file
/* The viewfinder is OFF while a cloud call runs and while the answer is on
* screen. Two reasons, both learned on the bench:
* 1. RAM: the live camera driver eats the internal memory a TLS handshake
* needs - with the preview running, connecting to OpenAI fails (HTTP -1).
* 2. Readability: the viewfinder repaints 25x/s and would wipe the answer.
* The answer STAYS on screen until the user taps CLEAR - no timer. */
bool preview_on = false;
/* The last spoken answer is KEPT in PSRAM so the REPLAY button can play it
* again without another cloud call. Overwritten by the next answer. */
uint8_t *last_audio = nullptr;
size_t last_audio_len = 0;
bool camStart(pixformat_t fmt, framesize_t size, int fb_count, int quality) {
camera_config_t c;
c.ledc_channel = LEDC_CHANNEL_0;
c.ledc_timer = LEDC_TIMER_0;
c.pin_d0 = CAM_PIN_D0; c.pin_d1 = CAM_PIN_D1;
c.pin_d2 = CAM_PIN_D2; c.pin_d3 = CAM_PIN_D3;
c.pin_d4 = CAM_PIN_D4; c.pin_d5 = CAM_PIN_D5;
c.pin_d6 = CAM_PIN_D6; c.pin_d7 = CAM_PIN_D7;
c.pin_xclk = CAM_PIN_XCLK;
c.pin_pclk = CAM_PIN_PCLK;
c.pin_vsync = CAM_PIN_VSYNC;
c.pin_href = CAM_PIN_HREF;
/* Share the I2C bus Wire already drives (the touch panel lives there too)
* instead of letting the camera install a second driver on the same pins -
* that kills touch. Wire.begin() must run before this function. */
c.pin_sccb_sda = -1;
c.pin_sccb_scl = -1;
c.sccb_i2c_port = 0; // Wire = I2C port 0
c.pin_pwdn = CAM_PIN_PWDN;
c.pin_reset = CAM_PIN_RESET;
c.xclk_freq_hz = 20000000;
c.frame_size = size;
c.pixel_format = fmt;
c.grab_mode = CAMERA_GRAB_WHEN_EMPTY;
c.fb_location = CAMERA_FB_IN_PSRAM;
c.jpeg_quality = quality;
c.fb_count = fb_count;
if (esp_camera_init(&c) != ESP_OK) return false;
sensor_t *s = esp_camera_sensor_get();
if (s) {
s->set_hmirror(s, mirrored ? 1 : 0);
s->set_vflip(s, mirrored ? 1 : 0);
s->set_brightness(s, 1);
}
return true;
}
bool camStartPreview() { return camStart(PIXFORMAT_RGB565, FRAMESIZE_240X240, 2, 12); }
/* Capture one SVGA JPEG into a PSRAM buffer the caller owns. */
uint8_t *captureJpeg(size_t *len_out) {
esp_camera_deinit();
delay(120);
if (!camStart(PIXFORMAT_JPEG, FRAMESIZE_SVGA, 1, 12)) { *len_out = 0; return nullptr; }
// a few warm-up frames so exposure settles
for (int i = 0; i < 3; i++) {
camera_fb_t *w = esp_camera_fb_get();
if (w) esp_camera_fb_return(w);
delay(100);
}
uint8_t *copy = nullptr;
*len_out = 0;
camera_fb_t *fb = esp_camera_fb_get();
if (fb && fb->len > 0) {
copy = (uint8_t *)ps_malloc(fb->len);
if (copy) { memcpy(copy, fb->buf, fb->len); *len_out = fb->len; }
}
if (fb) esp_camera_fb_return(fb);
/* Deliberately leave the camera OFF here. The caller restarts the preview
* after the cloud call - a running camera driver starves the TLS handshake
* of internal RAM (that was the "Vision HTTP -1" bug). */
esp_camera_deinit();
delay(120);
return copy;
}
/* ===========================================================================
* HTTP response reader — shared by the cloud calls below.
* Returns the status code and fills body_out. Handles chunked transfer
* encoding PROPERLY: the chunk-size markers must be stripped, or they end
* up embedded inside the JSON and the parse fails. (OpenAI pretty-prints
* its replies so they span several chunks - this bit us on the bench.)
* =========================================================================== */
/* Block until the connection has data (or the budget runs out). Every read
* below goes through this, because the server can go quiet for many seconds
* while it analyses the image - and a bare read() would simply time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
uint32_t t0 = millis();
while (!c.available()) {
if (!c.connected()) return false;
if (millis() - t0 > ms) return false;
delay(10);
}
return true;
}
static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
body_out = "";
if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
while (waitData(client, idle_ms)) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break; // blank line = end of headers
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
}
if (chunked) {
int blanks = 0;
while (true) {
/* Waiting here - instead of letting read() time out - is essential: a
* timed-out read looks exactly like "0" (final chunk), which silently
* truncates the body to nothing while the server is still thinking. */
if (!waitData(client, idle_ms)) {
Serial.println("HTTP: timed out waiting for the next chunk");
break;
}
String szline = client.readStringUntil('\n');
szline.trim();
if (szline.length() == 0) {
if (++blanks > 4) break;
continue;
}
blanks = 0;
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break; // genuine final chunk
long got = 0;
while (got < sz) {
if (!waitData(client, idle_ms)) break;
while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
}
if (waitData(client, 3000)) client.readStringUntil('\n');
if (got < sz) { Serial.println("HTTP: short chunk"); break; }
}
} else {
while (waitData(client, idle_ms))
while (client.available()) body_out += (char)client.read();
}
return code;
}
/* ===========================================================================
* OpenAI vision call
* The request body is built by hand in PSRAM rather than through ArduinoJson,
* because embedding a 55 KB base64 string in a JSON document would mean
* holding two copies. Base64 text never needs JSON escaping, so this is safe.
* =========================================================================== */
bool askVision(const uint8_t *jpg, size_t jpg_len, const char *prompt, String &answer_out) {
// --- base64 encode the image into PSRAM ---
size_t b64_cap = ((jpg_len + 2) / 3) * 4 + 16;
unsigned char *b64 = (unsigned char *)ps_malloc(b64_cap);
if (!b64) return false;
size_t b64_len = 0;
if (mbedtls_base64_encode(b64, b64_cap, &b64_len, jpg, jpg_len) != 0) {
free(b64);
return false;
}
// --- assemble the JSON request around it ---
/* Current OpenAI models reject the old "max_tokens" name - it must be
* "max_completion_tokens" (verified against the live API, July 2026). */
const char *head_fmt =
"{\"model\":\"%s\",\"max_completion_tokens\":%d,\"messages\":[{\"role\":\"user\","
"\"content\":[{\"type\":\"text\",\"text\":\"%s\"},"
"{\"type\":\"image_url\",\"image_url\":{\"url\":\"data:image/jpeg;base64,";
const char *tail = "\"}}]}]}";
char head[640];
int head_len = snprintf(head, sizeof(head), head_fmt, OPENAI_MODEL, LLM_MAX_TOKENS, prompt);
size_t body_len = head_len + b64_len + strlen(tail);
char *body = (char *)ps_malloc(body_len + 1);
if (!body) { free(b64); return false; }
memcpy(body, head, head_len);
memcpy(body + head_len, b64, b64_len);
strcpy(body + head_len + b64_len, tail);
free(b64);
// --- send it: manual HTTP, chunk-wise upload (proven pattern from 04) ---
WiFiClientSecure client;
client.setInsecure();
client.setTimeout(25);
if (!client.connect(OPENAI_HOST, 443)) {
Serial.println("Vision: TLS connect failed (is the camera still running?)");
free(body);
return false;
}
client.print(String("POST /v1/chat/completions HTTP/1.1\r\n"
"Host: " OPENAI_HOST "\r\n"
"Authorization: Bearer " OPENAI_KEY "\r\n"
"Content-Type: application/json\r\n"
"Connection: close\r\n"
"Content-Length: ") + String(body_len) + "\r\n\r\n");
size_t sent = 0;
while (sent < body_len) {
size_t n = min((size_t)4096, body_len - sent);
size_t w = client.write((uint8_t *)body + sent, n);
if (w == 0) {
delay(50);
w = client.write((uint8_t *)body + sent, n);
if (w == 0) break;
}
sent += w;
yield();
}
free(body);
if (sent < body_len) {
Serial.printf("Vision: upload stalled at %u/%u bytes\n",
(unsigned)sent, (unsigned)body_len);
client.stop();
return false;
}
/* read the reply with proper de-chunking */
String resp;
int code = readHttpResponse(client, resp, 30000);
client.stop();
if (code != 200) {
Serial.printf("Vision HTTP %d: %s\n", code, resp.c_str());
return false;
}
bool ok = false;
JsonDocument doc;
if (!deserializeJson(doc, resp)) {
const char *content = doc["choices"][0]["message"]["content"];
if (content) {
answer_out = String(content);
answer_out.trim();
ok = answer_out.length() > 0;
}
} else {
Serial.printf("Vision: JSON parse failed, %u bytes received\n", resp.length());
}
return ok;
}
/* ===========================================================================
* Azure TTS — same streaming trick as sketch 04: raw PCM into I2S.
* =========================================================================== */
#if SPEAK_REPLIES
void spkInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
.sample_rate = 16000,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = true,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_SPK_BCLK,
.ws_io_num = I2S_SPK_LRC,
.data_out_num = I2S_SPK_DOUT,
.data_in_num = I2S_PIN_NO_CHANGE
};
i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_SPK_PORT, &pins);
i2s_zero_dma_buffer(I2S_SPK_PORT);
}
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
size_t got = 0;
uint32_t t0 = millis();
while (got < n && millis() - t0 < 10000) {
int r = c.read(dst + got, n - got);
if (r > 0) { got += r; t0 = millis(); }
else if (!c.connected() && !c.available()) break;
else delay(2);
}
return got;
}
void playLastAnswer(); // defined below; explicit prototype for the IDE
void speak(const String &text) {
String safe = text;
safe.replace("&", "&");
safe.replace("<", "<");
safe.replace(">", ">");
String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
"<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";
/* Manual HTTP with proper de-chunking: Azure sends this audio chunked, and
* HTTPClient's raw stream leaks the ASCII chunk-size lines into the PCM -
* each one plays as an audible KNOCK. */
WiFiClientSecure client;
client.setInsecure();
client.setTimeout(20);
const char *host = AZURE_REGION ".tts.speech.microsoft.com";
if (!client.connect(host, 443)) { Serial.println("TTS: TLS connect failed"); return; }
client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
"Host: ") + host + "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: application/ssml+xml\r\n"
"X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
"User-Agent: MaTouchRobojax\r\n"
"Connection: close\r\n"
"Content-Length: " + String(ssml.length()) + "\r\n\r\n");
client.print(ssml);
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
long content_len = -1;
while (client.connected() || client.available()) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break;
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
}
if (code != 200) { Serial.printf("TTS HTTP %d\n", code); client.stop(); return; }
const size_t AUDIO_CAP = 1200 * 1024;
uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
if (!audio) { client.stop(); return; }
size_t alen = 0;
if (chunked) {
while (true) {
String szline = client.readStringUntil('\n');
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break;
if (alen + sz > AUDIO_CAP) break;
size_t got = readExact(client, audio + alen, sz);
alen += got;
client.readStringUntil('\n');
if (got < (size_t)sz) break;
}
} else if (content_len > 0) {
alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
} else {
uint32_t idle = millis();
while ((client.connected() || client.available()) && millis() - idle < 5000) {
int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
if (r > 0) { alen += r; idle = millis(); }
else delay(5);
}
}
client.stop();
if (alen > WAV_HEADER_LEN) {
/* keep this answer for the REPLAY button (replacing the previous one),
* then play it */
if (last_audio) free(last_audio);
last_audio = audio;
last_audio_len = alen;
playLastAnswer();
} else {
free(audio);
}
}
/* Play the kept answer from PSRAM - used right after download AND by REPLAY. */
void playLastAnswer() {
if (!last_audio || last_audio_len <= WAV_HEADER_LEN) return;
static const uint8_t lead_in[640] = {0}; // 20 ms silence pre-roll
size_t w = 0;
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, last_audio + WAV_HEADER_LEN,
last_audio_len - WAV_HEADER_LEN, &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
delay(150);
i2s_zero_dma_buffer(I2S_SPK_PORT);
}
#endif // SPEAK_REPLIES
/* ===========================================================================
* UI
* =========================================================================== */
void drawButton(int y, const char *l1, const char *l2, uint16_t colour) {
gfx->fillRoundRect(BTN_X, y, BTN_W, BTN_H, 6, colour);
gfx->drawRoundRect(BTN_X, y, BTN_W, BTN_H, 6, WHITE);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_X + 8, y + 14);
gfx->print(l1);
gfx->setCursor(BTN_X + 8, y + 28);
gfx->print(l2);
}
/* WiFi bars + dBm at the bottom of the side column, refreshed from the loop */
void drawWifi() {
gfx->fillRect(BTN_X, 188, BTN_W, 52, BLACK);
bool up = (WiFi.status() == WL_CONNECTED);
long rssi = up ? WiFi.RSSI() : -100;
int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;
for (int b = 0; b < 4; b++) {
int bh = 6 + b * 6;
uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
gfx->fillRect(BTN_X + 4 + b * 9, 216 - bh, 7, bh, col);
}
gfx->setTextSize(1);
gfx->setCursor(BTN_X + 44, 196);
if (up) {
gfx->setTextColor(CYAN);
gfx->printf("%ld", rssi);
} else {
gfx->setTextColor(RED);
gfx->print("DOWN");
}
gfx->setCursor(BTN_X + 44, 208);
gfx->setTextColor(gfx->color565(120, 120, 120));
gfx->print("dBm");
gfx->setCursor(BTN_X + 4, 228);
gfx->setTextColor(gfx->color565(120, 120, 120));
gfx->print("WiFi signal");
}
/* normal side column: the two ask buttons */
void drawButtons() {
gfx->fillRect(BTN_X, 0, BTN_W, 188, BLACK);
drawButton(BTN_OBJ_Y, "WHAT DO", "YOU SEE?", gfx->color565(0, 90, 160));
drawButton(BTN_PERSON_Y, "DESCRIBE", "PERSON", gfx->color565(120, 60, 140));
gfx->setTextColor(CYAN);
gfx->setCursor(BTN_X + 2, 128);
gfx->print("OpenAI eyes");
gfx->setCursor(BTN_X + 2, 140);
gfx->print("Azure voice");
gfx->setCursor(BTN_X + 2, 152);
gfx->print("Robojax.com");
drawWifi();
}
/* answer-mode side column: CLEAR (back to camera) + REPLAY (say it again).
* The answer stays on screen until CLEAR is tapped. */
void drawClearSide() {
gfx->fillRect(BTN_X, 0, BTN_W, 188, BLACK);
gfx->fillRoundRect(BTN_X, BTN_OBJ_Y, BTN_W, BTN_H, 6, gfx->color565(110, 35, 35));
gfx->drawRoundRect(BTN_X, BTN_OBJ_Y, BTN_W, BTN_H, 6, WHITE);
gfx->setTextSize(2);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_X + 10, BTN_OBJ_Y + 14);
gfx->print("CLEAR");
#if SPEAK_REPLIES
if (last_audio_len > 0) {
gfx->fillRoundRect(BTN_X, BTN_PERSON_Y, BTN_W, BTN_H, 6, gfx->color565(0, 110, 60));
gfx->drawRoundRect(BTN_X, BTN_PERSON_Y, BTN_W, BTN_H, 6, WHITE);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_X + 16, BTN_PERSON_Y + 20);
gfx->print("REPLAY");
}
#endif
gfx->setTextSize(1);
gfx->setTextColor(YELLOW);
gfx->setCursor(BTN_X + 2, 124);
gfx->print("CLEAR = camera");
gfx->setCursor(BTN_X + 2, 136);
gfx->print("REPLAY = again");
drawWifi();
}
/* Word-wrapped answer overlay across the bottom of the viewfinder. */
void showAnswer(const String &text) {
const int max_chars = 38;
int lines = (text.length() + max_chars - 1) / max_chars;
if (lines > 6) lines = 6;
int h = lines * 10 + 8;
int y = 240 - h;
gfx->fillRect(0, y, 240, h, gfx->color565(0, 0, 0));
gfx->drawRect(0, y, 240, h, CYAN);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
for (int i = 0; i < lines; i++) {
gfx->setCursor(4, y + 5 + i * 10);
gfx->print(text.substring(i * max_chars, min((int)text.length(), (i + 1) * max_chars)));
}
}
/* ===========================================================================
* One full "look and tell" cycle
* =========================================================================== */
void lookAndTell(const char *prompt, const char *label) {
led(255, 120, 0); // amber: working
preview_on = false; // camera goes OFF here
gfx->fillRect(0, 0, 240, 240, BLACK);
gfx->setTextColor(YELLOW);
gfx->setTextSize(1);
gfx->setCursor(30, 110);
gfx->printf("capturing photo (%s)...", label);
size_t jpg_len = 0;
uint32_t t_cap = millis();
uint8_t *jpg = captureJpeg(&jpg_len);
t_cap = millis() - t_cap;
if (!jpg || jpg_len == 0) {
if (jpg) free(jpg);
showAnswer("Capture failed - tap CLEAR to retry.");
led(255, 0, 0);
drawClearSide();
return;
}
Serial.printf("Captured %u KB in %lu ms\n", (unsigned)(jpg_len / 1024), (unsigned long)t_cap);
gfx->setCursor(30, 124);
gfx->printf("asking OpenAI (%uKB)...", (unsigned)(jpg_len / 1024));
String answer;
uint32_t t_ai = millis();
bool ok = askVision(jpg, jpg_len, prompt, answer);
t_ai = millis() - t_ai;
free(jpg);
if (!ok) {
showAnswer("No answer - see serial monitor for the reason.");
led(255, 0, 0);
drawClearSide();
return;
}
Serial.printf("Vision (%lu ms): %s\n", (unsigned long)t_ai, answer.c_str());
showAnswer(answer);
drawClearSide(); // CLEAR replaces the ask buttons
#if SPEAK_REPLIES
led(0, 255, 40); // green: speaking
speak(answer); // answer stays on screen
#endif
led(0, 0, 0);
/* the answer now stays until the user taps CLEAR */
}
/* ===========================================================================
* SETUP
* =========================================================================== */
void setup() {
Serial.begin(115200);
delay(400);
Serial.println("\n=== 05 Vision AI | Robojax.com ===");
pinMode(TFT_BLK, OUTPUT);
digitalWrite(TFT_BLK, LOW);
pinMode(SD_CS, OUTPUT);
digitalWrite(SD_CS, HIGH);
gfx->begin();
gfx->fillScreen(BLACK);
digitalWrite(TFT_BLK, HIGH);
bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
delay(50);
// The camera's SCCB shares this bus, so Wire must be up before the camera
Wire.begin(I2C_SDA, I2C_SCL, 100000);
delay(20);
rgb.begin();
rgb.setBrightness(LED_BRIGHTNESS);
led(0, 0, 0);
#if SPEAK_REPLIES
spkInit();
#endif
gfx->setTextSize(1);
gfx->setTextColor(YELLOW);
gfx->setCursor(4, 4);
gfx->printf("Connecting to %s ...", WIFI_SSID);
WiFi.mode(WIFI_STA);
WiFi.begin(WIFI_SSID, WIFI_PASS);
uint32_t t0 = millis();
while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) delay(300);
gfx->setCursor(4, 16);
if (WiFi.status() == WL_CONNECTED) {
gfx->setTextColor(GREEN);
gfx->print("WiFi ok");
} else {
gfx->setTextColor(RED);
gfx->print("WiFi FAILED (2.4GHz only! check secrets.h)");
}
gfx->setCursor(4, 28);
gfx->setTextColor(YELLOW);
gfx->print("starting camera...");
if (!camStartPreview()) {
gfx->fillScreen(RED);
gfx->setTextColor(WHITE);
gfx->setTextSize(2);
gfx->setCursor(20, 100);
gfx->print("CAMERA FAILED");
gfx->setTextSize(1);
gfx->setCursor(20, 130);
gfx->print("Press RESET (camera reset = board reset)");
while (1) delay(1000);
}
delay(400);
gfx->fillScreen(BLACK);
drawButtons();
preview_on = true;
Serial.println("Running. Touch a button to have the AI look through the camera.");
}
/* ===========================================================================
* LOOP
* =========================================================================== */
void loop() {
static uint32_t last_touch = 0;
if (preview_on) {
camera_fb_t *fb = esp_camera_fb_get();
if (fb) {
gfx->draw16bitBeRGBBitmap(0, 0, (uint16_t *)fb->buf, fb->width, fb->height);
esp_camera_fb_return(fb);
}
} else {
delay(20); // answer on screen, waiting for CLEAR
}
/* live WiFi signal, refreshed every 2 s in both modes */
static uint32_t last_wifi = 0;
if (millis() - last_wifi > 2000) {
last_wifi = millis();
drawWifi();
}
/* Edge-detected touch: an action fires only on a NEW finger-down, and the
* finger must fully lift (4 consecutive empty reads) before anything can
* fire again. This is what stops one tap on CLEAR from also triggering the
* capture button that appears in the same spot a moment later. */
static bool touch_down = false;
static uint8_t release_count = 0;
uint16_t x, y;
if (getTouch(&x, &y)) {
release_count = 0;
if (!touch_down) {
touch_down = true; // new tap - act exactly once
if (!preview_on) {
/* answer mode: CLEAR returns to the camera, REPLAY says it again */
if (x >= BTN_X && y >= BTN_OBJ_Y && y < BTN_OBJ_Y + BTN_H) {
if (camStartPreview()) {
preview_on = true;
drawButtons();
} else {
showAnswer("Camera restart failed - press RESET.");
}
}
#if SPEAK_REPLIES
else if (x >= BTN_X && y >= BTN_PERSON_Y && y < BTN_PERSON_Y + BTN_H) {
led(0, 255, 40); // green while speaking
playLastAnswer();
led(0, 0, 0);
}
#endif
} else if (x >= BTN_X && y >= BTN_OBJ_Y && y < BTN_OBJ_Y + BTN_H) {
lookAndTell(PROMPT_OBJECTS, "objects");
} else if (x >= BTN_X && y >= BTN_PERSON_Y && y < BTN_PERSON_Y + BTN_H) {
lookAndTell(PROMPT_PERSON, "person");
}
}
} else if (touch_down) {
if (++release_count >= 4) { touch_down = false; release_count = 0; }
}
(void)last_touch;
}
Nghĩ bạn có thể cần
-
KhácProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
Tài nguyên & tài liệu tham khảo
-
Tài liệuMakerfabs MaTouch ESP32-S3 2.8" Camera and Touchscreen: user's manualwiki.makerfabs.com
-
Tài liệu
-
Tài liệuProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
-
Tải xuốngArduino GFX Library on Githubgithub.com
Tập tin📁
Tệp Yêu Cầu (.h)
Các Tệp Khác
Sơ đồ
-
MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematicThe latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF0.15 MB