This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
The latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
دې ته وګوره: د ESP32-S3 ۲.۸ انچه کامرې سره Makerfabs MaTouch ESP32-S3 AI لید: کامرې ته اشاره وکړئ او وپوښتئ چې دا څه ګوري
بورډ یوه انځور اخلي، AI تشریح کوي چې څه ویني، او ځواب په لوړ غږ وايي
کیمره په یو څه وګرځوئ، یو تڼۍ ټک کړئ، او AI تاسو ته وايي چې دا څه ګوري - په سکرین باندې چاپ شوی او د سپیکر له لارې ویل شوی. دوه حالتونه: په نظر کې موجود شیان نومول، یا د کیمرې مخې ته د شخص تشریح کول.
AI ویژن - دا په MaTouch AI ESP32-S3 بورډ باندې په یو څه باندې اشاره کوي
انځور څنګه هلته رسیږي
کیمره د 800×600 JPEG تولیدوي چې شاوخوا 20-40 KB وي. بورډ دا په PSRAM کې base64-کوډ کوي، کوم چې دا د متن په توګه شاوخوا دریمه برخه لوی کوي، او دا د data URI په توګه په غوښتنه کې ځای پر ځای کوي. 8 MB PSRAM هغه څه دي چې دا آرامه کوي.
تشریح، هیڅکله پیژندنه نه
د PERSON حالت تشریح کوي چې څه لیدلی شي - مزاج، عینکې، یو څوک څه کوي لکه څنګه چې ښکاري. دا به تاسو ته ونه وايي چې یو څوک څوک دی، او دا قصدي ده. د مخ په واسطه د خلکو پیژندنه د OpenAI د کارونې پالیسیو خلاف ده، او د مایکروسافټ معادل د مخ پیژندنې خدمت د تصویب پروسې شاته تړل شوی چې عادي پراختیا کونکي یوازې نشي کولی پکې ګډون وکړي. هر هغه ټیوټوریل چې تاسو ته د کلاوډ مخ پیژندنې ژمنه درکوي هغه څه تشریح کوي چې په عمومي توګه شتون نلري.
که تاسو یو بورډ غواړئ چې ځانګړي خلک وپیژني، پروژه 03b وکاروئ - دا دا آفلاین کوي او هیڅکله د چا مخ چیرته نه لیږي.
کیمره باید د شبکې لپاره ودریږي
ژوندی ویوفایندر بند شوی دی پداسې حال کې چې پوښتنه پوښتل کیږي. دا کاسمیټیک نه دی: روان کیمره ډرایور هغه داخلي حافظه ساتي چې خوندي اړیکې ورته اړتیا لري، او د مخکتنې سره چلولو سره اړیکه په ساده ډول ناکامیږي. د کیمرې بندول دا آزادوي. دا هم معنی لري چې ځواب د لوستلو وړ وي پرځای د 25 ځله په یوه ثانیه کې بیا رنګ کیدو.
ځواب تر هغه پورې پاتې کیږي تر څو چې تاسو یې رد کړئ
ځواب په سکرین باندې پاتې کیږي تر هغه چې تاسو CLEAR ټک کړئ - هیڅ ټایمر نشته. پداسې حال کې چې دا ښودل کیږي، د REPLAY تڼۍ ویل شوی ځواب بیا له حافظې څخه غږوي، پرته له دوهم API کال او پرته له اضافي لګښت.
د MaTouch AI ESP32-S3 2.8" بورډ په اړه
پدې پاڼه کې هره پروژه په MaTouch AI ESP32-S3 2.8" TFT ST7789V د Makerfabs څخه چلیږي. دا یو ټول-په-یو بورډ دی: یو رنګه ټ�چ سکرین، یوه 3 میګاپکسل کیمره، دوه مایکروفونونه او یو ریښتینی سپیکر امپلیفیر، ټول د ESP32-S3 لخوا د 8 MB PSRAM سره پرمخ وړل کیږي. دا ترکیب هغه څه دي چې دا AI پروژې په یوه واحده بورډ باندې د بل هیڅ شی سره نه نښلول ممکن کوي.
8 MB PSRAM دلته د بل هر شمیر څخه ډیر مهم دی. دا هغه څه دي چې بورډ ته اجازه ورکوي چې د کیمرې چوکاټ، د ثبت شوي آډیو څو ثانیې، یا د base64-کوډ شوی انځور په حافظه کې په ورته وخت کې وساتي - چې هیڅ یو یې د ESP32 په عادي RAM کې نه ځای کیږي.
د تولید کونکي اسناد: د Makerfabs ویکي پاڼه.
کلیدي مشخصات
پروسسر: ESP32-S3، دوه کور 240 MHz، WiFi 2.4 GHz + Bluetooth 5.0
حافظه: 16 MB فلش، 8 MB PSRAM (دلته د نږدې هرې پروژې لخوا اړین دی)
ښودنه: 2.8" IPS، 320×240، ST7789V ډرایور، SPI
ټچ: GT911 capacitive، په یو وخت کې 5 ګوتې تعقیبوي
کیمره: OV3660، 3 میګاپکسل، تر 2048×1536 پورې
مایکروفونونه: دوه INMP441 I2S ډیجیټل مایکونه (یو ریښتینی سټیریو جوړه)
سپیکر: MAX98357A کلاس-D امپلیفیر، 3.2 W په 4 Ω کې
ذخیره: microSD کارت سلاټ (SPI حالت)
بریښنا: USB-C، JST بیټرۍ نښلونکی، TP4056 چارجر، د بریښنا سویچ
همدارنګه په بورډ کې: WS2812B RGB LED، PCF8563T د بیټرۍ ملاتړ شوی ریښتینی وخت ساعت، او د MAX17048 بیټرۍ د تیلو ګیج چې نه په رسمي مشخصاتو کې لیست شوی
دوه USB-C بندرونه یو شان ندي. د بورډ سپیکر خپل سیګنال پنونه (IO19 او IO20) د اصلي USB بندر سره شریکوي، ځکه چې دا پنونه د ESP32-S3 هارډوایر شوي USB ډیټا کرښې دي. تل د CH340K USB-C بندر له لارې اپلوډ او بریښنا وکړئ (هغه چې د RESET تڼۍ تر څنګ دی)، او USB CDC On Boot په Disabled وټاکئ. غلط بندر وکاروئ او آډیو به ناسم چلند وکړي یا اپلوډونه به ناکام شي.
د Arduino IDE ترتیبات
دا ترتیبات مهم دي. ډیری ستونزې چې خلک د دې بورډ سره راپور ورکوي د دې څخه یوه غلطه وي، او دوی کله چې تاسو د کور نسخه بدل کړئ بیا تنظیمیږي، نو د هر بدلون وروسته یې بیا وګورئ.
ترتیب | ارزښت |
|---|---|
بورډ | ESP32S3 Dev Module |
د ESP32 کور نسخه | 2.0.17 |
PSRAM | OPI PSRAM |
د فلش اندازه | 16MB (128Mb) |
د پارټیشن سکیم | 16M Flash (3MB APP/9.9MB FATFS) |
USB CDC On Boot | Disabled |
د اپلوډ سرعت | 921600 |
د اپلوډ دمخه ټول فلش پاک کړئ | Disabled |
بندر | د CH340K USB-C بندر |
د ESP32 کور 2.0.17 وکاروئ، نه 3.x. Espressif په کور 3 کې د وسیلې پر مخ د مخ پیژندنې ماډلونه لرې کړل، نو د مخ پروژې به هلته کمپایل نشي. د 2.0.17 پن کول دا یقیني کوي چې پدې پاڼه کې هره پروژه د یو ترتیب سره کار کوي. د Boards Manager کې، د نسخې ډراپ-ډاون تاسو ته اجازه درکوي چې هر وخت چې وغواړئ شا او خوا تبدیلي وکړئ.
د اردوینو لپاره د GFX کتابتون نسخه 1.5.6 وکاروئ، نه 1.6.x. د 1.6 خپرونې د ESP32 کور 3 لپاره جوړې شوې دي او کولی شي په کور 2.0.17 کې په پیل کې ځړول شي. که ستاسو سکرین د اپلوډ کولو وروسته تور پاتې شي، دا لومړی شی دی چې باید وګورئ.
اړین کتابتونونه
دا د اردوینو IDE کې د Tools → Manage Libraries له لارې نصب کړئ. د نسخې شمیرې مهمې دي - مهرباني وکړئ هغه وکاروئ چې لیست شوي دي.
کتابتون | نسخه | لیکوال |
|---|---|---|
د اردوینو لپاره د GFX کتابتون | 1.5.6 | moononournation |
bb_captouch | 1.3.1 | Larry Bank |
ArduinoJson | 7.x | Benoit Blanchon |
Adafruit NeoPixel | هر وروستی | Adafruit |
د secrets.h تنظیمول
ستاسو د وای فای معلومات او هر API کیلي په secrets.h کې ځي، کوم چې د ځای نیونکو ارزښتونو سره په ډاونلوډ کې شامل دی. په اردوینو IDE کې هغه ټب خلاص کړئ او د خپلو معلوماتو سره یې ځای په ځای کړئ.
وای فای باید 2.4 GHz وي. ESP32-S3 کولی شي د 5 GHz شبکه په هیڅ ډول ونه ګوري. که ستاسو روټر دواړه بانډونه د یو نوم لاندې سره یوځای کړي (Asus دا Smart Connect بولي)، یا دا بند کړئ یا د 2.4 GHz بانډ ته خپل نوم ورکړئ او په secrets.h کې هغه وکاروئ.
ستاسو د API کیلي ترلاسه کول
دا پروژه د کلاوډ AI خدمت سره خبرې کوي، نو تاسو خپله کیلي ته اړتیا لرئ. که تاسو مخکې دا هیڅکله نه وي کړي، اندیښنه مه کوئ - دا د پاسورډ په څیر ورته نظر دی چې ستاسو حساب خدمت ته پیژني. دا یو څو دقیقې وخت نیسي، یو ځل.
یوه کیلي د ویب پاڼې ګډون نه دی. د مثال په توګه، د ChatGPT Plus لپاره پیسې ورکول تاسو ته API کیلي نه درکوي - دواړه جلا محصولات دي چې جلا بیل لري. تاسو د پراختیا کونکي پلیټ فارم کې حساب ته اړتیا لرئ، لکه څنګه چې لاندې تشریح شوي.
OpenAI - سترګې
OpenAI هغه لید ماډل چمتو کوي چې انځورونه ګوري. یوازې هغه پروژې چې انځور لیږي دې ته اړتیا لري.
platform.openai.com/api-keys ته لاړ شئ او ننوځئ یا حساب جوړ کړئ.
په Create new secret key کلیک وکړئ، ورته یو نوم ورکړئ، او یې جوړ کړئ.
سمدستي یې کاپي کړئ - لکه DeepSeek، دا یوازې یو ځل ښودل کیږي.
Billing خلاص کړئ او لږ مقدار کریډیټ اضافه کړئ. API مخکې له مخکې تادیه شوی دی او د هر هغه ChatGPT ګډون څخه جلا دی چې تاسو یې دمخه لرئ.
کیلي په secrets.h کې د OPENAI_KEY په توګه واچوئ.
Microsoft Azure Speech - د اوریدو او خبرو کولو لپاره
Azure ستاسو خبرې متن ته اړوي او ځواب بیرته غږ ته اړوي. وړیا کچه د دې پاڼې د هرڅه لپاره کافي ده.
portal.azure.com ته لاړ شئ او د مایکروسافټ حساب سره ننوځئ (وړیا حساب ښه دی).
که تاسو مخکې Azure نه وي کارولی، تاسو به د Welcome to Azure سکرین وګورئ چې درې انتخابونه وړاندې کوي. Start with an Azure free trial غوره کړئ - تاسو ګډون ته اړتیا لرئ مخکې لدې چې Azure تاسو ته اجازه درکړي هرڅه جوړ کړئ. (زده کونکي باید پرځای یې Azure for Students غوره کړي: ورته پایله، کارت ته اړتیا نشته.) Manage Microsoft Entra ID له پامه غورځوئ، کوم چې په بشپړ ډول بل شی دی.
په Create a resource کلیک وکړئ، د Speech لپاره وپلټئ، او د مایکروسافټ لخوا خپور شوی Speech service غوره کړئ.
فورمه ډکه کړئ: هر ډول سرچینه ګروپ، هر نوم، او ستاسو ته نږدې Region غوره کړئ - هغه سیمه په سمه توګه ولیکئ لکه څنګه چې ښکاري، د مثال په توګه
eastus.د Pricing tier لپاره F0 (Free) غوره کړئ. دا په میاشت کې شاوخوا پنځه ساعته د خبرو څخه متن ته او نیم ملیون حروف د متن څخه خبرو ته اجازه ورکوي.
په Review + create کلیک وکړئ، بیا Create. شاوخوا یوه دقیقه انتظار وکړئ، بیا په Go to resource کلیک وکړئ.
په کیڼ اړخ مینو کې Keys and Endpoint خلاص کړئ. KEY 1 او Location/Region کاپي کړئ.
هغه په secrets.h کې د AZURE_SPEECH_KEY او AZURE_REGION په توګه واچوئ. د AZURE_STT_HOST لپاره، <region>.stt.speech.microsoft.com وکاروئ - نو د eastus سیمې سره دا eastus.stt.speech.microsoft.com دی.
د کریډیټ کارت په اړه. د Azure وړیا آزموینه ستاسو د هویت تصدیق کولو لپاره کارت غوښتنه کوي. دا تاسو ته پیسې نه اخلي. تاسو د 30 ورځو لپاره $200 کریډیټ ترلاسه کوئ، او وروسته حساب د Pay-As-You-Go ته ځي - مګر د F0 خبرې کچه وړیا پاتې کیږي، میاشت په میاشت، او په دې پروژو کې هرڅه په آرامۍ سره دننه فټ کیږي. که تاسو غواړئ په هیڅ ډول کارت ورنکړئ او زده کونکی یاست، د Azure for Students اختیار تاسو ته پرته له کارت څخه کریډیټ درکوي.
دا باید د "Speech service" سرچینه وي. د Translator، Language، یا عمومي Cognitive Services سرچینې څخه کیلي ورته ښکاري او په بشپړ ډول معتبره ده - مګر هر د خبرو غوښتنه 401 تېروتنه بیرته راګرځوي. دا موږ د ازموینې پرمهال ونیول او یو ساعت یې مصرف کړ. که خبرې د 401 سره ناکامې شي پداسې حال کې چې کیلي سمه ښکاري، وګورئ چې تاسو کوم ډول سرچینه جوړه کړې ده.
د چلولو لګښت څومره دی
ډیر لږ، مګر وړیا نه دی، او تاسو باید په نږدې توګه پوه شئ چې څومره مصرف کوئ مخکې لدې چې پروژه روانه پریږدئ.
خدمت | نږدې لګښت |
|---|---|
Azure Speech | وړیا کچه په میاشت کې شاوخوا ۵ ساعته اوریدل او ۰.۵ میلیونه توري خبرې کول پوښي |
DeepSeek | د هر ځواب لپاره د یو سینټ یوه برخه - د څو ډالرو لپاره زرګونه ځوابونه |
OpenAI vision | د هر انځور لپاره تقریبا یو یا دوه سینټه، د ماډل پورې اړه لري |
بیې بدلیږي، نو دا د نرخ پر ځای د لارښود په توګه وګورئ. د دې هر خدمت یوه کارونې پاڼه لري چې هلته تاسو کولی شئ وګورئ چې څومره مصرف کړی، او ټول تاسو ته اجازه درکوي چې د لګښت حد وټاکئ - کوم چې په لومړۍ ورځ کول ارزښت لري.
خپل کليډونه شخصي وساتئ. هر څوک چې دوی ولري کولی شي ستاسو پیسې مصرف کړي. دوی په ویډیو، سکرین شاټ، فورم پوسټ یا عامه کوډ ذخیره کې مه اچوئ. که کوم کليډ ښکاره شي، د چمتو کونکي په ویب پاڼه کې یې ړنګ کړئ او نوی جوړ کړئ - دا یوازې څو ثانیې وخت نیسي، او دا یوازینی ریښتینی حل دی.
ستونزې حل کول
نښه | علت او حل |
|---|---|
سکرین تور پاتې کیږي | د GFX کتابتون غلطه نسخه (۱.۵.۶ وکاروئ) یا د بورډ غلط تنظیمات. |
|
|
هیڅ نه پورته کیږي / د COM بندر نشته | غلط USB-C بندر، یا د CH340 ډرایور نصب شوی نه دی. |
کیمره ناکامیږي او بیرته نه راګرځي | د کیمرې ریسیټ کرښه د بورډ RESET تڼۍ سره تړلې ده، نو سافټویر نشي کولی دا بیا پیل کړي. RESET فشار کړئ. که بیا هم ناکامه شي، د کیمرې ریبن کیبل بیرته ځای کې کېږدئ. |
کوډ ډاونلوډ کړئ
د دې پروژې بشپړ Arduino سکیچ، د pins.h او نورو ټولو اړینو شیانو سره، په وړیا توګه ډاونلوډ لپاره شتون لري.
دا انزیپ کړئ، د .ino فایل د Arduino IDE کې خلاص کړئ، پورته تنظیمات وګورئ، او د CH340K USB-C بندر له لارې پورته یې کړئ.
This tutorial is part of: Makerfabs MaTouch AI ESP32S3 2.8" Camera
/*
* ===========================================================================
* 05_Vision_AI — MaTouch AI ESP32-S3 2.8" TFT ST7789V
* ===========================================================================
*
----------
* ROBOJAX.COM - MaTouch AI ESP32-S3 2.8" project series
*
* WATCH THE VIDEO
* https://youtu.be/6AL3g3tC_Hk
*
* WRITTEN TUTORIALS - every project, with photos and full explanation
* Camera and touchscreen.... https://robojax.com/RTJ849
* Offline face recognition.. https://robojax.com/RTJ850
* AI voice assistant........ https://robojax.com/RTJ851
* AI vision................. https://robojax.com/RTJ852
*
* GET THE BOARD - SAVE $5 with coupon code: Robojax_Makerfab
* https://www.makerfabs.com/matouch-ai-esp32s3-2-8-tft-st7789v.html
* (enter the code at checkout)
*
* All of this code is free. If it helped you, a subscribe on YouTube is
* the best way to support more of it.
*
*
* Watching the video first will save you time - it shows the Arduino IDE
* settings and the library versions being set up step by step.
* ---------------------------------------------------------------------------
*
* The board SEES. Point the camera at something, touch a button, and an
* OpenAI vision model tells you what it is looking at - drawn on the screen
* and (optionally) spoken out of the board's own speaker via Azure TTS.
*
* Two modes, two buttons:
*
* OBJECTS "What do you see?" - names the things in front of the lens.
* PERSON Describes the person in frame: glasses, expression, what
* they are doing. DESCRIPTION ONLY - it will not and must not
* try to say WHO someone is. Identifying people by face is
* against OpenAI's usage policies, and it is the right call:
* say this in the video, it is worth 15 honest seconds.
*
* WHY OPENAI FOR THIS DEMO AND NOT DEEPSEEK: DeepSeek's API is text-only.
* It cannot accept an image at all. Azure OpenAI could do it, but plain
* OpenAI is one endpoint with no deployment setup - simplest to follow.
*
* HOW THE IMAGE TRAVELS: the OV3660 gives us a JPEG directly (800x600,
* ~40 KB). We base64-encode it in PSRAM (~55 KB of text) and embed it in
* the JSON request as a data: URI. The 8 MB PSRAM makes this trivial.
*
* ---------------------------------------------------------------------------
* FILL IN secrets.h BEFORE FLASHING
* (needs WIFI_*, OPENAI_*; AZURE_* only if SPEAK_REPLIES is 1).
*
* BOARD SETTINGS (Tools menu - EVERY line matters, wrong = black screen
* or compile errors. These reset when you switch cores - recheck them!)
*
* Board : ESP32S3 Dev Module
* ESP32 core : 2.0.17
* PSRAM : OPI PSRAM <-- required, image buffers live there
* Flash Size : 16MB (128Mb)
* Partition Scheme : 16M Flash (3MB APP/9.9MB FATFS)
* USB CDC On Boot : Disabled <-- speaker shares pins with native USB
* Upload Speed : 921600
* Port : the CH340K USB-C port (the one near RESET)
*
* LIBRARIES
* GFX Library for Arduino v1.5.6 (NOT 1.6.x - that pairs with core 3)
* bb_captouch v1.3.1
* ArduinoJson v7.x
* Adafruit NeoPixel any recent
*
* ---------------------------------------------------------------------------
* FUNCTIONS IN THIS SKETCH
* led(r,g,b) set the RGB status LED colour
* getTouch(&x,&y) read the touch panel, mapped to screen coordinates
* camStart(...) start the camera in a given format/size
* camStartPreview() start the 240x240 RGB565 live-view camera
* captureJpeg(&len) take one 800x600 JPEG into PSRAM (camera stays OFF
* afterwards - caller restarts the preview)
* readHttpResponse() read an HTTPS reply, de-chunking it properly
* askVision(...) send photo + prompt to OpenAI, return the answer
* spkInit() configure the I2S speaker output
* readExact(...) read exactly N bytes from a TLS connection
* speak(text) Azure TTS -> download voice to PSRAM -> play it
* playLastAnswer() replay the kept voice from PSRAM (REPLAY button)
* drawButton(...) draw one side-column button
* drawWifi() WiFi signal bars + dBm readout
* drawButtons() normal side column (the two ask buttons)
* drawClearSide() answer-mode side column (CLEAR + REPLAY)
* showAnswer(text) word-wrapped answer overlay on the viewfinder
* lookAndTell(...) one full cycle: capture -> ask -> show -> speak
* setup() / loop() boot sequence / viewfinder + touch handling
*
* Robojax.com
* ===========================================================================
*/
#define SPEAK_REPLIES 1 // 1 = read the answer aloud with Azure TTS, 0 = screen only
/* Status LED brightness, 0-255. The WS2812 runs from the power rail and is
* uncomfortably bright at full power - 25 is plenty visible on camera. */
#define LED_BRIGHTNESS 15
/* CAMERA ORIENTATION
* 1 = camera faces the SAME way as the screen (how the board ships) - the
* image is mirrored so it looks natural when you point it at yourself.
* 0 = you folded the ribbon so the lens faces AWAY from the screen
* (phone-style, screen to you / camera to the subject).
* If the picture looks left-right reversed, flip this number. */
#define CAMERA_FACES_USER 1
#include <Arduino_GFX_Library.h>
#include <bb_captouch.h>
#include <Adafruit_NeoPixel.h>
#include <ArduinoJson.h>
#include <Wire.h>
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <HTTPClient.h>
#include "mbedtls/base64.h"
#include "esp_camera.h"
#include "pins.h"
#include "secrets.h"
#if SPEAK_REPLIES
#include "driver/i2s.h"
#define WAV_HEADER_LEN 44
#endif
Arduino_ESP32SPI *bus = new Arduino_ESP32SPI(
TFT_DC, TFT_CS, TFT_SCLK, TFT_MOSI, TFT_MISO, HSPI, true);
Arduino_GFX *gfx = new Arduino_ST7789(bus, TFT_RES, 1, true);
BBCapTouch bbct;
Adafruit_NeoPixel rgb(RGB_LED_NUM, RGB_LED_PIN, NEO_GRB + NEO_KHZ800);
/* --- the two vision prompts ------------------------------------------------
* Keep answers short: they must fit a 320x240 screen and, if spoken, must not
* leave the presenter waiting awkwardly on camera. */
const char *PROMPT_OBJECTS =
"Look at this photo from a small camera. In ONE short sentence of at most "
"20 words, plain text only, name the main object(s) you see.";
const char *PROMPT_PERSON =
"Look at this photo. If there is a person, describe them in ONE short "
"sentence of at most 20 words: mood, glasses or not, what they are doing. "
"Never guess who they are. If no person is visible, say so. Plain text only.";
/* --- layout: viewfinder left, buttons right ------------------------------- */
#define BTN_X 242
#define BTN_W 78
#define BTN_H 52
#define BTN_OBJ_Y 4
#define BTN_PERSON_Y 62
void led(uint8_t r, uint8_t g, uint8_t b) { rgb.setPixelColor(0, rgb.Color(r, g, b)); rgb.show(); }
/* ===========================================================================
* Touch
* =========================================================================== */
bool getTouch(uint16_t *x, uint16_t *y) {
TOUCHINFO ti;
if (!bbct.getSamples(&ti)) return false;
if (ti.count < 1) return false;
*x = ti.y[0];
*y = (ti.x[0] > 240) ? 0 : (240 - ti.x[0]);
return true;
}
/* ===========================================================================
* Camera — RGB565 for the live view; JPEG capture happens by restarting
* the driver, same technique as sketch 02.
* =========================================================================== */
bool mirrored = CAMERA_FACES_USER; // see the define at the top of the file
/* The viewfinder is OFF while a cloud call runs and while the answer is on
* screen. Two reasons, both learned on the bench:
* 1. RAM: the live camera driver eats the internal memory a TLS handshake
* needs - with the preview running, connecting to OpenAI fails (HTTP -1).
* 2. Readability: the viewfinder repaints 25x/s and would wipe the answer.
* The answer STAYS on screen until the user taps CLEAR - no timer. */
bool preview_on = false;
/* The last spoken answer is KEPT in PSRAM so the REPLAY button can play it
* again without another cloud call. Overwritten by the next answer. */
uint8_t *last_audio = nullptr;
size_t last_audio_len = 0;
bool camStart(pixformat_t fmt, framesize_t size, int fb_count, int quality) {
camera_config_t c;
c.ledc_channel = LEDC_CHANNEL_0;
c.ledc_timer = LEDC_TIMER_0;
c.pin_d0 = CAM_PIN_D0; c.pin_d1 = CAM_PIN_D1;
c.pin_d2 = CAM_PIN_D2; c.pin_d3 = CAM_PIN_D3;
c.pin_d4 = CAM_PIN_D4; c.pin_d5 = CAM_PIN_D5;
c.pin_d6 = CAM_PIN_D6; c.pin_d7 = CAM_PIN_D7;
c.pin_xclk = CAM_PIN_XCLK;
c.pin_pclk = CAM_PIN_PCLK;
c.pin_vsync = CAM_PIN_VSYNC;
c.pin_href = CAM_PIN_HREF;
/* Share the I2C bus Wire already drives (the touch panel lives there too)
* instead of letting the camera install a second driver on the same pins -
* that kills touch. Wire.begin() must run before this function. */
c.pin_sccb_sda = -1;
c.pin_sccb_scl = -1;
c.sccb_i2c_port = 0; // Wire = I2C port 0
c.pin_pwdn = CAM_PIN_PWDN;
c.pin_reset = CAM_PIN_RESET;
c.xclk_freq_hz = 20000000;
c.frame_size = size;
c.pixel_format = fmt;
c.grab_mode = CAMERA_GRAB_WHEN_EMPTY;
c.fb_location = CAMERA_FB_IN_PSRAM;
c.jpeg_quality = quality;
c.fb_count = fb_count;
if (esp_camera_init(&c) != ESP_OK) return false;
sensor_t *s = esp_camera_sensor_get();
if (s) {
s->set_hmirror(s, mirrored ? 1 : 0);
s->set_vflip(s, mirrored ? 1 : 0);
s->set_brightness(s, 1);
}
return true;
}
bool camStartPreview() { return camStart(PIXFORMAT_RGB565, FRAMESIZE_240X240, 2, 12); }
/* Capture one SVGA JPEG into a PSRAM buffer the caller owns. */
uint8_t *captureJpeg(size_t *len_out) {
esp_camera_deinit();
delay(120);
if (!camStart(PIXFORMAT_JPEG, FRAMESIZE_SVGA, 1, 12)) { *len_out = 0; return nullptr; }
// a few warm-up frames so exposure settles
for (int i = 0; i < 3; i++) {
camera_fb_t *w = esp_camera_fb_get();
if (w) esp_camera_fb_return(w);
delay(100);
}
uint8_t *copy = nullptr;
*len_out = 0;
camera_fb_t *fb = esp_camera_fb_get();
if (fb && fb->len > 0) {
copy = (uint8_t *)ps_malloc(fb->len);
if (copy) { memcpy(copy, fb->buf, fb->len); *len_out = fb->len; }
}
if (fb) esp_camera_fb_return(fb);
/* Deliberately leave the camera OFF here. The caller restarts the preview
* after the cloud call - a running camera driver starves the TLS handshake
* of internal RAM (that was the "Vision HTTP -1" bug). */
esp_camera_deinit();
delay(120);
return copy;
}
/* ===========================================================================
* HTTP response reader — shared by the cloud calls below.
* Returns the status code and fills body_out. Handles chunked transfer
* encoding PROPERLY: the chunk-size markers must be stripped, or they end
* up embedded inside the JSON and the parse fails. (OpenAI pretty-prints
* its replies so they span several chunks - this bit us on the bench.)
* =========================================================================== */
/* Block until the connection has data (or the budget runs out). Every read
* below goes through this, because the server can go quiet for many seconds
* while it analyses the image - and a bare read() would simply time out. */
static bool waitData(WiFiClientSecure &c, uint32_t ms) {
uint32_t t0 = millis();
while (!c.available()) {
if (!c.connected()) return false;
if (millis() - t0 > ms) return false;
delay(10);
}
return true;
}
static int readHttpResponse(WiFiClientSecure &client, String &body_out, uint32_t idle_ms) {
body_out = "";
if (!waitData(client, idle_ms)) { Serial.println("HTTP: no response at all"); return 0; }
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
while (waitData(client, idle_ms)) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break; // blank line = end of headers
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
}
if (chunked) {
int blanks = 0;
while (true) {
/* Waiting here - instead of letting read() time out - is essential: a
* timed-out read looks exactly like "0" (final chunk), which silently
* truncates the body to nothing while the server is still thinking. */
if (!waitData(client, idle_ms)) {
Serial.println("HTTP: timed out waiting for the next chunk");
break;
}
String szline = client.readStringUntil('\n');
szline.trim();
if (szline.length() == 0) {
if (++blanks > 4) break;
continue;
}
blanks = 0;
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break; // genuine final chunk
long got = 0;
while (got < sz) {
if (!waitData(client, idle_ms)) break;
while (client.available() && got < sz) { body_out += (char)client.read(); got++; }
}
if (waitData(client, 3000)) client.readStringUntil('\n');
if (got < sz) { Serial.println("HTTP: short chunk"); break; }
}
} else {
while (waitData(client, idle_ms))
while (client.available()) body_out += (char)client.read();
}
return code;
}
/* ===========================================================================
* OpenAI vision call
* The request body is built by hand in PSRAM rather than through ArduinoJson,
* because embedding a 55 KB base64 string in a JSON document would mean
* holding two copies. Base64 text never needs JSON escaping, so this is safe.
* =========================================================================== */
bool askVision(const uint8_t *jpg, size_t jpg_len, const char *prompt, String &answer_out) {
// --- base64 encode the image into PSRAM ---
size_t b64_cap = ((jpg_len + 2) / 3) * 4 + 16;
unsigned char *b64 = (unsigned char *)ps_malloc(b64_cap);
if (!b64) return false;
size_t b64_len = 0;
if (mbedtls_base64_encode(b64, b64_cap, &b64_len, jpg, jpg_len) != 0) {
free(b64);
return false;
}
// --- assemble the JSON request around it ---
/* Current OpenAI models reject the old "max_tokens" name - it must be
* "max_completion_tokens" (verified against the live API, July 2026). */
const char *head_fmt =
"{\"model\":\"%s\",\"max_completion_tokens\":%d,\"messages\":[{\"role\":\"user\","
"\"content\":[{\"type\":\"text\",\"text\":\"%s\"},"
"{\"type\":\"image_url\",\"image_url\":{\"url\":\"data:image/jpeg;base64,";
const char *tail = "\"}}]}]}";
char head[640];
int head_len = snprintf(head, sizeof(head), head_fmt, OPENAI_MODEL, LLM_MAX_TOKENS, prompt);
size_t body_len = head_len + b64_len + strlen(tail);
char *body = (char *)ps_malloc(body_len + 1);
if (!body) { free(b64); return false; }
memcpy(body, head, head_len);
memcpy(body + head_len, b64, b64_len);
strcpy(body + head_len + b64_len, tail);
free(b64);
// --- send it: manual HTTP, chunk-wise upload (proven pattern from 04) ---
WiFiClientSecure client;
client.setInsecure();
client.setTimeout(25);
if (!client.connect(OPENAI_HOST, 443)) {
Serial.println("Vision: TLS connect failed (is the camera still running?)");
free(body);
return false;
}
client.print(String("POST /v1/chat/completions HTTP/1.1\r\n"
"Host: " OPENAI_HOST "\r\n"
"Authorization: Bearer " OPENAI_KEY "\r\n"
"Content-Type: application/json\r\n"
"Connection: close\r\n"
"Content-Length: ") + String(body_len) + "\r\n\r\n");
size_t sent = 0;
while (sent < body_len) {
size_t n = min((size_t)4096, body_len - sent);
size_t w = client.write((uint8_t *)body + sent, n);
if (w == 0) {
delay(50);
w = client.write((uint8_t *)body + sent, n);
if (w == 0) break;
}
sent += w;
yield();
}
free(body);
if (sent < body_len) {
Serial.printf("Vision: upload stalled at %u/%u bytes\n",
(unsigned)sent, (unsigned)body_len);
client.stop();
return false;
}
/* read the reply with proper de-chunking */
String resp;
int code = readHttpResponse(client, resp, 30000);
client.stop();
if (code != 200) {
Serial.printf("Vision HTTP %d: %s\n", code, resp.c_str());
return false;
}
bool ok = false;
JsonDocument doc;
if (!deserializeJson(doc, resp)) {
const char *content = doc["choices"][0]["message"]["content"];
if (content) {
answer_out = String(content);
answer_out.trim();
ok = answer_out.length() > 0;
}
} else {
Serial.printf("Vision: JSON parse failed, %u bytes received\n", resp.length());
}
return ok;
}
/* ===========================================================================
* Azure TTS — same streaming trick as sketch 04: raw PCM into I2S.
* =========================================================================== */
#if SPEAK_REPLIES
void spkInit() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_TX),
.sample_rate = 16000,
.bits_per_sample = I2S_BITS_PER_SAMPLE_16BIT,
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 8,
.dma_buf_len = 256,
.use_apll = false,
.tx_desc_auto_clear = true,
.fixed_mclk = 0
};
i2s_pin_config_t pins = {
.mck_io_num = I2S_PIN_NO_CHANGE,
.bck_io_num = I2S_SPK_BCLK,
.ws_io_num = I2S_SPK_LRC,
.data_out_num = I2S_SPK_DOUT,
.data_in_num = I2S_PIN_NO_CHANGE
};
i2s_driver_install(I2S_SPK_PORT, &cfg, 0, NULL);
i2s_set_pin(I2S_SPK_PORT, &pins);
i2s_zero_dma_buffer(I2S_SPK_PORT);
}
/* Read exactly n bytes from a client (or until timeout). */
static size_t readExact(WiFiClientSecure &c, uint8_t *dst, size_t n) {
size_t got = 0;
uint32_t t0 = millis();
while (got < n && millis() - t0 < 10000) {
int r = c.read(dst + got, n - got);
if (r > 0) { got += r; t0 = millis(); }
else if (!c.connected() && !c.available()) break;
else delay(2);
}
return got;
}
void playLastAnswer(); // defined below; explicit prototype for the IDE
void speak(const String &text) {
String safe = text;
safe.replace("&", "&");
safe.replace("<", "<");
safe.replace(">", ">");
String ssml = "<speak version='1.0' xml:lang='" AZURE_TTS_LANG "'>"
"<voice name='" AZURE_TTS_VOICE "'>" + safe + "</voice></speak>";
/* Manual HTTP with proper de-chunking: Azure sends this audio chunked, and
* HTTPClient's raw stream leaks the ASCII chunk-size lines into the PCM -
* each one plays as an audible KNOCK. */
WiFiClientSecure client;
client.setInsecure();
client.setTimeout(20);
const char *host = AZURE_REGION ".tts.speech.microsoft.com";
if (!client.connect(host, 443)) { Serial.println("TTS: TLS connect failed"); return; }
client.print(String("POST /cognitiveservices/v1 HTTP/1.1\r\n"
"Host: ") + host + "\r\n"
"Ocp-Apim-Subscription-Key: " AZURE_SPEECH_KEY "\r\n"
"Content-Type: application/ssml+xml\r\n"
"X-Microsoft-OutputFormat: riff-16khz-16bit-mono-pcm\r\n"
"User-Agent: MaTouchRobojax\r\n"
"Connection: close\r\n"
"Content-Length: " + String(ssml.length()) + "\r\n\r\n");
client.print(ssml);
String status_line = client.readStringUntil('\n');
int code = 0;
sscanf(status_line.c_str(), "HTTP/%*s %d", &code);
bool chunked = false;
long content_len = -1;
while (client.connected() || client.available()) {
String h = client.readStringUntil('\n');
if (h == "\r" || h.length() <= 1) break;
h.toLowerCase();
if (h.startsWith("transfer-encoding:") && h.indexOf("chunked") >= 0) chunked = true;
if (h.startsWith("content-length:")) content_len = h.substring(15).toInt();
}
if (code != 200) { Serial.printf("TTS HTTP %d\n", code); client.stop(); return; }
const size_t AUDIO_CAP = 1200 * 1024;
uint8_t *audio = (uint8_t *)ps_malloc(AUDIO_CAP);
if (!audio) { client.stop(); return; }
size_t alen = 0;
if (chunked) {
while (true) {
String szline = client.readStringUntil('\n');
long sz = strtol(szline.c_str(), NULL, 16);
if (sz <= 0) break;
if (alen + sz > AUDIO_CAP) break;
size_t got = readExact(client, audio + alen, sz);
alen += got;
client.readStringUntil('\n');
if (got < (size_t)sz) break;
}
} else if (content_len > 0) {
alen = readExact(client, audio, min((size_t)content_len, AUDIO_CAP));
} else {
uint32_t idle = millis();
while ((client.connected() || client.available()) && millis() - idle < 5000) {
int r = client.read(audio + alen, min((size_t)2048, AUDIO_CAP - alen));
if (r > 0) { alen += r; idle = millis(); }
else delay(5);
}
}
client.stop();
if (alen > WAV_HEADER_LEN) {
/* keep this answer for the REPLAY button (replacing the previous one),
* then play it */
if (last_audio) free(last_audio);
last_audio = audio;
last_audio_len = alen;
playLastAnswer();
} else {
free(audio);
}
}
/* Play the kept answer from PSRAM - used right after download AND by REPLAY. */
void playLastAnswer() {
if (!last_audio || last_audio_len <= WAV_HEADER_LEN) return;
static const uint8_t lead_in[640] = {0}; // 20 ms silence pre-roll
size_t w = 0;
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, last_audio + WAV_HEADER_LEN,
last_audio_len - WAV_HEADER_LEN, &w, portMAX_DELAY);
i2s_write(I2S_SPK_PORT, lead_in, sizeof(lead_in), &w, portMAX_DELAY);
delay(150);
i2s_zero_dma_buffer(I2S_SPK_PORT);
}
#endif // SPEAK_REPLIES
/* ===========================================================================
* UI
* =========================================================================== */
void drawButton(int y, const char *l1, const char *l2, uint16_t colour) {
gfx->fillRoundRect(BTN_X, y, BTN_W, BTN_H, 6, colour);
gfx->drawRoundRect(BTN_X, y, BTN_W, BTN_H, 6, WHITE);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_X + 8, y + 14);
gfx->print(l1);
gfx->setCursor(BTN_X + 8, y + 28);
gfx->print(l2);
}
/* WiFi bars + dBm at the bottom of the side column, refreshed from the loop */
void drawWifi() {
gfx->fillRect(BTN_X, 188, BTN_W, 52, BLACK);
bool up = (WiFi.status() == WL_CONNECTED);
long rssi = up ? WiFi.RSSI() : -100;
int bars = rssi > -55 ? 4 : rssi > -65 ? 3 : rssi > -75 ? 2 : rssi > -85 ? 1 : 0;
for (int b = 0; b < 4; b++) {
int bh = 6 + b * 6;
uint16_t col = (b < bars) ? GREEN : gfx->color565(60, 60, 60);
gfx->fillRect(BTN_X + 4 + b * 9, 216 - bh, 7, bh, col);
}
gfx->setTextSize(1);
gfx->setCursor(BTN_X + 44, 196);
if (up) {
gfx->setTextColor(CYAN);
gfx->printf("%ld", rssi);
} else {
gfx->setTextColor(RED);
gfx->print("DOWN");
}
gfx->setCursor(BTN_X + 44, 208);
gfx->setTextColor(gfx->color565(120, 120, 120));
gfx->print("dBm");
gfx->setCursor(BTN_X + 4, 228);
gfx->setTextColor(gfx->color565(120, 120, 120));
gfx->print("WiFi signal");
}
/* normal side column: the two ask buttons */
void drawButtons() {
gfx->fillRect(BTN_X, 0, BTN_W, 188, BLACK);
drawButton(BTN_OBJ_Y, "WHAT DO", "YOU SEE?", gfx->color565(0, 90, 160));
drawButton(BTN_PERSON_Y, "DESCRIBE", "PERSON", gfx->color565(120, 60, 140));
gfx->setTextColor(CYAN);
gfx->setCursor(BTN_X + 2, 128);
gfx->print("OpenAI eyes");
gfx->setCursor(BTN_X + 2, 140);
gfx->print("Azure voice");
gfx->setCursor(BTN_X + 2, 152);
gfx->print("Robojax.com");
drawWifi();
}
/* answer-mode side column: CLEAR (back to camera) + REPLAY (say it again).
* The answer stays on screen until CLEAR is tapped. */
void drawClearSide() {
gfx->fillRect(BTN_X, 0, BTN_W, 188, BLACK);
gfx->fillRoundRect(BTN_X, BTN_OBJ_Y, BTN_W, BTN_H, 6, gfx->color565(110, 35, 35));
gfx->drawRoundRect(BTN_X, BTN_OBJ_Y, BTN_W, BTN_H, 6, WHITE);
gfx->setTextSize(2);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_X + 10, BTN_OBJ_Y + 14);
gfx->print("CLEAR");
#if SPEAK_REPLIES
if (last_audio_len > 0) {
gfx->fillRoundRect(BTN_X, BTN_PERSON_Y, BTN_W, BTN_H, 6, gfx->color565(0, 110, 60));
gfx->drawRoundRect(BTN_X, BTN_PERSON_Y, BTN_W, BTN_H, 6, WHITE);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
gfx->setCursor(BTN_X + 16, BTN_PERSON_Y + 20);
gfx->print("REPLAY");
}
#endif
gfx->setTextSize(1);
gfx->setTextColor(YELLOW);
gfx->setCursor(BTN_X + 2, 124);
gfx->print("CLEAR = camera");
gfx->setCursor(BTN_X + 2, 136);
gfx->print("REPLAY = again");
drawWifi();
}
/* Word-wrapped answer overlay across the bottom of the viewfinder. */
void showAnswer(const String &text) {
const int max_chars = 38;
int lines = (text.length() + max_chars - 1) / max_chars;
if (lines > 6) lines = 6;
int h = lines * 10 + 8;
int y = 240 - h;
gfx->fillRect(0, y, 240, h, gfx->color565(0, 0, 0));
gfx->drawRect(0, y, 240, h, CYAN);
gfx->setTextSize(1);
gfx->setTextColor(WHITE);
for (int i = 0; i < lines; i++) {
gfx->setCursor(4, y + 5 + i * 10);
gfx->print(text.substring(i * max_chars, min((int)text.length(), (i + 1) * max_chars)));
}
}
/* ===========================================================================
* One full "look and tell" cycle
* =========================================================================== */
void lookAndTell(const char *prompt, const char *label) {
led(255, 120, 0); // amber: working
preview_on = false; // camera goes OFF here
gfx->fillRect(0, 0, 240, 240, BLACK);
gfx->setTextColor(YELLOW);
gfx->setTextSize(1);
gfx->setCursor(30, 110);
gfx->printf("capturing photo (%s)...", label);
size_t jpg_len = 0;
uint32_t t_cap = millis();
uint8_t *jpg = captureJpeg(&jpg_len);
t_cap = millis() - t_cap;
if (!jpg || jpg_len == 0) {
if (jpg) free(jpg);
showAnswer("Capture failed - tap CLEAR to retry.");
led(255, 0, 0);
drawClearSide();
return;
}
Serial.printf("Captured %u KB in %lu ms\n", (unsigned)(jpg_len / 1024), (unsigned long)t_cap);
gfx->setCursor(30, 124);
gfx->printf("asking OpenAI (%uKB)...", (unsigned)(jpg_len / 1024));
String answer;
uint32_t t_ai = millis();
bool ok = askVision(jpg, jpg_len, prompt, answer);
t_ai = millis() - t_ai;
free(jpg);
if (!ok) {
showAnswer("No answer - see serial monitor for the reason.");
led(255, 0, 0);
drawClearSide();
return;
}
Serial.printf("Vision (%lu ms): %s\n", (unsigned long)t_ai, answer.c_str());
showAnswer(answer);
drawClearSide(); // CLEAR replaces the ask buttons
#if SPEAK_REPLIES
led(0, 255, 40); // green: speaking
speak(answer); // answer stays on screen
#endif
led(0, 0, 0);
/* the answer now stays until the user taps CLEAR */
}
/* ===========================================================================
* SETUP
* =========================================================================== */
void setup() {
Serial.begin(115200);
delay(400);
Serial.println("\n=== 05 Vision AI | Robojax.com ===");
pinMode(TFT_BLK, OUTPUT);
digitalWrite(TFT_BLK, LOW);
pinMode(SD_CS, OUTPUT);
digitalWrite(SD_CS, HIGH);
gfx->begin();
gfx->fillScreen(BLACK);
digitalWrite(TFT_BLK, HIGH);
bbct.init(TOUCH_SDA, TOUCH_SCL, TOUCH_RST, TOUCH_INT);
delay(50);
// The camera's SCCB shares this bus, so Wire must be up before the camera
Wire.begin(I2C_SDA, I2C_SCL, 100000);
delay(20);
rgb.begin();
rgb.setBrightness(LED_BRIGHTNESS);
led(0, 0, 0);
#if SPEAK_REPLIES
spkInit();
#endif
gfx->setTextSize(1);
gfx->setTextColor(YELLOW);
gfx->setCursor(4, 4);
gfx->printf("Connecting to %s ...", WIFI_SSID);
WiFi.mode(WIFI_STA);
WiFi.begin(WIFI_SSID, WIFI_PASS);
uint32_t t0 = millis();
while (WiFi.status() != WL_CONNECTED && millis() - t0 < 20000) delay(300);
gfx->setCursor(4, 16);
if (WiFi.status() == WL_CONNECTED) {
gfx->setTextColor(GREEN);
gfx->print("WiFi ok");
} else {
gfx->setTextColor(RED);
gfx->print("WiFi FAILED (2.4GHz only! check secrets.h)");
}
gfx->setCursor(4, 28);
gfx->setTextColor(YELLOW);
gfx->print("starting camera...");
if (!camStartPreview()) {
gfx->fillScreen(RED);
gfx->setTextColor(WHITE);
gfx->setTextSize(2);
gfx->setCursor(20, 100);
gfx->print("CAMERA FAILED");
gfx->setTextSize(1);
gfx->setCursor(20, 130);
gfx->print("Press RESET (camera reset = board reset)");
while (1) delay(1000);
}
delay(400);
gfx->fillScreen(BLACK);
drawButtons();
preview_on = true;
Serial.println("Running. Touch a button to have the AI look through the camera.");
}
/* ===========================================================================
* LOOP
* =========================================================================== */
void loop() {
static uint32_t last_touch = 0;
if (preview_on) {
camera_fb_t *fb = esp_camera_fb_get();
if (fb) {
gfx->draw16bitBeRGBBitmap(0, 0, (uint16_t *)fb->buf, fb->width, fb->height);
esp_camera_fb_return(fb);
}
} else {
delay(20); // answer on screen, waiting for CLEAR
}
/* live WiFi signal, refreshed every 2 s in both modes */
static uint32_t last_wifi = 0;
if (millis() - last_wifi > 2000) {
last_wifi = millis();
drawWifi();
}
/* Edge-detected touch: an action fires only on a NEW finger-down, and the
* finger must fully lift (4 consecutive empty reads) before anything can
* fire again. This is what stops one tap on CLEAR from also triggering the
* capture button that appears in the same spot a moment later. */
static bool touch_down = false;
static uint8_t release_count = 0;
uint16_t x, y;
if (getTouch(&x, &y)) {
release_count = 0;
if (!touch_down) {
touch_down = true; // new tap - act exactly once
if (!preview_on) {
/* answer mode: CLEAR returns to the camera, REPLAY says it again */
if (x >= BTN_X && y >= BTN_OBJ_Y && y < BTN_OBJ_Y + BTN_H) {
if (camStartPreview()) {
preview_on = true;
drawButtons();
} else {
showAnswer("Camera restart failed - press RESET.");
}
}
#if SPEAK_REPLIES
else if (x >= BTN_X && y >= BTN_PERSON_Y && y < BTN_PERSON_Y + BTN_H) {
led(0, 255, 40); // green while speaking
playLastAnswer();
led(0, 0, 0);
}
#endif
} else if (x >= BTN_X && y >= BTN_OBJ_Y && y < BTN_OBJ_Y + BTN_H) {
lookAndTell(PROMPT_OBJECTS, "objects");
} else if (x >= BTN_X && y >= BTN_PERSON_Y && y < BTN_PERSON_Y + BTN_H) {
lookAndTell(PROMPT_PERSON, "person");
}
}
} else if (touch_down) {
if (++release_count >= 4) { touch_down = false; release_count = 0; }
}
(void)last_touch;
}
Things you might need
-
OtherProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
Resources & references
-
DocumentationMakerfabs MaTouch ESP32-S3 2.8" Camera and Touchscreen: user's manualwiki.makerfabs.com
-
Documentation
-
DocumentationProduct page for MaTouch AI ESP32S3 2.8" TFT ST7789Vmakerfabs.com
-
DownloadArduino GFX Library on Githubgithub.com
Files📁
Required File (.h)
Other Files
Schematic
-
MaTouch_AI 2.8“ MaTouch AI ESP32S3 2.8" TFT ST7789V schematicThe latest MaTouch AI board integrate I2S voice input/I2S speaker/ 3 million camera OV3660/ 320*240 resolution display, with ESP32S3 strong processor& Wifi ability, to make this board a good tool/platform for AI development with ESP32.
MaTouch_AI 2.8“ SPI TFT ST7789V V1.1.PDF0.15 MB