Roni Bandini β Buenos Aires, Argentina β August 2023
Machine Learning Doorbell is a compact Computer Vision experiment that replaces the physical push button of a conventional doorbell with automatic face detection.
A Seeed Studio XIAO ESP32S3 Sense continuously captures images and runs an Edge Impulse FOMO object-detection model locally.
When one or more faces are detected, the device:
- π· Captures the current camera image.
- π§ Runs the Edge Impulse model.
- πΎ Saves the JPEG to microSD.
- π‘ Activates the onboard LED.
- π² Sends the image to a Telegram chat.
- π Sends a text notification containing the detection confidence.
- β±οΈ Waits before resuming detection.
No mechanical doorbell button is required.
- πͺ Buttonless doorbell concept
- π· Integrated camera
- π§ On-device Computer Vision inference
- π― Edge Impulse FOMO face detection
- π₯ Supports multiple detections in one frame
- πΎ Automatic JPEG storage to microSD
- π² Telegram image notifications
- π Detection confidence in the message
- π‘ Onboard LED feedback
- π‘ Wi-Fi connectivity
- β‘ ESP32-S3 dual-core processor
- π§ 8 MB PSRAM
- π¦ Edge Impulse Arduino library included
- π Single compact board + Sense expansion module
flowchart LR
VISITOR["π€ Visitor"]
CAMERA["π· XIAO Camera"]
ESP["ESP32-S3"]
FRAME["320Γ240 JPEG"]
EI["π§ Edge Impulse<br/>FOMO"]
FACE{"Face detected?"}
SD["πΎ microSD"]
LED["π‘ LED"]
WIFI["π‘ Wi-Fi"]
TG["π² Telegram"]
VISITOR --> CAMERA
CAMERA --> ESP
ESP --> FRAME
FRAME --> SD
FRAME --> EI
EI --> FACE
FACE -->|"Yes"| LED
FACE -->|"Yes"| WIFI
WIFI --> TG
Image capture, preprocessing and ML inference run directly on the ESP32-S3.
Telegram is used only for notification delivery.
The project uses the Seeed Studio XIAO ESP32S3 Sense.
Current hardware specifications include:
| Feature | Specification |
|---|---|
| MCU | ESP32-S3R8 |
| CPU | Dual-core Xtensa LX7 |
| Clock | Up to 240 MHz |
| PSRAM | 8 MB |
| Flash | 8 MB |
| Wi-Fi | 2.4 GHz |
| Bluetooth | BLE 5.0 |
| Camera | Sense camera expansion |
| Microphone | Digital microphone |
| Storage | microSD |
| USB | USB-C |
The board itself measures roughly:
21 Γ 17.8 mm
making it particularly suitable for compact embedded vision projects.
Official documentation:
π XIAO ESP32S3 Getting Started
π XIAO ESP32S3 Sense Camera Usage
Main firmware:
π XiaoESP32S3SenseDoorbell.ino
The camera is configured as:
#define CAMERA_MODEL_XIAO_ESP32S3
#define EI_CAMERA_RAW_FRAME_BUFFER_COLS 320
#define EI_CAMERA_RAW_FRAME_BUFFER_ROWS 240
#define EI_CAMERA_FRAME_BYTE_SIZE 3Capture configuration:
.pixel_format = PIXFORMAT_JPEG,
.frame_size = FRAMESIZE_QVGA,
.jpeg_quality = 12,
.fb_count = 1,
.fb_location = CAMERA_FB_IN_PSRAMCamera clock:
.xclk_freq_hz = 20000000or:
20 MHz
The source uses the XIAO ESP32S3 Sense camera mapping:
#define XCLK_GPIO_NUM 10
#define SIOD_GPIO_NUM 40
#define SIOC_GPIO_NUM 39
#define Y9_GPIO_NUM 48
#define Y8_GPIO_NUM 11
#define Y7_GPIO_NUM 12
#define Y6_GPIO_NUM 14
#define Y5_GPIO_NUM 16
#define Y4_GPIO_NUM 18
#define Y3_GPIO_NUM 17
#define Y2_GPIO_NUM 15
#define VSYNC_GPIO_NUM 38
#define HREF_GPIO_NUM 47
#define PCLK_GPIO_NUM 13These pins correspond to the Sense expansion-board camera interface.
Current reference:
π Seeed Camera Interface Documentation
Each loop allocates an RGB frame buffer:
snapshot_buf =
(uint8_t*)malloc(
EI_CAMERA_RAW_FRAME_BUFFER_COLS *
EI_CAMERA_RAW_FRAME_BUFFER_ROWS *
EI_CAMERA_FRAME_BYTE_SIZE
);The camera first captures a JPEG.
That JPEG is then:
Camera
β
JPEG 320Γ240
β
Save to microSD
β
Convert JPEG β RGB888
β
Resize / crop
β
Edge Impulse input
Conversion:
fmt2rgb888(
fb->buf,
fb->len,
PIXFORMAT_JPEG,
snapshot_buf
);If the model input size differs from QVGA:
ei::image::processing::
crop_and_interpolate_rgb888(...)resizes the frame for inference.
Every captured camera image is stored on microSD.
The filenames are generated as:
sprintf(
filename,
"/image%d.jpg",
myCounter
);producing:
/image0.jpg
/image1.jpg
/image2.jpg
/image3.jpg
...
The file is written before inference:
writeFile(
SD,
filename,
fb->buf,
fb->len
);
myCounter++;This means the microSD also acts as a chronological image archive.
Initialization:
if (!SD.begin(21)) {
Serial.println(
"Card mount failed"
);
}The code retries once if the first mount fails.
Current Seeed documentation recommends:
microSD β€ 32 GB
FAT32
π XIAO ESP32S3 Sense microSD / Camera Guide
The original build used an:
8 GB microSD card
The repository includes the exported model:
π ei-face-detection-arduino-1.0.19.zip
Archive size:
4.33 MB
The sketch loads it through:
#include <Face_detection_inferencing.h>and executes:
run_classifier(
&signal,
&result,
debug_nn
);The model uses FOMO β Faster Objects, More Objects, Edge Impulse's lightweight object-detection architecture designed for constrained devices.
Instead of merely classifying the whole image as:
face
or
no face
FOMO can return multiple detections and their locations.
Conceptually:
Camera frame
β
96Γ96 model input
β
Image processing
β
FOMO
β
face #1
face #2
...
Current documentation:
π Edge Impulse FOMO
The original project documents the following retraining workflow:
~400 face images
β
Bounding-box annotations
label = face
β
96 Γ 96 image impulse
β
Image processing
β
FOMO object detection
β
70 training cycles
learning rate 0.00015
β
Arduino Library
Suggested settings:
| Parameter | Value |
|---|---|
| Approx. images | 400 |
| Input | 96 Γ 96 |
| Label | face |
| Labeling | Bounding boxes |
| Learning block | Object Detection |
| Algorithm | FOMO |
| Training cycles | 70 |
| Learning rate | 0.00015 |
Full original instructions:
π Machine Learning Doorbell β Hackster.io
The original build requires a specific change in the generated Edge Impulse Arduino library.
Open:
Arduino/libraries/
Face_detection_inferencing/
src/edge-impulse-sdk/
classifier/
ei_classifier_config.h
Find:
#define EI_CLASSIFIER_TFLITE_ENABLE_ESP_NN 1and change it to:
#define EI_CLASSIFIER_TFLITE_ENABLE_ESP_NN 0This configuration is explicitly documented by the original project for this XIAO ESP32S3 Sense deployment.
Install the Espressif Arduino core using:
https://raw.githubusercontent.com/espressif/arduino-esp32/gh-pages/package_esp32_dev_index.json
Then select:
Board:
XIAO_ESP32S3
and enable:
PSRAM:
OPI PSRAM
PSRAM is important because the code places the camera frame buffer in external PSRAM:
.fb_location =
CAMERA_FB_IN_PSRAMThe firmware iterates through:
result.bounding_boxesFor each non-zero detection:
ei_printf(
"%s (%f) "
"[ x: %u, y: %u, "
"width: %u, height: %u ]",
bb.label,
bb.value,
bb.x,
bb.y,
bb.width,
bb.height
);Example:
face (0.87)
[x: 24, y: 16, width: 8, height: 8]
The source can therefore process more than one detected face in the same inference result.
The project uses:
π Universal Arduino Telegram Bot
Includes:
#include <WiFiClientSecure.h>
#include <UniversalTelegramBot.h>Initialization:
WiFiClientSecure secured_client;
UniversalTelegramBot bot(
BOT_TOKEN,
secured_client
);Create a bot using:
π Telegram BotFather
Then configure:
#define BOT_TOKEN ""
String chat_id = "";The chat can be:
Personal chat
or
Family / household group
The original project used a family Telegram group.
After a detection, the most recently saved image is reopened:
myFile = SD.open(
filename
);and streamed directly to Telegram:
bot.sendPhotoByBinary(
chat_id,
"image/jpeg",
myFile.size(),
isMoreDataAvailable,
getNextByte,
nullptr,
nullptr
);The SD file is therefore used as the data source for the Telegram upload.
The accompanying notification is:
bot.sendMessage(
chat_id,
"There is someone at the door " +
String(bb.value) +
"% - Powered by XIAO ESP32S3 Sense"
);The bb.value variable is the model's detection score.
Configure:
#define WIFI_SSID ""
#define WIFI_PASSWORD ""The board connects during startup:
WiFi.begin(
WIFI_SSID,
WIFI_PASSWORD
);and waits until:
WiFi.status() ==
WL_CONNECTEDThe local IP address is then printed to Serial.
The onboard LED is configured as:
pinMode(
LED_BUILTIN,
OUTPUT
);Helper functions:
void lightOn() {
digitalWrite(
LED_BUILTIN,
HIGH
);
}
void lightOff() {
digitalWrite(
LED_BUILTIN,
LOW
);
}The LED provides immediate local feedback while a detection is being handled.
The configurable interval is:
int delayAfterDetection =
10000;or:
10 seconds
After sending a detection notification:
ei_sleep(
delayAfterDetection
);pauses the detection workflow.
flowchart TD
START["Power On"]
CAM["π· Initialize Camera"]
WIFI["π‘ Connect Wi-Fi"]
SD["πΎ Mount microSD"]
CAPTURE["Capture QVGA JPEG"]
SAVE["Save /imageN.jpg"]
RGB["Convert JPEG β RGB888"]
ML["π§ FOMO Inference"]
FACE{"Face found?"}
NEXT["Capture Next Frame"]
LED["π‘ LED On"]
PHOTO["π² Send JPEG"]
MSG["π Send Detection Score"]
WAIT["Wait 10 seconds"]
START --> CAM
CAM --> WIFI
WIFI --> SD
SD --> CAPTURE
CAPTURE --> SAVE
SAVE --> RGB
RGB --> ML
ML --> FACE
FACE -->|"No"| NEXT
NEXT --> CAPTURE
FACE -->|"Yes"| LED
LED --> PHOTO
PHOTO --> MSG
MSG --> WAIT
WAIT --> CAPTURE
The original prototype deliberately focuses on:
Detection
+
Image capture
+
Remote notification
rather than reproducing the mechanical chime.
The original project notes several possible extensions:
Local buzzer
Relay + bell
Remote ESP32 sounder
Bluetooth-connected sounder
This keeps the first version centered on the Computer Vision experiment.
| Component | Quantity |
|---|---|
| XIAO ESP32S3 Sense | 1 |
| Sense camera module | 1 |
| U.FL Wi-Fi antenna | 1 |
| microSD card | 1 |
| USB-C cable / power supply | 1 |
Original microSD:
8 GB
FAT32
No separate:
PIR
button
Raspberry Pi
computer
external camera
is required after deployment.
The sketch uses:
#include <Face_detection_inferencing.h>
#include "edge-impulse-sdk/dsp/image/image.hpp"
#include "esp_camera.h"
#include "FS.h"
#include "SD.h"
#include "SPI.h"
#include <WiFi.h>
#include <WiFiClientSecure.h>
#include <UniversalTelegramBot.h>Install:
π ei-face-detection-arduino-1.0.19.zip
π witnessmenow/Universal-Arduino-Telegram-Bot
The Telegram library also requires:
π ArduinoJson
Camera, Wi-Fi, SD and SPI support come from the ESP32 Arduino platform.
git clone \
https://github.com/ronibandini/xiaoesp32s3Sensedoorbell.git
cd xiaoesp32s3SensedoorbellRepository:
π github.com/ronibandini/xiaoesp32s3Sensedoorbell
In Arduino IDE:
File
β Preferences
β Additional Boards Manager URLs
add:
https://raw.githubusercontent.com/espressif/arduino-esp32/gh-pages/package_esp32_dev_index.json
Then install the ESP32 board package.
Choose:
XIAO_ESP32S3
and enable:
OPI PSRAM
Install through Library Manager:
UniversalTelegramBot
ArduinoJson
Add:
π ei-face-detection-arduino-1.0.19.zip
with:
Sketch
β Include Library
β Add .ZIP Library
Then apply the documented ESP-NN setting:
#define EI_CLASSIFIER_TFLITE_ENABLE_ESP_NN 0Format as:
FAT32
and insert it into the Sense expansion board.
Open:
π XiaoESP32S3SenseDoorbell.ino
Set:
#define BOT_TOKEN ""
String chat_id = "";
#define WIFI_SSID ""
#define WIFI_PASSWORD ""Upload the firmware and open Serial Monitor at:
115200 baud
Startup output includes:
ESP32S3 Sense Machine Learning
doorbell started
Camera initialized
Connecting to Wifi
WiFi connected.
IP address: ...
Device ready...
xiaoesp32s3Sensedoorbell/
β
βββ XiaoESP32S3SenseDoorbell.ino
βββ ei-face-detection-arduino-1.0.19.zip
βββ README.md
- πͺ
XiaoESP32S3SenseDoorbell.inoβ complete camera, inference, SD and Telegram application - π§
ei-face-detection-arduino-1.0.19.zipβ compiled Edge Impulse face detector - π
README.mdβ original short project notes
Complete original build published in August 2023, including the hardware, Edge Impulse model workflow, Telegram configuration and deployment process.
π Machine Learning Doorbell with XIAO ESP32S3 Sense
Seeed Studio published a dedicated feature about the project:
π XIAO ESP32S3 Sense-Powered Doorbell Enhanced with Machine Learning
The project was also selected for the company's August community roundup:
π Seeed Monthly Wrap-up for August
Official board documentation:
π Getting Started with XIAO ESP32S3
Camera and microSD:
π Camera Usage for XIAO ESP32S3 Sense
Current FOMO documentation:
π FOMO β Faster Objects, More Objects
End-to-end tutorial:
π Object Detection with Centroids
Arduino Telegram client used by the project:
π Universal Arduino Telegram Bot
Bot creation:
Later-generation ESP32-S3 AI doorbell combining Computer Vision, audio, Telegram and AI-agent interactions.
π github.com/ronibandini/aicamdoorbell
ESP32-S3 camera project combining local Computer Vision, microphone recording and AI processing.
π github.com/ronibandini/guampapp
Edge Impulse Computer Vision object detection connected to a physical control system.
π github.com/ronibandini/TIAM62AITrafficLight
Computer Vision anomaly detection with Edge Impulse and automatic physical part rejection.
π github.com/ronibandini/visualAnomalyGroveV2
Earlier Arduino doorbell experiment using non-contact infrared temperature sensing and audio alerts.
π github.com/ronibandini/CoronavirusDoorbell
Contracultura Maker is a book by Roni Bandini about maker culture, experimental electronics, AI, physical computing and technological autonomy.
π Contracultura Maker β GitHub repository
π Download Contracultura Maker PDF
Roni Bandini Maker Β· AI Developer Β· Writer Buenos Aires, Argentina
- π GitHub β @ronibandini
- π Medium β @ronibandini
- π X / Twitter β @RoniBandini
- πΈ Instagram β @ronibandini
βΆοΈ YouTube β @RoniBandini- πΌ LinkedIn β Roni Bandini
Built with πͺ + ESP32-S3 + FOMO + Telegram.