Skip to content

About

Machine Learning Doorbell with Seeed XIAO ESP32S3 Sense

Resources

Stars

2 stars

Watchers

1 watching

Forks

Latest commit

Β 

History

5 Commits

Folders and files

Repository files navigation

πŸšͺπŸ“· Machine Learning Doorbell

Buttonless face-detection doorbell with XIAO ESP32S3 Sense, Edge Impulse FOMO and Telegram

Seeed Studio Edge Impulse ESP32 Telegram License: MIT

Roni Bandini β€” Buenos Aires, Argentina β€” August 2023

Machine Learning Doorbell is a compact Computer Vision experiment that replaces the physical push button of a conventional doorbell with automatic face detection.

A Seeed Studio XIAO ESP32S3 Sense continuously captures images and runs an Edge Impulse FOMO object-detection model locally.

When one or more faces are detected, the device:

  1. πŸ“· Captures the current camera image.
  2. 🧠 Runs the Edge Impulse model.
  3. πŸ’Ύ Saves the JPEG to microSD.
  4. πŸ’‘ Activates the onboard LED.
  5. πŸ“² Sends the image to a Telegram chat.
  6. πŸ“ Sends a text notification containing the detection confidence.
  7. ⏱️ Waits before resuming detection.

No mechanical doorbell button is required.


✨ Features

  • πŸšͺ Buttonless doorbell concept
  • πŸ“· Integrated camera
  • 🧠 On-device Computer Vision inference
  • 🎯 Edge Impulse FOMO face detection
  • πŸ‘₯ Supports multiple detections in one frame
  • πŸ’Ύ Automatic JPEG storage to microSD
  • πŸ“² Telegram image notifications
  • πŸ“Š Detection confidence in the message
  • πŸ’‘ Onboard LED feedback
  • πŸ“‘ Wi-Fi connectivity
  • ⚑ ESP32-S3 dual-core processor
  • 🧠 8 MB PSRAM
  • πŸ“¦ Edge Impulse Arduino library included
  • πŸ”Œ Single compact board + Sense expansion module

πŸ—οΈ Architecture

flowchart LR
    VISITOR["πŸ‘€ Visitor"]
    CAMERA["πŸ“· XIAO Camera"]
    ESP["ESP32-S3"]
    FRAME["320Γ—240 JPEG"]
    EI["🧠 Edge Impulse<br/>FOMO"]
    FACE{"Face detected?"}
    SD["πŸ’Ύ microSD"]
    LED["πŸ’‘ LED"]
    WIFI["πŸ“‘ Wi-Fi"]
    TG["πŸ“² Telegram"]

    VISITOR --> CAMERA
    CAMERA --> ESP
    ESP --> FRAME

    FRAME --> SD
    FRAME --> EI

    EI --> FACE

    FACE -->|"Yes"| LED
    FACE -->|"Yes"| WIFI
    WIFI --> TG
Loading

Image capture, preprocessing and ML inference run directly on the ESP32-S3.

Telegram is used only for notification delivery.


🧠 XIAO ESP32S3 Sense

The project uses the Seeed Studio XIAO ESP32S3 Sense.

Current hardware specifications include:

Feature Specification
MCU ESP32-S3R8
CPU Dual-core Xtensa LX7
Clock Up to 240 MHz
PSRAM 8 MB
Flash 8 MB
Wi-Fi 2.4 GHz
Bluetooth BLE 5.0
Camera Sense camera expansion
Microphone Digital microphone
Storage microSD
USB USB-C

The board itself measures roughly:

21 Γ— 17.8 mm

making it particularly suitable for compact embedded vision projects.

Official documentation:

πŸ‘‰ XIAO ESP32S3 Getting Started

πŸ‘‰ XIAO ESP32S3 Sense Camera Usage


πŸ“· Camera Configuration

Main firmware:

πŸ‘‰ XiaoESP32S3SenseDoorbell.ino

The camera is configured as:

#define CAMERA_MODEL_XIAO_ESP32S3

#define EI_CAMERA_RAW_FRAME_BUFFER_COLS 320
#define EI_CAMERA_RAW_FRAME_BUFFER_ROWS 240
#define EI_CAMERA_FRAME_BYTE_SIZE       3

Capture configuration:

.pixel_format = PIXFORMAT_JPEG,
.frame_size   = FRAMESIZE_QVGA,
.jpeg_quality = 12,
.fb_count     = 1,
.fb_location  = CAMERA_FB_IN_PSRAM

Camera clock:

.xclk_freq_hz = 20000000

or:

20 MHz

πŸ“Œ Camera Pin Mapping

The source uses the XIAO ESP32S3 Sense camera mapping:

#define XCLK_GPIO_NUM  10
#define SIOD_GPIO_NUM  40
#define SIOC_GPIO_NUM  39

#define Y9_GPIO_NUM    48
#define Y8_GPIO_NUM    11
#define Y7_GPIO_NUM    12
#define Y6_GPIO_NUM    14
#define Y5_GPIO_NUM    16
#define Y4_GPIO_NUM    18
#define Y3_GPIO_NUM    17
#define Y2_GPIO_NUM    15

#define VSYNC_GPIO_NUM 38
#define HREF_GPIO_NUM  47
#define PCLK_GPIO_NUM  13

These pins correspond to the Sense expansion-board camera interface.

Current reference:

πŸ‘‰ Seeed Camera Interface Documentation


πŸ–ΌοΈ Image Capture Pipeline

Each loop allocates an RGB frame buffer:

snapshot_buf =
    (uint8_t*)malloc(
        EI_CAMERA_RAW_FRAME_BUFFER_COLS *
        EI_CAMERA_RAW_FRAME_BUFFER_ROWS *
        EI_CAMERA_FRAME_BYTE_SIZE
    );

The camera first captures a JPEG.

That JPEG is then:

Camera
  ↓
JPEG 320Γ—240
  ↓
Save to microSD
  ↓
Convert JPEG β†’ RGB888
  ↓
Resize / crop
  ↓
Edge Impulse input

Conversion:

fmt2rgb888(
    fb->buf,
    fb->len,
    PIXFORMAT_JPEG,
    snapshot_buf
);

If the model input size differs from QVGA:

ei::image::processing::
crop_and_interpolate_rgb888(...)

resizes the frame for inference.


πŸ’Ύ Automatic Photo Storage

Every captured camera image is stored on microSD.

The filenames are generated as:

sprintf(
    filename,
    "/image%d.jpg",
    myCounter
);

producing:

/image0.jpg
/image1.jpg
/image2.jpg
/image3.jpg
...

The file is written before inference:

writeFile(
    SD,
    filename,
    fb->buf,
    fb->len
);

myCounter++;

This means the microSD also acts as a chronological image archive.


πŸ’Ύ microSD

Initialization:

if (!SD.begin(21)) {
    Serial.println(
        "Card mount failed"
    );
}

The code retries once if the first mount fails.

Current Seeed documentation recommends:

microSD ≀ 32 GB
FAT32

πŸ‘‰ XIAO ESP32S3 Sense microSD / Camera Guide

The original build used an:

8 GB microSD card

🧠 Edge Impulse Model

The repository includes the exported model:

πŸ‘‰ ei-face-detection-arduino-1.0.19.zip

Archive size:

4.33 MB

The sketch loads it through:

#include <Face_detection_inferencing.h>

and executes:

run_classifier(
    &signal,
    &result,
    debug_nn
);

🎯 FOMO Object Detection

The model uses FOMO β€” Faster Objects, More Objects, Edge Impulse's lightweight object-detection architecture designed for constrained devices.

Instead of merely classifying the whole image as:

face
or
no face

FOMO can return multiple detections and their locations.

Conceptually:

Camera frame
     ↓
96Γ—96 model input
     ↓
Image processing
     ↓
FOMO
     ↓
face #1
face #2
...

Current documentation:

πŸ‘‰ Edge Impulse FOMO


πŸ§ͺ Original Training Recipe

The original project documents the following retraining workflow:

~400 face images
        ↓
Bounding-box annotations
label = face
        ↓
96 Γ— 96 image impulse
        ↓
Image processing
        ↓
FOMO object detection
        ↓
70 training cycles
learning rate 0.00015
        ↓
Arduino Library

Suggested settings:

Parameter Value
Approx. images 400
Input 96 Γ— 96
Label face
Labeling Bounding boxes
Learning block Object Detection
Algorithm FOMO
Training cycles 70
Learning rate 0.00015

Full original instructions:

πŸ‘‰ Machine Learning Doorbell β€” Hackster.io


βš™οΈ ESP32-S3 / Edge Impulse Setting

The original build requires a specific change in the generated Edge Impulse Arduino library.

Open:

Arduino/libraries/
Face_detection_inferencing/
src/edge-impulse-sdk/
classifier/
ei_classifier_config.h

Find:

#define EI_CLASSIFIER_TFLITE_ENABLE_ESP_NN 1

and change it to:

#define EI_CLASSIFIER_TFLITE_ENABLE_ESP_NN 0

This configuration is explicitly documented by the original project for this XIAO ESP32S3 Sense deployment.


πŸ“¦ Arduino Board Configuration

Install the Espressif Arduino core using:

https://raw.githubusercontent.com/espressif/arduino-esp32/gh-pages/package_esp32_dev_index.json

Then select:

Board:
XIAO_ESP32S3

and enable:

PSRAM:
OPI PSRAM

PSRAM is important because the code places the camera frame buffer in external PSRAM:

.fb_location =
    CAMERA_FB_IN_PSRAM

πŸ“Š Detection Output

The firmware iterates through:

result.bounding_boxes

For each non-zero detection:

ei_printf(
    "%s (%f) "
    "[ x: %u, y: %u, "
    "width: %u, height: %u ]",
    bb.label,
    bb.value,
    bb.x,
    bb.y,
    bb.width,
    bb.height
);

Example:

face (0.87)
[x: 24, y: 16, width: 8, height: 8]

The source can therefore process more than one detected face in the same inference result.


πŸ“² Telegram Notifications

The project uses:

πŸ‘‰ Universal Arduino Telegram Bot

Includes:

#include <WiFiClientSecure.h>
#include <UniversalTelegramBot.h>

Initialization:

WiFiClientSecure secured_client;

UniversalTelegramBot bot(
    BOT_TOKEN,
    secured_client
);

πŸ”‘ Telegram Configuration

Create a bot using:

πŸ‘‰ Telegram BotFather

Then configure:

#define BOT_TOKEN ""

String chat_id = "";

The chat can be:

Personal chat
or
Family / household group

The original project used a family Telegram group.


πŸ“· Sending the Photo

After a detection, the most recently saved image is reopened:

myFile = SD.open(
    filename
);

and streamed directly to Telegram:

bot.sendPhotoByBinary(
    chat_id,
    "image/jpeg",
    myFile.size(),
    isMoreDataAvailable,
    getNextByte,
    nullptr,
    nullptr
);

The SD file is therefore used as the data source for the Telegram upload.


πŸ“ Detection Message

The accompanying notification is:

bot.sendMessage(
    chat_id,
    "There is someone at the door " +
    String(bb.value) +
    "% - Powered by XIAO ESP32S3 Sense"
);

The bb.value variable is the model's detection score.


πŸ“‘ Wi-Fi

Configure:

#define WIFI_SSID ""
#define WIFI_PASSWORD ""

The board connects during startup:

WiFi.begin(
    WIFI_SSID,
    WIFI_PASSWORD
);

and waits until:

WiFi.status() ==
    WL_CONNECTED

The local IP address is then printed to Serial.


πŸ’‘ LED Feedback

The onboard LED is configured as:

pinMode(
    LED_BUILTIN,
    OUTPUT
);

Helper functions:

void lightOn() {
    digitalWrite(
        LED_BUILTIN,
        HIGH
    );
}

void lightOff() {
    digitalWrite(
        LED_BUILTIN,
        LOW
    );
}

The LED provides immediate local feedback while a detection is being handled.


⏱️ Detection Delay

The configurable interval is:

int delayAfterDetection =
    10000;

or:

10 seconds

After sending a detection notification:

ei_sleep(
    delayAfterDetection
);

pauses the detection workflow.


πŸ”„ Runtime Flow

flowchart TD
    START["Power On"]
    CAM["πŸ“· Initialize Camera"]
    WIFI["πŸ“‘ Connect Wi-Fi"]
    SD["πŸ’Ύ Mount microSD"]
    CAPTURE["Capture QVGA JPEG"]
    SAVE["Save /imageN.jpg"]
    RGB["Convert JPEG β†’ RGB888"]
    ML["🧠 FOMO Inference"]
    FACE{"Face found?"}
    NEXT["Capture Next Frame"]
    LED["πŸ’‘ LED On"]
    PHOTO["πŸ“² Send JPEG"]
    MSG["πŸ“ Send Detection Score"]
    WAIT["Wait 10 seconds"]

    START --> CAM
    CAM --> WIFI
    WIFI --> SD

    SD --> CAPTURE
    CAPTURE --> SAVE
    SAVE --> RGB
    RGB --> ML
    ML --> FACE

    FACE -->|"No"| NEXT
    NEXT --> CAPTURE

    FACE -->|"Yes"| LED
    LED --> PHOTO
    PHOTO --> MSG
    MSG --> WAIT
    WAIT --> CAPTURE
Loading

πŸ”Š Why There Is No Physical Bell

The original prototype deliberately focuses on:

Detection
+
Image capture
+
Remote notification

rather than reproducing the mechanical chime.

The original project notes several possible extensions:

Local buzzer
Relay + bell
Remote ESP32 sounder
Bluetooth-connected sounder

This keeps the first version centered on the Computer Vision experiment.


πŸ› οΈ Hardware

Component Quantity
XIAO ESP32S3 Sense 1
Sense camera module 1
U.FL Wi-Fi antenna 1
microSD card 1
USB-C cable / power supply 1

Original microSD:

8 GB
FAT32

No separate:

PIR
button
Raspberry Pi
computer
external camera

is required after deployment.


πŸ“š Required Libraries

The sketch uses:

#include <Face_detection_inferencing.h>
#include "edge-impulse-sdk/dsp/image/image.hpp"

#include "esp_camera.h"

#include "FS.h"
#include "SD.h"
#include "SPI.h"

#include <WiFi.h>
#include <WiFiClientSecure.h>

#include <UniversalTelegramBot.h>

Install:

Edge Impulse Model

πŸ‘‰ ei-face-detection-arduino-1.0.19.zip

Universal Telegram Bot

πŸ‘‰ witnessmenow/Universal-Arduino-Telegram-Bot

The Telegram library also requires:

πŸ‘‰ ArduinoJson

Camera, Wi-Fi, SD and SPI support come from the ESP32 Arduino platform.


πŸš€ Installation

1. Clone the Repository

git clone \
https://github.com/ronibandini/xiaoesp32s3Sensedoorbell.git

cd xiaoesp32s3Sensedoorbell

Repository:

πŸ‘‰ github.com/ronibandini/xiaoesp32s3Sensedoorbell


2. Install ESP32 Support

In Arduino IDE:

File
β†’ Preferences
β†’ Additional Boards Manager URLs

add:

https://raw.githubusercontent.com/espressif/arduino-esp32/gh-pages/package_esp32_dev_index.json

Then install the ESP32 board package.


3. Select the Board

Choose:

XIAO_ESP32S3

and enable:

OPI PSRAM

4. Install Telegram Libraries

Install through Library Manager:

UniversalTelegramBot
ArduinoJson

5. Install the ML Model

Add:

πŸ‘‰ ei-face-detection-arduino-1.0.19.zip

with:

Sketch
β†’ Include Library
β†’ Add .ZIP Library

Then apply the documented ESP-NN setting:

#define EI_CLASSIFIER_TFLITE_ENABLE_ESP_NN 0

6. Prepare microSD

Format as:

FAT32

and insert it into the Sense expansion board.


7. Configure Credentials

Open:

πŸ‘‰ XiaoESP32S3SenseDoorbell.ino

Set:

#define BOT_TOKEN ""
String chat_id = "";

#define WIFI_SSID ""
#define WIFI_PASSWORD ""

8. Upload

Upload the firmware and open Serial Monitor at:

115200 baud

Startup output includes:

ESP32S3 Sense Machine Learning
doorbell started

Camera initialized

Connecting to Wifi

WiFi connected.
IP address: ...

Device ready...

πŸ“ Repository Structure

xiaoesp32s3Sensedoorbell/
β”‚
β”œβ”€β”€ XiaoESP32S3SenseDoorbell.ino
β”œβ”€β”€ ei-face-detection-arduino-1.0.19.zip
└── README.md

🌐 External References

πŸ› οΈ Hackster.io

Complete original build published in August 2023, including the hardware, Edge Impulse model workflow, Telegram configuration and deployment process.

πŸ‘‰ Machine Learning Doorbell with XIAO ESP32S3 Sense


🌱 Seeed Studio

Seeed Studio published a dedicated feature about the project:

πŸ‘‰ XIAO ESP32S3 Sense-Powered Doorbell Enhanced with Machine Learning

The project was also selected for the company's August community roundup:

πŸ‘‰ Seeed Monthly Wrap-up for August


πŸ“· XIAO ESP32S3 Sense

Official board documentation:

πŸ‘‰ Getting Started with XIAO ESP32S3

Camera and microSD:

πŸ‘‰ Camera Usage for XIAO ESP32S3 Sense


🧠 Edge Impulse FOMO

Current FOMO documentation:

πŸ‘‰ FOMO β€” Faster Objects, More Objects

End-to-end tutorial:

πŸ‘‰ Object Detection with Centroids


πŸ“² Telegram

Arduino Telegram client used by the project:

πŸ‘‰ Universal Arduino Telegram Bot

Bot creation:

πŸ‘‰ Telegram Bot Tutorial


πŸ”— Related GitHub Projects

πŸ”” AI Bell

Later-generation ESP32-S3 AI doorbell combining Computer Vision, audio, Telegram and AI-agent interactions.

πŸ‘‰ github.com/ronibandini/aicamdoorbell

πŸ‘οΈ GuampApp

ESP32-S3 camera project combining local Computer Vision, microphone recording and AI processing.

πŸ‘‰ github.com/ronibandini/guampapp

🚦 TI AM62A AI Traffic Light

Edge Impulse Computer Vision object detection connected to a physical control system.

πŸ‘‰ github.com/ronibandini/TIAM62AITrafficLight

βš™οΈ Visual Anomaly β€” Grove Vision AI V2

Computer Vision anomaly detection with Edge Impulse and automatic physical part rejection.

πŸ‘‰ github.com/ronibandini/visualAnomalyGroveV2

🌑️ Coronavirus Doorbell

Earlier Arduino doorbell experiment using non-contact infrared temperature sensing and audio alerts.

πŸ‘‰ github.com/ronibandini/CoronavirusDoorbell


πŸ“• Contracultura Maker

Contracultura Maker is a book by Roni Bandini about maker culture, experimental electronics, AI, physical computing and technological autonomy.

πŸ“‚ Contracultura Maker β€” GitHub repository

πŸ“• Download Contracultura Maker PDF


πŸ“¬ Contact

Roni Bandini Maker Β· AI Developer Β· Writer Buenos Aires, Argentina


Built with πŸšͺ + ESP32-S3 + FOMO + Telegram.

About

Machine Learning Doorbell with Seeed XIAO ESP32S3 Sense

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages