Note
This directory contains Ansible playbooks and comprehensive documentation for deploying and managing the Forward Email infrastructure.
ansible/
├── README.md # This file - Complete documentation index
├── docs/ # Documentation guides
│ ├── MONITORING.md # Security monitoring system guide
│ ├── MONITORING_TESTING.md # Comprehensive monitoring testing guide
│ ├── SYSTEMD_STATUS_INCIDENTS.md # Private-by-design public incident bridge
│ ├── PM2_MONITORING.md # PM2 health monitoring guide
│ ├── SYSTEM_OPTIMIZATION.md # System-wide optimization (tmpfs, mount options)
│ ├── AMD_RYZEN_NUMA.md # AMD Ryzen NUMA optimization guide
│ ├── SYSCTL_HIGH_TRAFFIC.md # High-traffic server kernel tuning guide
│ ├── IO_FILESYSTEM_TUNING.md # I/O scheduler and filesystem optimization guide
│ ├── UFW_ALLOWLIST.md # UFW IP allowlist management guide
│ ├── REMOVE_POSTFIX.md # Postfix removal guide (cleanup tool)
│ ├── README_MONGO_REDIS.md # MongoDB & Redis/Valkey deployment
│ ├── MONGODB_OPERATIONS_GUIDE.md
│ ├── MONGODB_PERFORMANCE_TUNING.md
│ ├── REDIS_PERFORMANCE_TUNING.md
│ ├── DISASTER_RECOVERY.md
│ ├── SERVICE_USER_AUDIT.md
│ └── LSYNCD_STORAGE_MIRRORING.md # Real-time storage mirroring guide
├── playbooks/ # Ansible playbooks
│ ├── security.yml # Security baseline, monitoring, and send-only alerts
│ ├── node.yml # Node.js & PM2 deployment
│ ├── chrony-timesync.yml # Reusable time synchronization (chrony)
│ ├── system-optimization.yml # Reusable system-wide optimization (tmpfs, mount options)
│ ├── amd-ryzen-numa.yml # Reusable AMD Ryzen NUMA optimization
│ ├── sysctl-high-traffic.yml # Reusable high-traffic server kernel tuning
│ ├── io-filesystem-tuning.yml # Reusable I/O scheduler and filesystem optimization
│ ├── ufw-allowlist.yml # Reusable UFW IP allowlist management
│ ├── remove-postfix.yml # Remove Postfix from servers (cleanup tool)
│ ├── mongo.yml # MongoDB deployment
│ ├── logs.yml # Logs MongoDB deployment (separate instance)
│ ├── redis.yml # Redis/Valkey deployment
│ ├── bree.yml # Bree job scheduler
│ ├── http.yml # HTTP/API servers
│ ├── smtp.yml # SMTP server
│ ├── imap.yml # IMAP server
│ ├── pop3.yml # POP3 server
│ ├── mx1.yml # MX1 mail exchanger
│ ├── mx2.yml # MX2 mail exchanger
│ ├── unbound.yml # Unbound DNS resolver
│ ├── sqlite.yml # SQLite server
│ ├── sqlite-mirror.yml # Storage mirroring (cron rsync)
│ ├── certificates.yml # SSL/TLS certificates
│ ├── dkim.yml # DKIM key deployment
│ ├── env.yml # Environment variables
│ ├── ecosystem.yml # PM2 ecosystem config
│ ├── fonts.yml # Font deployment
│ ├── gapp-creds.yml # Google App credentials
│ ├── gpg-security-key.yml # GPG security keys
│ ├── ssh-keys.yml # SSH key deployment
│ ├── deployment-keys.yml # Deployment keys
│ └── patch-dns-role.yml # DNS role patches
└── requirements.yml # Ansible collections and roles
- Getting Started
- Deployment Guides
- System-Wide Optimization
- AMD Ryzen NUMA Optimization
- UFW IP Allowlist Management
- High-Traffic Server Kernel Tuning
- Monitoring & Alerting
- Operations & Maintenance
- Performance Tuning
- Real-time Storage Mirroring
- Disaster Recovery
- Security & Auditing
- Common Commands
- Related Resources
Important
Before deploying any services, ensure you have:
- Ansible 2.9+ installed
- SSH access to target servers
- Required environment variables configured
- SSL/TLS certificates ready
# Install Ansible
pip install ansibleNote
No Ansible Galaxy dependencies required! Our playbooks use custom installations:
- MongoDB 6.0.27: Installed directly from official MongoDB repository
- Valkey: Compiled from source
This eliminates external role dependencies and gives us full control.
Infrastructure alerts use a host-local, send-only Postfix queue. Postfix delivers directly to each recipient domain's MX over outbound TCP port 25. It does not use credentials, the Forward Email SMTP service, the application mail queue, or MongoDB.
Configure the alert sender and recipients before deployment:
export ALERT_EMAIL_FROM=mailerdaemon@forwardemail.net
export ALERT_EMAIL_RECIPIENTS=security@forwardemail.netMSMTP_RCPTS remains a deprecated fallback for ALERT_EMAIL_RECIPIENTS during migration. MSMTP_USERNAME, MSMTP_PASSWORD, SMTP_HOST, and SMTP_PORT are not used by the alert transport.
The security.yml playbook fails closed unless SPF passes for both the envelope sender and the host's HELO identity. Before rollout:
-
Publish one logical TXT record at
_spf-alerts.forwardemail.netwith the approved IPv4 egress addresses:v=spf1 ip4:121.127.44.61 ip4:121.127.44.68 ip4:121.127.44.70 ip4:121.127.44.71 ip4:121.127.44.72 ip4:121.127.44.75 ip4:121.127.44.86 ip4:121.127.44.92 ip4:121.127.44.99 ip4:121.127.44.103 -all -
Add
include:_spf-alerts.forwardemail.netto the existingforwardemail.netSPF record somailerdaemon@forwardemail.netpasses from those hosts. Keep exactly one SPF record for the domain. -
Publish
v=spf1 a -allat every configured sending HELO hostname. This includes the database host variables and every Node.js host variable:BREE_HOST,WEB_HOST,API_HOST,CALDAV_HOST,CARDDAV_HOST,IMAP_HOST,MX1_HOST,MX2_HOST,POP3_HOST,SMTP_HOST, andSQLITE_HOST. -
Confirm outbound TCP port 25 is allowed. In inventory, set
alert_public_ipwhenansible_hostis not the real egress IPv4.direct_alert_public_ipremains a temporary compatibility fallback.
Each service playbook sets the server hostname before security.yml runs. Postfix uses that hostname for HELO, and deployment stops if the configured hostname does not match the live server hostname. This prevents one server from advertising another server's name and avoids using a provider PTR name by mistake. Use alert_hostname only when it matches the server hostname; direct_alert_helo_identity is accepted only as a compatibility fallback.
The transport is IPv4-only. Revalidate the address list before every DNS change; do not weaken or bypass the Ansible SPF gates. The Postfix configuration disables every inet master service, accepts no TCP connections even on loopback, permits local queue submission only from root, has no relayhost or trusted client network, and disables local, virtual, and relay-domain delivery.
# 1. Deploy security baseline and monitoring
ansible-playbook ansible/playbooks/security.yml -i hosts.yml
# 2. Deploy Node.js and PM2
ansible-playbook ansible/playbooks/node.yml -i hosts.yml
# 3. Deploy MongoDB (primary)
ansible-playbook ansible/playbooks/mongo.yml -i hosts.yml
# 3b. Deploy Logs MongoDB (separate instance)
ansible-playbook ansible/playbooks/logs.yml -i hosts.yml
# 4. Deploy Redis/Valkey
ansible-playbook ansible/playbooks/redis.yml -i hosts.yml
MongoDB & Redis/Valkey Deployment Guide
Complete guide for deploying MongoDB v6 and Valkey (Redis fork) with:
- ✅ SSL/TLS encryption
- ✅ UFW firewall configuration
- ✅ Automated backups to Cloudflare R2
- ✅ Email alerting system
- ✅ Security hardening
Warning
MongoDB is LOCKED to v6.0.27 - Do not upgrade to v7 or v8 due to severe performance regressions. See MONGODB_OPERATIONS_GUIDE.md for details.
Tip
Start here if you're deploying database services for the first time.
Automated system-wide optimizations applied to ALL servers via security.yml:
- 🚀 tmpfs /tmp: RAM-based temporary storage (2GB, auto-configured)
- 🔒 /dev/shm hardening: Secured with noexec (1GB limit, blocks malware execution)
- 💾 Mount options: noatime, nodiratime, discard (TRIM for SSDs)
- 🔄 Automated fstab editing: Automatically updates /etc/fstab and remounts
- 🛡️ LUKS/LVM support: Works with all device formats (UUID, /dev/disk/by-id/, dm-uuid, etc.)
- ✅ Idempotent: Safe to run multiple times
Imported by: security.yml (applies to all servers)
Benefits:
- ⚡ Faster temporary file operations
- 🔒 Enhanced security (noexec on /dev/shm blocks malware)
- 📉 Reduced SSD wear (noatime, nodiratime)
- 🔄 Extended SSD lifespan (TRIM support)
- 🧹 Automatic /tmp cleanup on reboot
Note
All filesystem changes are fully automated - no manual intervention required.
AMD Ryzen NUMA Optimization Guide
Critical NUMA optimizations for AMD Ryzen/EPYC processors with automatic CPU detection:
- 🎯 zone_reclaim_mode=0: Prevents 10-100x tail latency spikes
- ⚖️ numa_balancing=0: Reduces latency variance by 20-50%
- 🚫 THP disabled: Eliminates 10-100ms THP-related stalls
- 🔍 Auto-detection: Only applies if AMD Ryzen/EPYC detected
Imported by: security.yml (applies to all servers with auto-detection)
Performance Impact:
- 10-100x reduction in tail latency
- 20-50% reduction in latency variance
- Elimination of THP-related latency spikes
Important
These optimizations are critical for database servers (MongoDB, Redis) running on AMD hardware.
UFW Allowlist Management Guide
Reusable UFW (Uncomplicated Firewall) IP allowlist management system for database and service security:
- 🔒 Automated IP allowlist updates - Fetches approved IPs from central source
- 🧹 Orphaned rule cleanup - Removes outdated rules from previous deployments
- 📧 Email notifications - Alerts for all changes with detailed reports
- 🔄 Retry logic - Network failure resilience (3 attempts, 10s timeout)
- ⏱️ Systemd timer automation - Updates every 10 minutes
- 🛡️ Safety features - Graceful failure handling, no connection drops
Integrated Services:
- 🔴 Redis/Valkey - Port 6380 (TLS)
- 🍃 MongoDB - Port 27017
- 🗄️ SQLite - Port 3456
Architecture:
# Reusable playbook with service-specific variables
- name: Import UFW allowlist playbook for Redis
import_playbook: ufw-allowlist.yml
vars:
target_hosts: redis
service_name: Redis
service_port_var: REDIS_PORT
service_port_default: "6380"
service_identifier: redis
ufw_comment: "Auto-whitelist Redis TLS"Benefits:
- ✅ Single source of truth - One playbook for all services
- ✅ Easy maintenance - Update once, applies everywhere
- ✅ Consistency guaranteed - Same logic across all services
- ✅ Extensible - Easy to add new services
Tip
The UFW allowlist system is automatically integrated into mongo.yml, redis.yml, and sqlite.yml playbooks.
Note
IP allowlists are fetched from https://forwardemail.net/ips/v4.txt and validated before application.
High-Traffic Server Kernel Tuning Guide
Kernel parameter optimizations for high-traffic servers (databases, APIs, mail servers):
- 🌐 BBR congestion control: Modern TCP congestion control algorithm
- 📊 Auto-scaled buffers: Automatically scales based on available RAM (256MB cap)
- 🔌 Connection tracking: Optimized for high connection counts
- 🚀 TCP optimizations: Fast open, window scaling, timestamps
Integrated into: mongo.yml, logs.yml, redis.yml, sqlite.yml
Benefits:
- 📈 Better throughput under load
- 📉 Lower latency
- 🔄 Improved connection handling
Security Monitoring System Guide
Comprehensive automated monitoring with email notifications for:
- 📊 System Resource Monitoring - CPU/Memory/Disk at 75%, 80%, 90%, 95%, 100% thresholds
- 🔐 SSH Security Monitoring - ALL SSH activity (successful/failed logins, logged in users, commands)
- 🔌 USB Device Monitoring - Unknown device detection with whitelisting
- 👤 Root Access Monitoring - Sudo, su, and direct root login tracking
- 🔍 Lynis System Audit - Daily security audits with hardening index
- 📦 Package Installation Monitoring - Track package installations, upgrades, removals
- 🔓 Open Ports Monitoring - Detect unexpected listening services
- 📜 SSL Certificate Monitoring - Expiration alerts (30/14/7 days before expiry)
Features:
- ⏱️ Systemd timers - Reliable periodic execution
- 📧 Email alerts - Detailed notifications through the shared, rate-limited direct-MX sender
- 🔒 Whitelisting - Authorized IPs, users, devices, sudo users
- 🚦 Rate limiting - Intelligent alert throttling to prevent flooding
- 📝 Comprehensive logs - All events logged to
/var/log/*-monitor.log
Note
All monitoring submits through /usr/local/bin/send-rate-limited-email.sh to the root-only, send-only local Postfix queue.
Testing Guide: See MONITORING_TESTING.md for comprehensive testing procedures.
The MongoDB, Logs, and Valkey/Redis hosts can publish minimal public incidents for sustained process failures while retaining diagnostics only in private alert channels. See SYSTEMD_STATUS_INCIDENTS.md for the privacy contract, token permissions, debounce behavior, rollout, and environment-driven end-to-end test.
Automated PM2 process health monitoring for Node.js applications:
- ✅ Process inventory - Reports status and uptime for every managed process
- 🔴 Errored and stopped detection - Fails the health invocation when a process is unhealthy
- ⏰ Saved-state drift detection - Detects missing or unexpected managed processes
- 📧 Email and SMS notifications - PM2 startup and health failures use the centralized direct-MX email path plus optional eligible SMS
- ⏱️ Scheduled checks - Runs every 10 minutes through a systemd timer
Integrated into: node.yml
Environment Variables:
ALERT_EMAIL_RECIPIENTS- Comma-separated email recipients for alerts
Tip
Configure ALERT_EMAIL_RECIPIENTS before running security.yml. MSMTP_RCPTS is accepted only as a deprecated migration fallback.
Complete MongoDB operations manual:
- 🔧 Version management (LOCKED to v6.0.27)
- 💾 Backup and restore procedures
- 🔄 Replica set management
- 📊 Performance monitoring
- 🐛 Troubleshooting
SQLite Storage Mirroring Guide
Periodic rsync-based mirroring between primary and secondary encrypted storage volumes:
- 🔄 Cron rsync - Mirrors every 2 minutes with near-zero memory usage
- 🛡️ Safety checks - Prevents accidental data loss
- 📧 Email notifications - Alerts for sync errors and failures
- ⏱️ Health monitoring - Systemd timer checks every 5 minutes
- 🔒 LUKS support - Works with encrypted volumes
Usage:
# Deploy sqlite-mirror to SQLite servers
ansible-playbook ansible/playbooks/sqlite-mirror.yml -l sqlite
# With custom source/target
MIRROR_SOURCE=/mnt/primary MIRROR_TARGET=/mnt/backup \
ansible-playbook ansible/playbooks/sqlite-mirror.yml -l sqliteEnvironment Variables:
| Variable | Description | Default |
|---|---|---|
MIRROR_SOURCE |
Source directory to mirror | /mnt/storage_do_1 |
MIRROR_TARGET |
Target directory for mirror | /mnt/storage_do_2 |
MIRROR_INTERVAL |
Cron interval in minutes | 2 |
ALERT_EMAIL_RECIPIENTS |
Email recipients for alerts | security@forwardemail.net |
MIRROR_SKIP_SAFETY |
Skip safety checks | false |
Warning
The playbook will fail if the target directory contains existing data to prevent accidental deletion. Set MIRROR_SKIP_SAFETY=true to bypass this check after verifying the target data is expendable.
Tip
Run this playbook after setting up LUKS-encrypted storage volumes (see README.md "Bare Metal Advice" section).
Comprehensive disaster recovery procedures:
- 💾 Backup strategies
- 🔄 Restore procedures
- 🚨 Emergency response
- 📋 Recovery checklists
MongoDB Performance Tuning Guide
Optimize MongoDB for production workloads:
- 📊 Index optimization
- 💾 WiredTiger configuration
- 🔄 Connection pooling
- 📈 Query optimization
Redis Performance Tuning Guide
Optimize Redis/Valkey for high-performance caching:
- 💾 Memory management
- 🔄 Persistence configuration
- 📊 Monitoring and metrics
- ⚡ Performance best practices
Security audit procedures for service accounts:
- 👤 User permission audits
- 🔐 SSH key management
- 📝 Access logging
- 🛡️ Security hardening
An address the sshd jail bans on one server is banned on every server (fail2ban-cluster.yml, run by security.yml):
- Bans go through a Redis stream (the application's Redis, TLS), so a server that was down catches up and a new server gets every earlier ban.
- Each server applies them in its
sshd-clusterjail: SSH only, permanent. IPv6 addresses are banned as their /64, and an IPv4 /24 is banned once 3 of its addresses have been. - Nothing overlapping loopback, private ranges, an inventory host or the operators' addresses is banned, in any jail. Adding an address later lifts its bans.
It needs REDIS_HOST, REDIS_PORT, REDIS_PASSWORD and the operators' addresses, so a mistyped login cannot lock you out of every server. Without them the play is skipped and an existing installation is left as it is. Keep the addresses in the inventory (group_vars/all):
fail2ban_ignore_ips:
- 203.0.113.10
- 2001:db8:1234::/48or set FAIL2BAN_IGNORE_IPS="203.0.113.10, 2001:db8:1234::/48" (none if there are none).
On a server:
fail2ban-client status sshd-cluster # bans applied here
fail2ban-cluster target 198.51.100.7 # what an address is banned as (or "ignored")
fail2ban-cluster unban 198.51.100.0/24 # lift overlapping bans on every server (/16 or narrower)
journalctl -u fail2ban-clusterTo apply the whole stream again on a server: systemctl stop fail2ban-cluster, rm -rf /var/lib/fail2ban-cluster, systemctl start fail2ban-cluster.
Test: sudo ansible/scripts/test-fail2ban-cluster.sh
# Deploy security baseline (includes system optimization and AMD Ryzen NUMA)
ansible-playbook ansible/playbooks/security.yml -i hosts.yml
# Deploy all database optimizations
ansible-playbook ansible/playbooks/mongo.yml -i hosts.yml
# 3b. Deploy Logs MongoDB (separate instance)
ansible-playbook ansible/playbooks/logs.yml -i hosts.yml
ansible-playbook ansible/playbooks/redis.yml -i hosts.yml
# Deploy Node.js with PM2 monitoring
ansible-playbook ansible/playbooks/node.yml -i hosts.yml# Check monitoring service status
sudo systemctl status system-resource-monitor.timer
sudo systemctl status ssh-security-monitor.timer
# View monitoring logs
sudo tail -f /var/log/system-resource-monitor.log
sudo tail -f /var/log/ssh-security-monitor.log
# Manually trigger monitoring check
sudo systemctl start system-resource-monitor.service# Validate Postfix configuration and the no-listener control
sudo postfix check
sudo postconf -h master_service_disable # Must print: inet
sudo ss -H -ltnp | grep -E 'master|smtpd|postscreen|smtp-sink' # Must print nothing
# Submit a test as root through the shared rate-limited path
sudo /usr/local/bin/send-rate-limited-email.sh \
manual-test \
"Direct-MX alert transport test" \
"Test from $(hostname)"
# Inspect queue and delivery logs
sudo postqueue -p
sudo journalctl -u postfix -f# Check MongoDB status
sudo systemctl status mongod
# Check Redis/Valkey status
sudo systemctl status valkey
# View database logs
sudo journalctl -u mongod -f
sudo journalctl -u valkey -f# Verify tmpfs /tmp is mounted
mount | grep /tmp
# Check mount options
mount | grep "on / "
# Verify AMD Ryzen NUMA settings (if applicable)
cat /proc/sys/vm/zone_reclaim_mode # Should be 0
cat /proc/sys/kernel/numa_balancing # Should be 0
cat /sys/kernel/mm/transparent_hugepage/enabled # Should be [never]The security baseline now requires its hardened, send-only Postfix transport on every managed host. Do not run remove-postfix.yml against a host managed by security.yml; doing so disables independent infrastructure alerts. The legacy cleanup playbook is retained only for decommissioned or unmanaged systems after their monitoring has been removed.