Unified Big Data Engineering Ecosystem for Docker, Google Cloud (Dataproc & GCE), VMware (Kali Linux), Hyper-V, WSL 2, Oracle VirtualBox & Bare-Metal Linux
Overview β’ Environments β’ Quick Start β’ Architecture β’ Web Consoles β’ Tutorials β’ Datasets β’ Structure β’ Docs
Welcome to the Apache Hadoop Enterprise Multi-Platform Lab β a unified, production-grade Big Data engineering workspace designed for distributed computing research, university courses, ETL prototyping, and performance benchmarking.
Unlike conventional single-purpose repositories, this project delivers a cross-platform deployment suite supporting:
- Containerized Cluster (Docker & Docker Compose): Single-node Hadoop 3.1.2 with automated daemon supervisors, named persistent storage, and built-in health probes.
- Google Cloud Platform (Dataproc & Compute Engine): Managed Apache Hadoop 3 + Spark clusters with decoupled Cloud Storage (
gs://), Component Gateway web consoles, Spot workers, and 1-click Compute Engine VM deployment. - Virtual Machine Workstations (VMware Workstation Pro & Kali Linux): Automated VMX hardware tuning (6GB RAM, 4 vCPUs, G1GC optimization, swappiness tuning, full-screen 1080p, programmatic mouse fix, and Hadoop 3.3.6 installer).
- Enterprise Type-1 Hypervisor (Microsoft Hyper-V Generation 2): 4 vCPUs, dynamic memory allocation, enhanced session mode (
HvSocketbidirectional clipboard), and automated NAT virtual switch recovery. - Near-Bare-Metal Windows Subsystem (WSL 2 Ubuntu): Ultra-fast I/O with XFCE4 visual desktop over RDP (port 3390) and zero-friction cluster startup.
- Open Source Virtualization (Oracle VirtualBox): Automated PowerShell VM orchestrator (
virtualbox-setup.ps1) with NAT port forwarding rules. - Native Linux & Bare Metal: Non-root systemd service unit configurations and user-space zero-sudo installers.
- Multi-Environment Orchestration: Launch Hadoop across Docker, VMware, Hyper-V, WSL 2, or VirtualBox with platform-tailored scripts.
- Complete Hadoop Daemon Stack:
- HDFS: NameNode, DataNode, SecondaryNameNode.
- YARN: ResourceManager, NodeManager.
- MapReduce: JobHistory Server.
- 1-Click Windows Launchers (
launchers/windows/):Start-Hadoop-Docker.bat&Stop-Hadoop-Docker.batfor instant Docker cluster control.Deploy-Hadoop-GCP.bat&Deploy-Hadoop-GCP.ps1for Google Cloud Dataproc & GCE orchestration.Launch-Kali-VMware.batfor automated VMX tuning & Kali boot.Fix-Lag-And-Start-VM.bat&Fix-VM-Internet.batfor Hyper-V management.Ubuntu-WSL-GUI.rdpfor instant Remote Desktop GUI access.
- Hands-On Big Data Tutorials (
examples/):- HDFS CLI: Comprehensive operations walkthrough (
demo-hdfs-operations.sh) covering block inspection, quotas, and SafeMode. - Python Hadoop Streaming: Automated mapper/reducer WordCount pipeline.
- Java Native MapReduce: Standalone WordCount application with automated compiler and runner.
- Apache Spark & PySpark: Direct HDFS Parquet & CSV DataFrame ingestion and aggregation.
- Apache Sqoop Ingestion (
examples/sqoop/): MySQL RDBMS <-> HDFS & Hive bulk import/export scripts and code generation. - Apache Oozie Workflows (
examples/oozie/): Production multi-action DAG pipeline (workflow.xml) and daily scheduler (coordinator.xml). - Apache Pig Latin (
examples/pig/): High-level data transformation, filtering, and country aggregation scripts (analytics.pig). - Apache Hive Warehouse (
examples/hive/): External table DDL (create-tables.hql) and window rank queries (analytics.hql). - Apache Flume & HBase (
examples/flume/,examples/hbase/): Spooling directory streaming agent and columnar NoSQL table scripts.
- HDFS CLI: Comprehensive operations walkthrough (
- Pre-Packaged Datasets (
datasets/):- Real-world unstructured text (
wordcount-sample.txt) and tabular records (employees.csv) for zero-setup experimentation.
- Real-world unstructured text (
- Enterprise Hardening & DevOps CI/CD:
- Dedicated non-root
hduser:hadoop(UID/GID 1000) execution. - Multi-stage Docker build pruning ~60,000 redundant Javadoc HTML files to bypass filesystem journal overhead.
- GitHub Actions CI matrix validating builds, container health checks, JPS daemons, and MapReduce jobs on every push.
- Dedicated non-root
π Click the diagram above to view the scalable, high-definition vector SVG version.
flowchart TB
subgraph Host["π» DEVELOPER WORKSTATION & BROWSER ACCESS LAYER"]
direction LR
DevHub["π <b>Unified Big Data Control Hub</b><br/><b>http://localhost:3030</b><br/>Single Pane of Glass UI"]
DevJupyter["πͺ <b>JupyterLab PySpark</b><br/><b>http://localhost:8888</b><br/>Interactive Data Pipelines"]
DevSpark["β‘ <b>Spark Master & History</b><br/><b>:8080 • :8081 • :18080</b><br/>Compute Consoles"]
DevHadoop["π <b>Hadoop HDFS & YARN</b><br/><b>:9870 • :8088 • :19888</b><br/>Storage & Scheduling"]
end
subgraph Deployments["π₯οΈ MULTI-PLATFORM CLUSTER RUNTIMES"]
direction LR
PlatDocker["π³ Docker Compose Full Stack<br/><b>Hadoop + Spark + Hive + Hub</b><br/>6 Synchronized Services"]
PlatVMware["π VMware Workstation<br/><b>Kali Linux 2026.2</b><br/>Hadoop 3.3.6 LTS<br/>6GB RAM / 4 vCPUs"]
PlatGCP["βοΈ Google Cloud<br/><b>Dataproc & GCE</b><br/>Decoupled gs://<br/>Auto-Idle Teardown"]
PlatHyperV["πͺ Microsoft Hyper-V<br/><b>Ubuntu 24.04 Gen 2</b><br/>4 vCPUs / Dynamic RAM<br/>HvSocket Clipboard"]
PlatWSL["π§ WSL 2 Ubuntu<br/><b>Windows 11 Native</b><br/>XFCE GUI Desktop<br/>Port 3390 (RDP)"]
end
subgraph CoreEngine["π APACHE HADOOP & SPARK DISTRIBUTED ECOSYSTEM"]
direction TB
subgraph HDFS["ποΈ HDFS DISTRIBUTED STORAGE LAYER"]
direction TB
NN["π NameNode (Master)<br/><b>Port: 9870 (Web / WebHDFS) | 9000 (RPC)</b><br/>Inodes Namespace Graph & WAL Journal"]
SNN["π SecondaryNameNode<br/><b>Port: 9868 (HTTP)</b><br/>Consolidates fsimage.ckpt Checkpoints"]
DN["π¦ DataNode (Worker)<br/><b>Port: 9864 (Web) | 9866 (Data)</b><br/>128MB Blocks • CRC32C Checksums"]
NN <-->|"Heartbeats (3s) & Block Reports"| DN
NN <-->|"Checkpoint Sync"| SNN
end
subgraph ComputeGrid["βοΈ MULTI-ENGINE COMPUTE & SCHEDULING"]
direction TB
SparkM["β‘ Spark Master • Port: 8080 (Web) | 7077 (RPC)<br/>In-Memory DAG Scheduling & Stages"]
SparkW["π¨ Spark Workers (:8081) • In-Memory Task Executors"]
SparkH["β±οΈ Spark History Server (:18080) • Event Logs"]
HiveMS["π Apache Hive Warehouse • Metastore (:9083) | JDBC (:10000)"]
RM["π§ YARN ResourceManager • Port: 8088 (Web) | 8032 (IPC)"]
NM["π· YARN NodeManager • Port: 8042 (Web) | cgroups Slots"]
JHS["π MapReduce JobHistory • Port: 19888 (Web)"]
SparkM <--> SparkW
SparkM -.-> SparkH
RM <--> NM
NM --> JHS
end
ComputeGrid -.->|"Data Locality Read/Write"| DN
end
subgraph Storage["πΎ DURABLE PERSISTENT STORAGE TIER (ZERO DATA LOSS)"]
direction LR
V_Docker["π Docker Named Volumes<br/>hadoop_namenode_data<br/>hadoop_datanode_data"]
V_Spark["β‘ spark_event_logs_data<br/>/spark-logs • /opt/spark/events"]
V_Jupyter["πͺ jupyter_notebooks_data<br/>/home/jovyan/work"]
V_GCS["βοΈ Google Cloud Storage<br/>gs://bucket/data & staging"]
end
Host ==>|"β Submit Pipelines & Queries"| Deployments
Deployments ==>|"Dispatch to Compute Grid"| CoreEngine
NN ==>|"Persist Inodes"| Storage
DN ==>|"Store 128MB Blocks"| Storage
ComputeGrid ==>|"Stream Logs & Events"| Storage
classDef hostStyle fill:#0f172a,stroke:#38bdf8,stroke-width:2px,color:#f8fafc;
classDef platStyle fill:#1e1b4b,stroke:#818cf8,stroke-width:2px,color:#f8fafc;
classDef hdfsStyle fill:#0c2d48,stroke:#00a8e8,stroke-width:2px,color:#f8fafc;
classDef yarnStyle fill:#064e3b,stroke:#10b981,stroke-width:2px,color:#f8fafc;
classDef sparkStyle fill:#4c0519,stroke:#f43f5e,stroke-width:2px,color:#f8fafc;
classDef storageStyle fill:#450a0a,stroke:#f87171,stroke-width:2px,color:#f8fafc;
class DevHub,DevJupyter,DevSpark,DevHadoop hostStyle;
class PlatDocker,PlatVMware,PlatGCP,PlatHyperV,PlatWSL platStyle;
class NN,SNN,DN hdfsStyle;
class RM,NM,JHS,HiveMS yarnStyle;
class SparkM,SparkW,SparkH sparkStyle;
class V_Docker,V_Spark,V_Jupyter,V_GCS storageStyle;
All Big Data web consoles, interactive developer studios, and service endpoints are mapped to localhost:
| Service | Container Port | Host Port | Web Console URL | Description |
|---|---|---|---|---|
| π Unified Control Hub | 3000 |
3030 |
http://localhost:3030 | Single pane of glass dashboard: live health monitor, HDFS browser, Hive SQL studio, Sqoop builder, Oozie orchestrator, Pig sandbox & diagrams. |
| β‘ Apache Spark Master | 8080 |
8080 |
http://localhost:8080 | Standalone cluster coordinator, CPU cores, active workers, and running applications. |
| π¨ Apache Spark Worker | 8081 |
8081 |
http://localhost:8081 | Worker node execution slots, thread pools, and executor memory metrics. |
| β±οΈ Spark History Server | 18080 |
18080 |
http://localhost:18080 | Post-mortem Spark job diagnostics, DAG execution stages, and timeline metrics. |
| πͺ JupyterLab PySpark Studio | 8888 |
8888 |
http://localhost:8888 | Interactive Data Engineering notebooks pre-loaded with PySpark, Pandas, and Delta Lake. |
| π HDFS NameNode | 9870 |
9870 |
http://localhost:9870 | Browse HDFS filesystem, inspect cluster capacity, and WebHDFS REST API. |
| βοΈ YARN ResourceManager | 8088 |
8088 |
http://localhost:8088 | Monitor running YARN applications, cluster memory/vcore metrics, and queues. |
| π¦ HDFS DataNode | 9864 |
9864 |
http://localhost:9864 | Inspect DataNode volume status, block pools, and raw chunk metrics. |
| π· YARN NodeManager | 8042 |
8042 |
http://localhost:8042 | Container allocation and per-node execution details. |
| π MapReduce JobHistory | 19888 |
19888 |
http://localhost:19888 | Historical MapReduce task counters, logs, and execution timelines. |
| π Apache Hive Warehouse | 10002 |
10002 |
http://localhost:10002 | Schema Metastore (:9083) and HiveServer2 JDBC interface. |
| π Apache Oozie Engine | 11000 |
11000 |
http://localhost:11000/oozie |
Workflow DAG scheduler and coordinator pipeline engine. |
| π Apache Sqoop Ingestion | -- | 3030 |
Sqoop Studio | Bulk RDBMS <-> HDFS/Hive data transfer generator & simulator. |
| π· Apache Pig Latin | -- | 3030 |
Pig Studio | High-level dataflow Pig Latin compilation and execution sandbox. |
| π Spark Master RPC | 7077 |
7077 |
spark://localhost:7077 |
Cluster manager endpoint for PySpark & spark-submit. |
| π HDFS RPC Endpoint | 9000 |
9000 |
hdfs://localhost:9000 |
IPC protocol endpoint for external tools (Spark, Flink, PySpark). |
| π SSH Bastion | 22 |
22222 |
ssh -p 22222 hduser@localhost |
Direct SSH shell access (password: ubuntu). |
| π₯οΈ WSL 2 GUI Desktop | 3390 |
3390 |
localhost:3390 (RDP) |
XFCE4 graphical desktop session for WSL 2 (Ubuntu-WSL-GUI.rdp). |
Tip
Recommended Workflow: Open the Unified Big Data Control Hub in your browser. It automatically monitors and links to every service listed above with one-click access!
Note
Web consoles operate over plain HTTP (http://). If your browser auto-redirects to HTTPS, open an Incognito / Private Window using http://127.0.0.1:3030 or refer to Troubleshooting Runbook: Issue 9.
Choose your preferred deployment platform below:
Option 1: Docker Compose (1-Click or CLI) - Recommended
Double-click launchers/windows/Start-Hadoop-Docker.bat β it boots all Hadoop & Spark containers and automatically launches the Unified Big Data Control Hub at http://localhost:3030 in your default browser.
# Clone the repository
git clone https://github.com/Sohila-Khaled-Abbas/docker-hadoop.git
cd docker-hadoop
# Start all Big Data containers (Hadoop, Spark, JupyterLab, Control Hub)
docker compose up -d
# Open the Unified Control Hub in your browser
# http://localhost:3030
# Inspect cluster health and active daemons
docker compose ps
docker compose exec hadoop jpsTo stop the cluster:
docker compose down
# or double-click launchers/windows/Stop-Hadoop-Docker.batOption 2: Kali Linux on VMware Workstation Pro
- From Windows host, double-click
launchers/windows/Launch-Kali-VMware.bat(tunes VMX for 6GB RAM, 4 vCPUs, disables WHPX popups, and launches VMware). - Inside Kali Linux terminal (
user: kali,pass: kali):bash scripts/vmware/install-hadoop-kali.sh
- Read the complete Kali Linux & VMware Guide.
Option 3: Microsoft Hyper-V Generation 2 (Ubuntu)
- Double-click
launchers/windows/Fix-Lag-And-Start-VM.batto allocate 4 vCPUs and launch the VM. - If internet connection is lost, double-click
launchers/windows/Fix-VM-Internet.bat. - To enable clipboard & full-screen resizing, run
launchers/windows/Eject-ISO-And-Enable-Clipboard.bat. - Read the complete Hyper-V Ubuntu Guide.
Option 4: WSL 2 Ubuntu with Visual XFCE4 GUI
- Inside WSL 2 Ubuntu terminal:
bash scripts/wsl/install-hadoop-wsl.sh
- Double-click
launchers/windows/Ubuntu-WSL-GUI.rdpto connect to the desktop interface onlocalhost:3390. - Start the Hadoop cluster with
bash scripts/wsl/start-hadoop-cluster.sh. - Read the complete WSL 2 Ubuntu GUI Guide.
Option 5: Oracle VirtualBox Automated Setup
- Run the automated PowerShell VM creator:
powershell -ExecutionPolicy Bypass -File .\scripts\virtualbox\virtualbox-setup.ps1
- SSH into the VM:
ssh -p 2222 hadoopuser@localhost(password:hadoopuser). - Read the complete VirtualBox Ubuntu Guide.
Option 6: Google Cloud Platform (Dataproc & Compute Engine)
- Via 1-Click Windows Launcher: Double-click
launchers/windows/Deploy-Hadoop-GCP.batto interactively create Dataproc clusters, submit jobs togs://, deploy to Compute Engine, or teardown clusters. - Via Terminal CLI:
- Create an auto-terminating Dataproc cluster with Component Gateway:
bash scripts/gcp/create-dataproc-cluster.sh
- Submit a Python Hadoop Streaming WordCount job reading/writing from Cloud Storage:
bash scripts/gcp/submit-mapreduce-job.sh streaming
- Or deploy the Docker containerized Hadoop stack to a Google Compute Engine VM:
bash scripts/gcp/deploy-hadoop-gce.sh
- Teardown Dataproc cluster to prevent cloud charges:
bash scripts/gcp/teardown-dataproc-cluster.sh
- Create an auto-terminating Dataproc cluster with Component Gateway:
- Read the complete Google Cloud Dataproc & GCE Guide.
Run the comprehensive HDFS operations walkthrough inside the container:
docker compose exec hadoop bash < examples/hdfs-cli/demo-hdfs-operations.shRead the HDFS CLI Tutorial for complete command examples.
Execute mapper/reducer streaming pipeline on sample text:
make test-mr-python
# or
bash examples/mapreduce-python/run.shRead the Python Streaming Guide.
Compile and execute standalone Java MapReduce job:
make test-mr-java
# or
bash examples/mapreduce-java/compile-and-run.shRead the Java MapReduce Guide.
Run standalone Python script connecting to HDFS:
python examples/spark-pyspark/pyspark_hdfs_read_write.pyRead the PySpark Integration Guide.
Open http://localhost:8888 or browse the notebooks/ directory:
01-pyspark-hdfs-pipeline.ipynb: End-to-end ingestion, schema transformation, and partitioned Snappy Parquet write to HDFS.02-spark-sql-hive-analytics.ipynb: Window functions, revenue ranking, and Spark SQL queries over HDFS tables.03-realtime-streaming-simulation.ipynb: Structured Streaming micro-batch windowed aggregations.
Open http://localhost:3030 to view live cluster health, browse HDFS directories via WebHDFS, submit Spark and MapReduce jobs with live terminal output, and inspect architecture diagrams.
The repository includes pre-built test datasets in datasets/:
| Dataset | Format | Path | Purpose |
|---|---|---|---|
| Text Corpus | .txt |
datasets/wordcount-sample.txt |
WordCount benchmarking, tokenization, grep |
| Employees Data | .csv |
datasets/employees.csv |
PySpark DataFrames, aggregations, SQL queries |
To load them directly into HDFS:
docker cp datasets/employees.csv hadoop-master:/tmp/
docker compose exec hadoop hdfs dfs -mkdir -p /datasets
docker compose exec hadoop hdfs dfs -put -f /tmp/employees.csv /datasets/
docker compose exec hadoop hdfs dfs -cat /datasets/employees.csvSee Datasets Documentation for advanced ingestion recipes.
docker-hadoop/
βββ .github/ # GitHub Actions CI/CD & Issue Templates
βββ config/ # XML & Configuration Files
β βββ core-site.xml # Filesystem & temporary storage configuration
β βββ hadoop-env.sh # Environment exports & JVM options
β βββ hdfs-site.xml # NameNode, DataNode & WebHDFS settings
β βββ mapred-site.xml # MapReduce framework & JobHistory configuration
β βββ spark-defaults.conf # Spark HDFS & History Server defaults
β βββ yarn-site.xml # YARN ResourceManager & NodeManager settings
βββ datasets/ # Built-in sample datasets
β βββ employees.csv # Structured employee records for Spark/SQL
β βββ wordcount-sample.txt # Distributed systems text corpus
β βββ README.md # HDFS loading instructions & recipes
βββ docs/ # Comprehensive technical documentation
β βββ architecture.md # Internal architecture, HDFS & Spark deep dive
β βββ configuration-tuning.md # XML tuning & JVM GC optimization
β βββ data-engineering-patterns.md # Lakehouse, Medallion, & join patterns
β βββ ecosystem-integration.md # Spark, Hive, Presto, & Jupyter guides
β βββ getting-started.md # Fast onboarding guide
β βββ google-cloud-dataproc-hadoop-guide.md # GCP Dataproc & GCE deployment
β βββ hadoop-ecosystem-guide.md# Complete ecosystem, HDFS & fault-tolerance guide
β βββ hyperv-ubuntu-guide.md # Microsoft Hyper-V setup & optimization
β βββ images/ # High-definition vector SVGs & 4K PNG diagrams
β βββ kali-vmware-hadoop-guide.md # VMware Workstation & Kali Linux guide
β βββ mapreduce-guide.md # Comprehensive MapReduce manual
β βββ software-engineering-practices.md # 12-factor Big Data & DevOps
β βββ troubleshooting.md # Diagnostic runbook for cluster issues
β βββ virtualbox-ubuntu-guide.md # Oracle VirtualBox guide
β βββ wsl2-ubuntu-hadoop-guide.md# WSL 2 Ubuntu GUI & Hadoop setup
βββ examples/ # Hands-on Big Data examples
β β βββ README.md
β βββ mapreduce-python/ # Python Hadoop Streaming example
β β βββ mapper.py
β β βββ reducer.py
β β βββ run.sh
β β βββ sample.txt
β β βββ README.md
β βββ spark-pyspark/ # PySpark HDFS read/write integration
β βββ pyspark_hdfs_read_write.py
β βββ README.md
βββ launchers/ # Standalone 1-Click platform launchers
β βββ README.md # Launcher catalog and usage guide
β βββ windows/ # Windows 1-click desktop batch launchers
β βββ Start-Hadoop-Docker.bat # 1-Click Docker cluster startup
β βββ Stop-Hadoop-Docker.bat # 1-Click Docker cluster shutdown
β βββ Deploy-Hadoop-GCP.bat # Google Cloud Dataproc & GCE launcher
β βββ Launch-Kali-VMware.bat # VMware Workstation Kali launcher
β βββ Fix-Lag-And-Start-VM.bat# Hyper-V 4-vCPU & performance launcher
β βββ Fix-VM-Internet.bat # Hyper-V virtual switch network repair
β βββ Eject-ISO-And-Enable-Clipboard.bat # Hyper-V ISO & clipboard setup
β βββ Ubuntu-WSL-GUI.rdp # WSL 2 Remote Desktop profile
βββ scripts/ # Modular automation scripts
β βββ README.md # Script catalog and runtime architecture
β βββ docker/ # Docker container entrypoint & probes
β β βββ entrypoint.sh
β β βββ healthcheck.sh
β β βββ test-cluster.sh
β βββ gcp/ # Google Cloud Dataproc & GCE automation
β β βββ create-dataproc-cluster.sh
β β βββ submit-mapreduce-job.sh
β β βββ teardown-dataproc-cluster.sh
β β βββ deploy-hadoop-gce.sh
β βββ vmware/ # VMware & Kali Linux automation
β β βββ install-hadoop-kali.sh
β β βββ optimize-kali-vmx.ps1
β β βββ apply-guest-mouse-fix.ps1
β β βββ fix-mouse-in-guest.sh
β β βββ set-fullscreen-resolution.ps1
β βββ hyperv/ # Hyper-V host & guest automation
β β βββ configure-hyperv-host-enhanced-session.ps1
β β βββ enable-hyperv-enhanced-session.sh
β β βββ fix-vm-internet.ps1
β β βββ optimize-hyperv-vm.ps1
β βββ virtualbox/ # VirtualBox VM provisioning
β β βββ virtualbox-setup.ps1
β βββ wsl/ # WSL 2 Ubuntu automation
β β βββ install-hadoop-wsl.sh
β β βββ start-hadoop-cluster.sh
β β βββ sync-wsl-configs.sh
β β βββ fix-xrdp.sh
β βββ linux/ # Bare-metal & native Linux scripts
β βββ install-hadoop-ubuntu.sh
β βββ install-hadoop-user.sh
β βββ setup-hadoop-systemd.sh
β βββ start-daemons-direct.sh
βββ .dockerignore # Docker build exclusions
βββ .env.example # Port and environment variable templates
βββ .gitignore # Git exclusions (ISOs, 7z, and VM disks ignored)
βββ CHANGELOG.md # Semantic version history
βββ CODE_OF_CONDUCT.md # Community code of conduct
βββ CONTRIBUTING.md # Guidelines for contributing
βββ Dockerfile # Multi-stage Ubuntu 20.04 Hadoop 3.1.2 image
βββ docker-compose.yml # Multi-volume container orchestration
βββ LICENSE # Apache 2.0 License
βββ Makefile # Developer CLI shortcuts
βββ README.md # Project documentation
βββ SECURITY.md # Security policies & vulnerability reporting
| Command | Description |
|---|---|
make help |
Display available developer CLI commands |
make build |
Build the Hadoop Docker image locally |
make up |
Start the Hadoop Docker cluster in background |
make down |
Stop and remove the Docker container |
make restart |
Restart the cluster services |
make logs |
Stream container logs in real time |
make ps |
Inspect container health status |
make jps |
List running Java daemons inside the container |
make test |
Run built-in integration tests & Pi MapReduce |
make test-mr-python |
Run Python Streaming MapReduce WordCount |
make test-mr-java |
Compile & run Java Native MapReduce WordCount |
make test-hdfs-cli |
Run HDFS CLI interactive demo script |
make safemode-leave |
Force HDFS NameNode to leave SafeMode |
make hdfs-report |
Display HDFS storage capacity report |
make bash |
Open root shell inside the container |
make hdfs-shell |
Open interactive shell as hduser |
make clean |
Full teardown (removes containers, images, and named volumes) |
make vm-create |
Create & configure Ubuntu VM in Oracle VirtualBox |
make vm-start |
Start VirtualBox VM in GUI window |
make vm-ssh |
Connect to VirtualBox VM via SSH (port 2222) |
make vm-kali-optimize |
Tune Kali VMX hardware specs (6GB RAM, 4 vCPUs) |
make vm-kali-start |
Launch Kali Linux in VMware Workstation |
make gcp-dataproc-create |
Provision auto-terminating Dataproc cluster |
make gcp-dataproc-stream |
Submit Python Streaming WordCount to Dataproc |
make gcp-dataproc-java |
Submit Native Java MapReduce Pi to Dataproc |
make gcp-dataproc-delete |
Teardown Dataproc cluster (stop charges) |
make gcp-gce-deploy |
Deploy Docker Hadoop to Google Compute Engine |
Explore our comprehensive technical documentation and deep-dive guides:
| Document | Topic & Focus | Key Highlights |
|---|---|---|
| Google Cloud Dataproc Guide | Google Cloud (Dataproc & GCE) | Managed Hadoop 3 + Spark clusters, decoupled Cloud Storage (gs://), Component Gateway web consoles, Spot workers, and GCE deployment. |
| Hadoop Ecosystem Guide | Ecosystem & HDFS Architecture | Core Hadoop principles, component classification, HDFS block sizes, replication topology, fault tolerance, NameNode HA, and write pipelines. |
| System Architecture | Architecture & Daemon Internals | Comprehensive system breakdown, NameNode vs DataNode table, block management, heartbeat mechanisms, Secondary vs Standby NameNode, and network topology. |
| Getting Started | Fast Onboarding | Prerequisites, 3-minute quickstart, cluster verification, and basic data ingest. |
| Configuration & Tuning | Performance & GC Tuning | JVM G1GC optimizations, XML configuration recipes (hdfs-site.xml, yarn-site.xml), heap memory sizing. |
| MapReduce Engineering Manual | Compute Paradigms | Detailed MapReduce execution flow, combiners, partitioners, custom Writable comparators, and streaming pipelines. |
| Data Engineering Patterns | Enterprise Architecture | Medallion Lakehouse architecture (Bronze/Silver/Gold), idempotent pipelines, distributed joins, compaction, and data partitioning. |
| Ecosystem Integration | Modern Big Data Stack | Connecting Apache Spark, Hive, Presto/Trino, Kafka, and Jupyter notebooks to the containerized HDFS storage layer. |
| Troubleshooting Runbook | Diagnostics & RCA | 10+ categorized production issue resolutions (SafeMode, RPC connection refused, Java OutOfMemory, port collisions, browser empty responses). |
| Software Engineering Practices | Big Data DevOps | 12-factor Big Data principles, CI/CD with GitHub Actions, container health checks, and linting. |
| Kali Linux & VMware Guide | VMware Workstation Pro | Automated VMX tuning (6GB RAM, 4 vCPUs), resolution scaling, guest mouse integration, and native Hadoop 3.3.6 installation. |
| Hyper-V Ubuntu Guide | Microsoft Hyper-V Gen 2 | Dynamic memory, virtual switch recovery, enhanced session mode via HvSocket, and full-screen display. |
| WSL 2 Ubuntu GUI Guide | Windows Subsystem for Linux | Native I/O performance, XFCE4 desktop GUI over RDP (port 3390), and single-script cluster lifecycle. |
| Oracle VirtualBox Guide | VirtualBox Automation | PowerShell VM provisioning script, NAT port forwarding rules, and headless execution. |
Contributions are warmly welcomed! Please read our Contributing Guidelines and Code of Conduct before submitting pull requests.
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.