A hands-on, investigate-in-your-SIEM challenge for aspiring and practicing SOC analysts. You get the raw telemetry from a real-world-style intrusion at a fictional company, Meridian Group, and a set of questions to answer. Load the data into your own Splunk or Elasticsearch/Kibana, dig through the evidence, and work the questions in order — each stage builds on what you find in the last.
This pack contains the data and the questions — not the answers. Answers are graded by the community Discord bot (
/play,/submit,/score,/leaderboard,/hint). Nothing here spoils the investigation, so it's safe to read start to finish.
Meridian Group, a mid-size company, was compromised. Something arrived by email, someone opened it, and it went downhill from there — an endpoint foothold, hands-on-keyboard activity, movement onto the servers, and eventually a serious impact event. Your job is to reconstruct what happened from the logs.
The environment is realistic: alongside the real intrusion there's legitimate, benign activity that looks similar — an IT support tool, a routine vulnerability scan, a sanctioned backup job, ordinary business travel. Every question asks about the attacker's activity, so read carefully and rule out the noise.
scanned-document-468/
├── README.md ← you are here
├── questions.md ← the 26 questions (work them in order)
├── data-dictionary.md ← the 7 sourcetypes and what each carries — read this first
├── data/
│ ├── events.ndjson ← telemetry (one JSON event per line) — for the HEC loader
│ ├── events-single-upload.json ← ALL events in one file → routes to 7 sourcetypes on upload (no HEC)
│ └── by-sourcetype/ ← the same events split per sourcetype (alternate no-HEC upload)
│ ├── erebus_sysmon.json
│ ├── erebus_windows_security.json
│ └── … (one file per sourcetype)
├── splunk/
│ ├── indexes.conf ← the target index definition
│ ├── props.conf ← field-extraction + line-break/timestamp config (required)
│ ├── transforms.conf ← sourcetype router (only for the single-file upload)
│ └── load_splunk.py ← loads the data into Splunk via HEC
└── elastic/
└── load_elastic.py ← loads the data into Elasticsearch via _bulk
Every method below lands the data in the single iceid index across 7 sourcetypes — exactly the
same end state, whichever you pick. One-time setup first:
- Create the index. Copy
splunk/indexes.confinto$SPLUNK_HOME/etc/system/local/(or runsplunk add index iceid). - Add the parsing config. Copy the stanzas from
splunk/props.confinto your$SPLUNK_HOME/etc/system/local/props.conf. For the single-file upload (Option B) also copysplunk/transforms.confinto$SPLUNK_HOME/etc/system/local/transforms.conf. Restart Splunk. (Required — without these the fields won't extract, and for file upload the events won't line-break, get the right timestamp, or route to the correct sourcetype.) - Search — the loaders automatically shift timestamps so the last event lands near now.
Use a recent time range (e.g. Last 4 hours).
index=iceidshould show 479 events across 7 sourcetypes; confirm the spread withindex=iceid | stats count by sourcetype. (Pass--no-redateto the loader to keep the original 2025-07-01 timestamps if needed.)
Fastest if you can open the HEC port (8088).
- Enable an HEC token (Settings → Data inputs → HTTP Event Collector); make sure
iceidis in the token's allowed-index list. - Load:
python splunk/load_splunk.py --url http://localhost:8088 --token <YOUR_HEC_TOKEN>
No open ports — you just hand Splunk one file and it fans out to all 7 sourcetypes inside iceid,
the same as the HEC push. Each event in data/events-single-upload.json carries an st field naming
its sourcetype; the transforms.conf router rewrites the Splunk sourcetype at index time.
GUI: Settings → Add Data → Upload → choose data/events-single-upload.json → on Set Source
Type pick / type erebus:upload → Next → set Index = iceid → Review → Submit.
CLI:
splunk add oneshot data/events-single-upload.json -index iceid -sourcetype erebus:upload
That's it — one drop, one index, 7 sourcetypes.
Prefer not to install transforms.conf? Upload the pre-split files instead — Splunk assigns one
sourcetype per file, so the data also ships split under data/by-sourcetype/ (one file per
sourcetype). Same single iceid index, just one upload per sourcetype.
# sourcetype = "erebus:" + the file's base name
splunk add oneshot data/by-sourcetype/erebus_sysmon.json -index iceid -sourcetype erebus:sysmon
splunk add oneshot data/by-sourcetype/erebus_windows_security.json -index iceid -sourcetype erebus:windows_security
splunk add oneshot data/by-sourcetype/erebus_proxy_access.json -index iceid -sourcetype erebus:proxy_access
splunk add oneshot data/by-sourcetype/erebus_zeek_dns.json -index iceid -sourcetype erebus:zeek_dns
splunk add oneshot data/by-sourcetype/erebus_zeek_conn.json -index iceid -sourcetype erebus:zeek_conn
splunk add oneshot data/by-sourcetype/erebus_panos.json -index iceid -sourcetype erebus:panos
splunk add oneshot data/by-sourcetype/erebus_proofpoint_tap.json -index iceid -sourcetype erebus:proofpoint_tap
(GUI equivalent: the Add Data → Upload wizard, once per file, setting the matching source type.)
Why NDJSON, not CSV? These logs span seven sources whose fields don't line up in one table (and a few carry nested objects), so a single flat CSV can't hold them without losing data. The JSON files above are the drop-in "just hand Splunk a file" equivalent — no HEC, no ports.
python elastic/load_elastic.py --url http://localhost:9200 --index iceid
Elasticsearch 8+ defaults to HTTPS with a self-signed cert. If you get an SSL error, add
--no-verify (and credentials):
python elastic/load_elastic.py --url https://localhost:9200 --index iceid --user elastic --password <your-password> --no-verify
First export a Kibana-ready NDJSON file, then upload it through the UI:
python elastic/load_elastic.py --export events-elastic.ndjson
In Kibana: Machine Learning → Data Visualizer → Import data (or Upload a file). Select the
events-elastic.ndjson file you just exported, set the index name to iceid, and import.
Don't upload
data/events.ndjsondirectly — that file uses a Splunk-style envelope format that Kibana's upload UI can't parse. The--exportflag flattens each event into a clean Elasticsearch document first.
Then create a Kibana Data View over the iceid index and search with a recent time range
(e.g. Last 4 hours) — timestamps are shifted to land near now by default. Each doc carries a
sourcetype field (e.g. erebus:sysmon) so you can filter exactly as the questions describe,
plus an @timestamp.
-
Read
data-dictionary.mdto learn which sourcetype holds which kind of evidence. -
Work through
questions.mdin order (Stage 0 → Stage 4). Most questions have exactly one answer; a couple of "name one" questions accept any valid member of a small set. -
Answer in the community Discord. Run
/playto start: it serves one question at a time and remembers where you are. Answer with/answer <your answer>,/skipto move on, and/hint <question#>if you're stuck (a nudge, then stronger, then near-answer, at a small point cost). Prefer to jump around?/submit <question#> <answer>answers any question directly. Check/scoreand/leaderboardany time. The capstone (Q26) is mentor-reviewed.Every command except
/leaderboardreplies privately to you, so you can play in a shared channel without spoiling anything for anyone else.
Good hunting. 🔎
Synthetic training data. Any resemblance to real companies, users, or infrastructure is coincidental. Built for defensive security education.