Skip to content

About

A hands-on SOC analyst investigation challenge. Load synthetic intrusion telemetry into your own Splunk or Elasticsearch and work 26 questions about attacker activity at a fictional company. Graded in Discord.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SOC Analyst CTF — "Scanned Document 468"

A hands-on, investigate-in-your-SIEM challenge for aspiring and practicing SOC analysts. You get the raw telemetry from a real-world-style intrusion at a fictional company, Meridian Group, and a set of questions to answer. Load the data into your own Splunk or Elasticsearch/Kibana, dig through the evidence, and work the questions in order — each stage builds on what you find in the last.

This pack contains the data and the questions — not the answers. Answers are graded by the community Discord bot (/play, /submit, /score, /leaderboard, /hint). Nothing here spoils the investigation, so it's safe to read start to finish.

The scenario

Meridian Group, a mid-size company, was compromised. Something arrived by email, someone opened it, and it went downhill from there — an endpoint foothold, hands-on-keyboard activity, movement onto the servers, and eventually a serious impact event. Your job is to reconstruct what happened from the logs.

The environment is realistic: alongside the real intrusion there's legitimate, benign activity that looks similar — an IT support tool, a routine vulnerability scan, a sanctioned backup job, ordinary business travel. Every question asks about the attacker's activity, so read carefully and rule out the noise.

What's in this pack

scanned-document-468/
├── README.md              ← you are here
├── questions.md           ← the 26 questions (work them in order)
├── data-dictionary.md     ← the 7 sourcetypes and what each carries — read this first
├── data/
│   ├── events.ndjson              ← telemetry (one JSON event per line) — for the HEC loader
│   ├── events-single-upload.json  ← ALL events in one file → routes to 7 sourcetypes on upload (no HEC)
│   └── by-sourcetype/             ← the same events split per sourcetype (alternate no-HEC upload)
│       ├── erebus_sysmon.json
│       ├── erebus_windows_security.json
│       └── … (one file per sourcetype)
├── splunk/
│   ├── indexes.conf       ← the target index definition
│   ├── props.conf         ← field-extraction + line-break/timestamp config (required)
│   ├── transforms.conf    ← sourcetype router (only for the single-file upload)
│   └── load_splunk.py     ← loads the data into Splunk via HEC
└── elastic/
    └── load_elastic.py    ← loads the data into Elasticsearch via _bulk

Load it into Splunk

Every method below lands the data in the single iceid index across 7 sourcetypes — exactly the same end state, whichever you pick. One-time setup first:

  1. Create the index. Copy splunk/indexes.conf into $SPLUNK_HOME/etc/system/local/ (or run splunk add index iceid).
  2. Add the parsing config. Copy the stanzas from splunk/props.conf into your $SPLUNK_HOME/etc/system/local/props.conf. For the single-file upload (Option B) also copy splunk/transforms.conf into $SPLUNK_HOME/etc/system/local/transforms.conf. Restart Splunk. (Required — without these the fields won't extract, and for file upload the events won't line-break, get the right timestamp, or route to the correct sourcetype.)
  3. Search — the loaders automatically shift timestamps so the last event lands near now. Use a recent time range (e.g. Last 4 hours). index=iceid should show 479 events across 7 sourcetypes; confirm the spread with index=iceid | stats count by sourcetype. (Pass --no-redate to the loader to keep the original 2025-07-01 timestamps if needed.)

Option A — HTTP Event Collector (HEC)

Fastest if you can open the HEC port (8088).

  1. Enable an HEC token (Settings → Data inputs → HTTP Event Collector); make sure iceid is in the token's allowed-index list.
  2. Load: python splunk/load_splunk.py --url http://localhost:8088 --token <YOUR_HEC_TOKEN>

Option B — No HEC, one file (closest to the HEC injection)

No open ports — you just hand Splunk one file and it fans out to all 7 sourcetypes inside iceid, the same as the HEC push. Each event in data/events-single-upload.json carries an st field naming its sourcetype; the transforms.conf router rewrites the Splunk sourcetype at index time.

GUI: Settings → Add Data → Upload → choose data/events-single-upload.json → on Set Source Type pick / type erebus:upload → Next → set Index = iceid → Review → Submit.

CLI:

splunk add oneshot data/events-single-upload.json -index iceid -sourcetype erebus:upload

That's it — one drop, one index, 7 sourcetypes.

Option C — No HEC, per-sourcetype files (no router needed)

Prefer not to install transforms.conf? Upload the pre-split files instead — Splunk assigns one sourcetype per file, so the data also ships split under data/by-sourcetype/ (one file per sourcetype). Same single iceid index, just one upload per sourcetype.

# sourcetype = "erebus:" + the file's base name
splunk add oneshot data/by-sourcetype/erebus_sysmon.json           -index iceid -sourcetype erebus:sysmon
splunk add oneshot data/by-sourcetype/erebus_windows_security.json -index iceid -sourcetype erebus:windows_security
splunk add oneshot data/by-sourcetype/erebus_proxy_access.json     -index iceid -sourcetype erebus:proxy_access
splunk add oneshot data/by-sourcetype/erebus_zeek_dns.json         -index iceid -sourcetype erebus:zeek_dns
splunk add oneshot data/by-sourcetype/erebus_zeek_conn.json        -index iceid -sourcetype erebus:zeek_conn
splunk add oneshot data/by-sourcetype/erebus_panos.json            -index iceid -sourcetype erebus:panos
splunk add oneshot data/by-sourcetype/erebus_proofpoint_tap.json   -index iceid -sourcetype erebus:proofpoint_tap

(GUI equivalent: the Add Data → Upload wizard, once per file, setting the matching source type.)

Why NDJSON, not CSV? These logs span seven sources whose fields don't line up in one table (and a few carry nested objects), so a single flat CSV can't hold them without losing data. The JSON files above are the drop-in "just hand Splunk a file" equivalent — no HEC, no ports.

Load it into Elasticsearch / Kibana

Option A — Python script (recommended)

python elastic/load_elastic.py --url http://localhost:9200 --index iceid

Elasticsearch 8+ defaults to HTTPS with a self-signed cert. If you get an SSL error, add --no-verify (and credentials):

python elastic/load_elastic.py --url https://localhost:9200 --index iceid --user elastic --password <your-password> --no-verify

Option B — Kibana file upload (no terminal required)

First export a Kibana-ready NDJSON file, then upload it through the UI:

python elastic/load_elastic.py --export events-elastic.ndjson

In Kibana: Machine Learning → Data Visualizer → Import data (or Upload a file). Select the events-elastic.ndjson file you just exported, set the index name to iceid, and import.

Don't upload data/events.ndjson directly — that file uses a Splunk-style envelope format that Kibana's upload UI can't parse. The --export flag flattens each event into a clean Elasticsearch document first.


Then create a Kibana Data View over the iceid index and search with a recent time range (e.g. Last 4 hours) — timestamps are shifted to land near now by default. Each doc carries a sourcetype field (e.g. erebus:sysmon) so you can filter exactly as the questions describe, plus an @timestamp.

Play

  1. Read data-dictionary.md to learn which sourcetype holds which kind of evidence.

  2. Work through questions.md in order (Stage 0 → Stage 4). Most questions have exactly one answer; a couple of "name one" questions accept any valid member of a small set.

  3. Answer in the community Discord. Run /play to start: it serves one question at a time and remembers where you are. Answer with /answer <your answer>, /skip to move on, and /hint <question#> if you're stuck (a nudge, then stronger, then near-answer, at a small point cost). Prefer to jump around? /submit <question#> <answer> answers any question directly. Check /score and /leaderboard any time. The capstone (Q26) is mentor-reviewed.

    Every command except /leaderboard replies privately to you, so you can play in a shared channel without spoiling anything for anyone else.

Good hunting. 🔎


Synthetic training data. Any resemblance to real companies, users, or infrastructure is coincidental. Built for defensive security education.

About

A hands-on SOC analyst investigation challenge. Load synthetic intrusion telemetry into your own Splunk or Elasticsearch and work 26 questions about attacker activity at a fictional company. Graded in Discord.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages