This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
TCP TimeArcs is a network traffic visualization tool that extends the TimeArcs technique for visualizing temporal relationships in TCP/IP packet data. It processes network traffic captures (CSV format) and creates interactive arc-based visualizations showing connections between IP addresses over time, with special emphasis on attack traffic patterns.
The project combines Python data processing scripts with browser-based D3.js visualizations, supporting multiple data loading strategies for different dataset sizes (from thousands to millions of packets).
Flow: Raw PCAP/CSV → Python processors → Structured JSON/CSV → Browser visualization
- tcp_data_loader.py - Basic single-file processor for small datasets
- tcp_data_loader_all.py - Comprehensive processor with full TCP flow state machine
- tcp_data_loader_split.py - Generates folder structure with individual flow files (v1.0 format)
- tcp_data_loader_chunked.py - Generates chunked flow files (v2.0 format, preferred)
All Python loaders implement the same TCP flow detection logic:
- 3-way handshake detection (SYN → SYN+ACK → ACK)
- Flow state tracking (establishment, data transfer, closing)
- Proper flow termination (FIN/RST handling)
- IP address mapping to numeric IDs
Two parallel visualization systems:
- Direct CSV upload with multiple files
- Synchronous rendering
- Best for small datasets (<100k packets)
- Features lensing/magnification (mouse-based time zooming)
- Entry point: attack_timearcs.html
- Folder-based loading with File System Access API
- Progressive rendering with Web Workers
- Handles large datasets (1M+ packets)
- Entry point: index.html
Key JavaScript Modules:
- folder_loader.js - Handles File System Access API for folder-based loading
- folder_integration.js - Bridges folder loader with visualization
- ip_bar_diagram.js / ip_arc_diagram.js - Main visualization renderers
- overview_chart.js - Time-series overview with clickable bins
- legends.js - Legend rendering and filtering
- sidebar.js - IP selection sidebar
- config.js - Global configuration (bin counts, batch sizes)
High-performance data ingestion using Web Workers and browser storage:
Architecture: Parse Workers → Aggregator → Tile Store (OPFS/IndexedDB)
- src/ingest/parse-worker.ts - CSV parsing in parallel workers
- src/ingest/aggregate-worker.ts - Aggregates parsed data into time tiles
- src/ingest/worker-pool.ts - Orchestrates worker pool and data flow
- src/storage/tile-store.ts - Persists time-binned data (OPFS preferred, IndexedDB fallback)
- src/api.ts - Public API for ingestion and querying
Data model: Time-series points stored in 60-second tiles (configurable) as Float64Arrays
# Generate folder structure with chunked flows (recommended)
python tcp_data_loader_chunked.py \
--data set1_first90_minutes.csv \
--ip-map combined_pcap_data_set5_compressed_ip_map.json \
--output-dir output_folder
# Legacy single-file format
python tcp_data_loader.py \
--data set1_first90_minutes.csv \
--ip-map combined_pcap_data_set5_compressed_ip_map.json \
--output output.json--data FILE # Input CSV (supports .csv.gz)
--ip-map FILE # JSON mapping of IP→numeric ID
--output-dir DIR # Output directory for split files
--max-records N # Limit packets processed (for testing)
--chunk-size N # Flows per chunk file (default: 200)Input CSV columns: timestamp,length,src_ip,dst_ip,src_port,dst_port,flags,protocol,[attack]
- timestamp: Unix epoch (seconds/milliseconds/microseconds auto-detected)
- flags: TCP flags as integer (e.g., 2=SYN, 16=ACK, 18=SYN+ACK)
- attack: Optional attack type label
IP mapping JSON: {"192.168.1.1": 1, "10.0.0.2": 2, ...}
output_folder/
├── manifest.json # Dataset metadata, format version
├── packets.csv # Minimal packets for timearcs rendering
├── flows/
│ ├── flows_index.json # Flow metadata with chunk references
│ ├── chunk_00000.json # Flows 0-199 with full packet data
│ ├── chunk_00001.json # Flows 200-399
│ └── ...
├── indices/
│ └── bins.json # Time-based bins for range queries
└── ips/
├── ip_stats.json # Per-IP statistics
├── flag_stats.json # TCP flag distribution
└── unique_ips.json # IP address list
- Load
manifest.json(instant) - Load
packets.csvprogressively with worker pool - Load
flows_index.json(contains chunk references) - Load flow chunks on-demand when user clicks flows
- Cache chunks for reuse (200 flows cached together)
Why chunked? Loading flow 42 also loads flows 0-199 in same chunk, so subsequent clicks in that range are instant.
GLOBAL_BIN_COUNT = 300- Time bins for overview chart (affects all visualizations)MAX_FLOW_LIST_ITEMS = 500- Performance limit for flow listsFLOW_LIST_RENDER_BATCH = 200- DOM update batch sizeFLOW_RECONSTRUCT_BATCH = 5000- Worker progress update frequency
isLensing- Toggle magnification mode (line 60)lensingMul- Magnification factor, default 5x (line 62)lensingRange- Width of magnified region (line 64)- Keyboard shortcut: Shift+L to toggle
- Implementation:
xScaleWithLensing()wrapper function (lines 871-927)
tileMs = 60_000- Time window per tile (milliseconds)maxPointsPerFlush = 50_000- Batch size for persistencenumWorkers- Auto-detected fromnavigator.hardwareConcurrency - 1
- No build step required for basic usage - open HTML directly in browser
- For TypeScript workers: Build with bundler (Vite/Webpack) if modifying src/
- Chrome/Edge required for File System Access API (folder loading)
- Firefox/Safari: Use legacy CSV upload mode
# Create small test dataset
python tcp_data_loader_chunked.py \
--data set1_first90_minutes.csv \
--ip-map combined_pcap_data_set5_compressed_ip_map.json \
--output-dir test_output \
--max-records 10000 \
--chunk-size 100// Enable debug logging (in browser console)
localStorage.setItem('debug', 'true');
location.reload();
// Access current data state
folderLoader.manifest // Dataset metadata
folderLoader.packets // Loaded packets
folderLoader.flowsIndex // Flow summaries
folderLoader.loadedFlows // Cached flow details (Map)All Python loaders implement identical 3-phase flow detection:
- Establishment: SYN → SYN+ACK → ACK
- Data Transfer: Any packets with PSH, ACK flags
- Closing: FIN → FIN+ACK → ACK, or immediate RST
Flows are identified by bidirectional 5-tuple: (src_ip, dst_ip, src_port, dst_port, protocol)
Auto-detection of timestamp format:
- If timestamp > 1e6: Treated as absolute time (minutes since epoch)
- Otherwise: Relative time (displayed as t=0, t=1, etc.)
- Supports seconds, milliseconds, microseconds - normalized to minutes
The folder loader auto-detects format versions:
- Check
manifest.json→format: "chunked"(v2.0) - Check flow index entries for
chunk_fileproperty (v2.0) - Fallback to individual files if neither present (v1.0)
- Small datasets (<100k packets): Use CSV upload (attack_timearcs.html)
- Medium (100k-1M packets): Use folder loading with default chunks
- Large (1M+ packets): Increase chunk size (500-1000) or use tile-based system (src/)
- Memory: Each packet ~100-200 bytes in memory, flows ~500 bytes
- Update event_type_mapping.json with new type and color
- Regenerate data with updated mapping
- Legend auto-updates from mapping file
- Reduce
GLOBAL_BIN_COUNTin config.js (fewer time bins) - Increase
FLOW_LIST_RENDER_BATCHfor larger DOM updates - Increase chunk size in data generation (fewer files to load)
- Edit
lensingMulin attack_timearcs.js (line 62) - Or use slider in UI (range: 2x to 200x)
- Adjust
lensingRangefor wider/narrower focus area
*_loader.py- Python data processors*_loader.js- JavaScript data loading modules*_worker.js/*.ts- Web Worker implementations*_diagram.js- D3.js visualization renderers*_integration.js- Module bridges/connectorsREADME_*.md- Feature-specific documentationset*_*.csv- Network traffic datasets (numbered by capture session)
| Feature | Chrome | Edge | Firefox | Safari |
|---|---|---|---|---|
| CSV Upload | ✅ | ✅ | ✅ | ✅ |
| Folder Loading | ✅ 86+ | ✅ 86+ | ❌ | ❌ |
| Worker Pool | ✅ | ✅ | ✅ | ✅ |
| OPFS Storage | ✅ 102+ | ✅ 102+ | ❌ | |
| Lensing | ✅ | ✅ | ✅ | ✅ |
Fallbacks: IDB for OPFS, CSV for folder loading