Archive aged InfluxDB windows to compressed Parquet in object storage, keep a durable record of exactly what moved, verify it, and rehydrate it on demand.
InfluxDB is excellent at recent data and expensive at old data. A year of 10-second IoT telemetry sits in the same hot storage as this morning's readings, inflating your disk bill, your backup window and your compaction load — even though nobody has queried last March since last March. The usual answer is a retention policy, which solves the cost problem by destroying the data. That is fine right up until an auditor, a regulator or a model-training job asks for last March.
influx-tier-mover is the middle path. It moves aged windows out of InfluxDB into Parquet on S3-compatible object storage, where the same points cost roughly a fifteenth as much per gigabyte and stay directly queryable by DuckDB, Athena, Spark or Trino. What makes it safe to actually run in production is the bookkeeping: every archived window gets a manifest recording its exact time range, row count, byte size, Parquet schema, per-object SHA-256 checksums and an archive-level checksum derived from them. That manifest is written into the bucket alongside the data and mirrored into a local SQLite index. Nothing is ever deleted from InfluxDB unless a verification pass has re-read the objects and reproduced those checksums, and six other guardrails agree.
- Chunked and streaming. A window is cut into fixed time slices; peak memory is one slice, not one window. Archiving a year with a 1h chunk never holds more than an hour of points.
- Hive-partitioned Parquet.
measurement=/year=/month=/day=keys, so partition pruning works with no external catalog. - Lossless round trip. Tags come back as tags, fields with their original types, timestamps to the microsecond. The tag/field role map travels inside the Parquet metadata, so an object is rehydratable even if you lose the manifest.
- Guarded expiry. Seven independent checks, no override for six of them.
- Cron-shaped.
planemits a machine-readable list of due windows from a policy file; every command speaks--jsonand stable exit codes. - No credentials in config. The policy file is validated to reject credential-shaped keys, so it is safe to commit.
A complete lifecycle — plan, archive, verify, inspect, rehydrate, and a guarded expire — run against the bundled fake InfluxDB and a local object store. This is real captured output; reproduce it with python examples/demo.py.
The same session as text
$ python examples/demo.py # seeded 10,368 points across 2 measurements
$ python -m influx_tier_mover plan --now 2026-03-01T00:00:00Z
sensors-cold: bucket=iot_hot measurements=environment,power older-than=16w5d window=1d lookback=17w1d retain-days=7
POLICY MEASUREMENT WINDOW START WINDOW END STATE
------------ ----------- ---------------- ---------------- -----
sensors-cold environment 2025-11-01 00:00 2025-11-02 00:00 due
sensors-cold environment 2025-11-02 00:00 2025-11-03 00:00 due
sensors-cold environment 2025-11-03 00:00 2025-11-04 00:00 due
sensors-cold power 2025-11-01 00:00 2025-11-02 00:00 due
sensors-cold power 2025-11-02 00:00 2025-11-03 00:00 due
sensors-cold power 2025-11-03 00:00 2025-11-04 00:00 due
6 window(s) due
$ python -m influx_tier_mover archive --measurement environment --start 2025-11-01 --end 2025-11-02
[#####---------------] 1/4 2025-11-01 00:00 432 rows 8.6 KiB
[##########----------] 2/4 2025-11-01 06:00 432 rows 8.6 KiB
[###############-----] 3/4 2025-11-01 12:00 432 rows 8.5 KiB
[####################] 4/4 2025-11-01 18:00 432 rows 8.6 KiB
archived iot_hot.environment.20251101T000000Z_20251102T000000Z: 1,728 rows, 4 object(s), 34.2 KiB (zstd), sha256 82ab69d83466…
$ python -m influx_tier_mover --quiet archive --policy sensors-cold --now 2026-03-01T00:00:00Z
archived iot_hot.environment.20251102T000000Z_20251103T000000Z: 1,728 rows, 4 object(s), 34.2 KiB (zstd), sha256 bfd9ee06baf0…
archived iot_hot.environment.20251103T000000Z_20251104T000000Z: 1,728 rows, 4 object(s), 34.2 KiB (zstd), sha256 68623441b0b8…
archived iot_hot.power.20251101T000000Z_20251102T000000Z: 1,728 rows, 4 object(s), 33.1 KiB (zstd), sha256 8dad61ed2680…
archived iot_hot.power.20251102T000000Z_20251103T000000Z: 1,728 rows, 4 object(s), 32.9 KiB (zstd), sha256 f2686aeba60a…
archived iot_hot.power.20251103T000000Z_20251104T000000Z: 1,728 rows, 4 object(s), 33.2 KiB (zstd), sha256 b1fabb250c74…
$ python -m influx_tier_mover list
ARCHIVE ID MEASUREMENT WINDOW ROWS SIZE STATE
----------------------------------------------------- ----------- ---------- ----- -------- --------
iot_hot.power.20251103T000000Z_20251104T000000Z power 2025-11-03 1,728 33.2 KiB archived
iot_hot.environment.20251103T000000Z_20251104T000000Z environment 2025-11-03 1,728 34.2 KiB archived
iot_hot.power.20251102T000000Z_20251103T000000Z power 2025-11-02 1,728 32.9 KiB archived
iot_hot.environment.20251102T000000Z_20251103T000000Z environment 2025-11-02 1,728 34.2 KiB archived
iot_hot.power.20251101T000000Z_20251102T000000Z power 2025-11-01 1,728 33.1 KiB archived
iot_hot.environment.20251101T000000Z_20251102T000000Z environment 2025-11-01 1,728 34.2 KiB archived
$ python -m influx_tier_mover verify --deep
iot_hot.power.20251103T000000Z_20251104T000000Z: OK (4 objects, 1728 rows, deep check)
iot_hot.environment.20251103T000000Z_20251104T000000Z: OK (4 objects, 1728 rows, deep check)
iot_hot.power.20251102T000000Z_20251103T000000Z: OK (4 objects, 1728 rows, deep check)
iot_hot.environment.20251102T000000Z_20251103T000000Z: OK (4 objects, 1728 rows, deep check)
iot_hot.power.20251101T000000Z_20251102T000000Z: OK (4 objects, 1728 rows, deep check)
iot_hot.environment.20251101T000000Z_20251102T000000Z: OK (4 objects, 1728 rows, deep check)
$ python -m influx_tier_mover stats
BUCKET MEASUREMENT ARCHIVES ROWS HOT COLD RATIO
------- ----------- -------- ----- --------- --------- -----
iot_hot environment 3 5,184 345.9 KiB 102.7 KiB 3.4x
iot_hot power 3 5,184 209.0 KiB 99.2 KiB 2.1x
total 6 archives, 24 objects, 10,368 rows (6 verified, 0 expired)
footprint 555.0 KiB hot -> 202.0 KiB cold (2.7x smaller)
cost/month USD 0.0001 hot vs USD 0.0000 cold
saving USD 0.0001/month, USD 0.00/year (97.6%)
projected at this ratio 1.0 TiB hot -> 372.6 GiB cold, saving USD 229.93/month per TiB
$ python -m influx_tier_mover rehydrate iot_hot.environment.20251101T000000Z_20251102T000000Z --target-bucket iot_restore
[#####---------------] 1/4 432 rows iot_hot.environment.20251101T000000Z_20251102T000000Z-00000.parquet
[##########----------] 2/4 432 rows iot_hot.environment.20251101T000000Z_20251102T000000Z-00001.parquet
[###############-----] 3/4 432 rows iot_hot.environment.20251101T000000Z_20251102T000000Z-00002.parquet
[####################] 4/4 432 rows iot_hot.environment.20251101T000000Z_20251102T000000Z-00003.parquet
wrote 1,728 rows from 4 object(s) into bucket 'iot_restore'
$ python -m influx_tier_mover rehydrate iot_hot.environment.20251101T000000Z_20251102T000000Z --target-bucket iot_restore
skipped iot_hot.environment.20251101T000000Z_20251102T000000Z: already rehydrated into 'iot_restore' at 2026-07-20T13:32:47+00:00 (pass --force to repeat)
$ python -m influx_tier_mover expire iot_hot.power.20251103T000000Z_20251104T000000Z
error: refusing to expire iot_hot.power.20251103T000000Z_20251104T000000Z:
[REFUSE] confirm: expire deletes source data and requires --confirm
[exit 4]
$ python -m influx_tier_mover expire iot_hot.power.20251103T000000Z_20251104T000000Z --confirm --dry-run
[pass] confirm: --confirm supplied
[pass] not-already-expired: no prior expiry recorded
[pass] has-objects: 4 object(s), 1728 row(s) archived
[pass] verified: verified at 2026-07-20T13:32:47+00:00 (deep)
[pass] retain-days: window ended 2025-11-04T00:00:00+00:00, outside the 7-day safety margin
[pass] checksum: objects still hash to b1fabb250c742318…
[pass] row-drift: source still holds exactly the 1728 archived rows
iot_hot.power.20251103T000000Z_20251104T000000Z: all guardrails pass, expiry would proceed
$ python -m influx_tier_mover expire iot_hot.power.20251103T000000Z_20251104T000000Z --confirm
[pass] confirm: --confirm supplied
[pass] not-already-expired: no prior expiry recorded
[pass] has-objects: 4 object(s), 1728 row(s) archived
[pass] verified: verified at 2026-07-20T13:32:47+00:00 (deep)
[pass] retain-days: window ended 2025-11-04T00:00:00+00:00, outside the 7-day safety margin
[pass] checksum: objects still hash to b1fabb250c742318…
[pass] row-drift: source still holds exactly the 1728 archived rows
expired iot_hot.power.20251103T000000Z_20251104T000000Z: deleted 1,728 rows from InfluxDB
$ # iot_restore now holds 1,728 rehydrated pointsThe compression ratio in that run, measured rather than estimated:
There is no package to install. Clone the repo and run it as a module.
git clone https://github.com/taskautomation-org/influx-tier-mover.git
cd influx-tier-mover
python -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python -m influx_tier_mover --helpPython 3.11 or newer. Runtime dependencies are pyarrow, boto3, pydantic and PyYAML; requirements-dev.txt adds pytest, ruff and matplotlib for the test suite and the chart above.
Credentials are read only from the environment. They are never accepted in tiering.yaml — the config validator treats a credential-shaped key (token, secret_access_key, password, and a dozen others, at any nesting depth) as a hard error, so the policy file is safe to commit to the same repo as your code.
export INFLUX_TOKEN='...' # InfluxDB API token
# Object storage uses the standard AWS credential chain: environment,
# ~/.aws/credentials, or an instance/pod role. Nothing extra to configure.
export AWS_ACCESS_KEY_ID='...'
export AWS_SECRET_ACCESS_KEY='...'For Cloudflare R2, MinIO or Backblaze B2, set destination.endpoint_url in the config and use the same two environment variables.
Clients are built lazily, so the commands that only read the policy file and the local catalog — plan, list, show and stats — run with no credentials in the environment at all. That is deliberate: plan is the command you run first and the one you run from a locked-down scheduler, and it should not demand a token it never uses.
destination.provider: local archives to a directory instead of a bucket. Combined with the bundled demo it lets you exercise the whole tool before you point it at anything real:
.venv/bin/python examples/demo.py --keep ./demo-archive
find ./demo-archive/bucket -name '*.parquet' | head -3tiering.yaml (a complete, working example ships in the repo root):
version: 1
source:
url: http://localhost:8086
org: acme
bucket: iot_hot
destination:
provider: s3
bucket: telemetry-archive
prefix: influx-tier
# endpoint_url: https://<account-id>.r2.cloudflarestorage.com
region: auto
archive:
compression: zstd
compression_level: 5
chunk: 1h # one Parquet object per hour of source data
timestamp_precision: ns
policies:
- name: sensors-cold
bucket: iot_hot
measurements: [environment, power]
min_age: 90d # only touch data older than this
window: 1d # archive one day per run
max_lookback: 365d # never look further back than this
retain_days: 7 # extra safety margin before expire may delete
catalog:
path: ./tiering.db
cost:
hot_usd_per_gb_month: 0.23
cold_usd_per_gb_month: 0.015plan is pure computation over the policy and the catalog — it touches neither InfluxDB nor the bucket, so it is safe to run anywhere, and --now makes it deterministic for testing and backfills.
$ python -m influx_tier_mover plan --now 2026-03-01T00:00:00Z
sensors-cold: bucket=iot_hot measurements=environment,power older-than=16w5d window=1d lookback=17w1d retain-days=7
POLICY MEASUREMENT WINDOW START WINDOW END STATE
------------ ----------- ---------------- ---------------- -----
sensors-cold environment 2025-11-01 00:00 2025-11-02 00:00 due
sensors-cold environment 2025-11-02 00:00 2025-11-03 00:00 due
sensors-cold environment 2025-11-03 00:00 2025-11-04 00:00 due
sensors-cold power 2025-11-01 00:00 2025-11-02 00:00 due
sensors-cold power 2025-11-02 00:00 2025-11-03 00:00 due
sensors-cold power 2025-11-03 00:00 2025-11-04 00:00 due
6 window(s) dueWindow boundaries live on an absolute grid anchored to the Unix epoch, not on "now". A daily policy therefore always produces midnight-to-midnight UTC windows, so the archive id for a given day is identical whether the job ran on schedule, ran six hours late, or is backfilling from last year. That is what makes plan safe to re-run and archive idempotent.
Either one explicit window:
$ python -m influx_tier_mover archive --measurement environment --start 2025-11-01 --end 2025-11-02
[#####---------------] 1/4 2025-11-01 00:00 432 rows 8.6 KiB
[##########----------] 2/4 2025-11-01 06:00 432 rows 8.6 KiB
[###############-----] 3/4 2025-11-01 12:00 432 rows 8.5 KiB
[####################] 4/4 2025-11-01 18:00 432 rows 8.6 KiB
archived iot_hot.environment.20251101T000000Z_20251102T000000Z: 1,728 rows, 4 object(s), 34.2 KiB (zstd), sha256 82ab69d83466……or everything a policy says is due, which is the form you put in cron:
$ python -m influx_tier_mover --quiet archive --policy sensors-cold
archived iot_hot.environment.20251102T000000Z_20251103T000000Z: 1,728 rows, 4 object(s), 34.2 KiB (zstd), sha256 bfd9ee06baf0…
archived iot_hot.environment.20251103T000000Z_20251104T000000Z: 1,728 rows, 4 object(s), 34.2 KiB (zstd), sha256 68623441b0b8…
archived iot_hot.power.20251101T000000Z_20251102T000000Z: 1,728 rows, 4 object(s), 33.1 KiB (zstd), sha256 8dad61ed2680…Re-running is a no-op — a window already in the catalog is skipped unless you pass --force. --dry-run does every byte of the work (query, convert, compress, checksum) and uploads nothing, which is a genuine size estimate rather than a guess.
Shallow verification (the default) re-reads every object and checks size and SHA-256. --deep additionally parses the Parquet, counts rows per object and diffs the file's schema against the manifest.
$ python -m influx_tier_mover verify --deep
iot_hot.power.20251103T000000Z_20251104T000000Z: OK (4 objects, 1728 rows, deep check)
iot_hot.environment.20251103T000000Z_20251104T000000Z: OK (4 objects, 1728 rows, deep check)A failure names the object and what is wrong, and sets exit code 3:
$ python -m influx_tier_mover verify
iot_hot.environment.20251101T000000Z_20251102T000000Z: FAILED with 2 problem(s)
- influx-tier/measurement=environment/…-00000.parquet: size is 9 bytes, manifest says 8797
- influx-tier/measurement=environment/…-00000.parquet: sha256 is 40df3b6a4e69…, manifest says d2146f8e0035…show prints one manifest in full — the whole record of a single archived window:
$ python -m influx_tier_mover show iot_hot.environment.20251101T000000Z_20251102T000000Z
archive iot_hot.environment.20251101T000000Z_20251102T000000Z
bucket iot_hot
measurement environment
window 2025-11-01T00:00:00+00:00 .. 2025-11-02T00:00:00+00:00
rows 1,728 in 4 object(s)
size 34.2 KiB compressed (zstd), 115.3 KiB in memory
checksum 82ab69d834666ea3f8a8256b72c4db4ade156d4d6dac17ec5eca1cd817c8ce96
state archived
schema
time _time timestamp[ns, tz=UTC]
tag room string
tag sensor string
tag site string
field co2_ppm int64
field humidity_pct double
field temperature_c double
objects
influx-tier/measurement=environment/year=2025/month=11/day=01/iot_hot.environment.20251101T000000Z_20251102T000000Z-00000.parquet rows=432 size=8.6 KiB sha256=d2146f8e0035…
influx-tier/measurement=environment/year=2025/month=11/day=01/iot_hot.environment.20251101T000000Z_20251102T000000Z-00001.parquet rows=432 size=8.6 KiB sha256=1935e4be3a9f…
influx-tier/measurement=environment/year=2025/month=11/day=01/iot_hot.environment.20251101T000000Z_20251102T000000Z-00002.parquet rows=432 size=8.5 KiB sha256=fb9f1b3e25c6…
influx-tier/measurement=environment/year=2025/month=11/day=01/iot_hot.environment.20251101T000000Z_20251102T000000Z-00003.parquet rows=432 size=8.6 KiB sha256=fee4f9e4a2b6…list and stats summarise the catalog:
$ python -m influx_tier_mover stats
BUCKET MEASUREMENT ARCHIVES ROWS HOT COLD RATIO
------- ----------- -------- ----- --------- --------- -----
iot_hot environment 3 5,184 345.9 KiB 102.7 KiB 3.4x
iot_hot power 3 5,184 209.0 KiB 99.2 KiB 2.1x
total 6 archives, 24 objects, 10,368 rows (6 verified, 0 expired)
footprint 555.0 KiB hot -> 202.0 KiB cold (2.7x smaller)
cost/month USD 0.0001 hot vs USD 0.0000 cold
saving USD 0.0001/month, USD 0.00/year (97.6%)
projected at this ratio 1.0 TiB hot -> 372.6 GiB cold, saving USD 229.93/month per TiBThe projected line restates your measured ratio at a scale worth budgeting for; it invents no data, it just scales the observation. Cost figures use the prices in your config and cover storage only — not request charges or egress.
$ python -m influx_tier_mover rehydrate iot_hot.environment.20251101T000000Z_20251102T000000Z --target-bucket iot_restore
[#####---------------] 1/4 432 rows iot_hot.environment.20251101T000000Z_20251102T000000Z-00000.parquet
[####################] 4/4 432 rows iot_hot.environment.20251101T000000Z_20251102T000000Z-00003.parquet
wrote 1,728 rows from 4 object(s) into bucket 'iot_restore'Rehydration is idempotent. The manifest records where and when a window was restored, in the object store as well as locally, so a repeat is refused rather than silently double-written:
$ python -m influx_tier_mover rehydrate iot_hot.environment.20251101T000000Z_20251102T000000Z --target-bucket iot_restore
skipped iot_hot.environment.20251101T000000Z_20251102T000000Z: already rehydrated into 'iot_restore' at 2026-07-20T13:17:46+00:00 (pass --force to repeat)Restoring into a scratch bucket first (as above) is the recommended drill — see the disaster-recovery section of the operational runbook.
Expiry is the only destructive operation, and it is deliberately the hardest to trigger. Without --confirm it refuses and exits 4:
$ python -m influx_tier_mover expire iot_hot.power.20251103T000000Z_20251104T000000Z
error: refusing to expire iot_hot.power.20251103T000000Z_20251104T000000Z:
[REFUSE] confirm: expire deletes source data and requires --confirm--dry-run shows every check so you can see exactly what would happen:
$ python -m influx_tier_mover expire iot_hot.power.20251103T000000Z_20251104T000000Z --confirm --dry-run
[pass] confirm: --confirm supplied
[pass] not-already-expired: no prior expiry recorded
[pass] has-objects: 4 object(s), 1728 row(s) archived
[pass] verified: verified at 2026-07-20T13:17:46+00:00 (deep)
[pass] retain-days: window ended 2025-11-04T00:00:00+00:00, outside the 7-day safety margin
[pass] checksum: objects still hash to b1fabb250c742318…
[pass] row-drift: source still holds exactly the 1728 archived rows
iot_hot.power.20251103T000000Z_20251104T000000Z: all guardrails pass, expiry would proceed| Command | What it does |
|---|---|
plan |
Compute which windows a policy says are due. Reads nothing but the config and the catalog. |
archive |
Archive one explicit window, or every window a policy says is due. |
verify |
Re-read archived objects and check them against their manifest. |
show |
Print one manifest in full, from the catalog or from object storage. |
list |
List archives, filtered by bucket, measurement, time range or lifecycle state. |
stats |
Rows, bytes and estimated cost per tier, with a per-TiB projection. |
rehydrate |
Write an archived window back into an InfluxDB bucket. |
expire |
Delete a verified window from InfluxDB, behind seven guardrails. |
reindex |
Rebuild the local SQLite catalog from the manifests in object storage. |
| Flag | Default | Meaning |
|---|---|---|
-c, --config PATH |
tiering.yaml |
Policy file to load. |
-q, --quiet |
off | Suppress per-chunk progress output. |
--version |
— | Print the version and exit. |
Every reporting command also accepts --json, which replaces the human table with a machine-readable document on stdout.
| Command | Flag | Meaning |
|---|---|---|
archive |
--measurement, --start, --end |
Archive one explicit window. See time arguments for accepted formats. |
--bucket |
Override source.bucket. |
|
--policy NAME |
Archive every window that policy says is due. | |
--now |
Reference time for --policy. Defaults to now. |
|
--limit N |
Stop after N windows. | |
--force |
Re-archive windows already in the catalog. | |
--dry-run |
Do all the work; upload and index nothing. | |
verify |
[archive_id] |
Verify one archive. Omit to verify all of them. |
--deep |
Also parse the Parquet and check rows and schema. | |
rehydrate |
--target-bucket |
Destination bucket. Defaults to the source bucket. |
--precision |
ns, us, ms or s. Defaults to the archived precision. |
|
--dry-run |
Read and convert; write nothing. | |
--force |
Rehydrate again even though the manifest says it was already done. | |
expire |
--confirm |
Required. Expiry deletes source data. |
--retain-days N |
Refuse if the window ended fewer than N days ago. Default 7. | |
--allow-drift |
Accept a row-count difference between source and archive. | |
--dry-run |
Report every guardrail; delete nothing. | |
plan |
--policy NAME |
Restrict to one policy. Repeatable. |
--now, --all |
Reference time; include already-archived windows. | |
list |
--bucket, --measurement, --since, --until |
Filters. |
--state |
all, archived, verified, rehydrated or expired. |
|
--limit N |
Maximum rows. Default 50. | |
show |
--from-store |
Read the manifest from object storage rather than the local index. |
--start, --end, --now, --since and --until all accept the same three forms:
| Form | Example | Meaning |
|---|---|---|
| RFC 3339 | 2025-11-01T00:00:00Z |
An absolute instant. Naive values are read as UTC. |
| Date | 2025-11-01 |
Midnight UTC on that date. |
| Relative offset | -90d, -36h, -1h30m |
That duration before now. |
Relative offsets work as ordinary arguments — --start -90d — on every supported
Python. (The --start=-90d form is equivalent, if you prefer being explicit.)
| Code | Meaning |
|---|---|
0 |
Success. |
1 |
Unexpected error (unknown archive id, missing object, InfluxDB failure). |
2 |
Configuration problem (bad policy file, unparseable time, missing argument). |
3 |
Verification failed. |
4 |
A destructive operation was refused by a guardrail. |
A refused expire is exit 4, not exit 1, precisely so a CronJob can tell "the tool did its job and declined" apart from "the tool broke".
| Key | Type | Default | Meaning |
|---|---|---|---|
version |
int | 1 |
Policy file format version. Only 1 is accepted. |
source.url |
str | http://localhost:8086 |
InfluxDB base URL. |
source.org |
str | "" |
Organisation name. |
source.bucket |
str | "" |
Default source bucket. |
source.timeout_seconds |
float | 60 |
Per-request HTTP timeout. |
destination.provider |
s3 | local |
s3 |
Object store backend. |
destination.bucket |
str | "" |
Bucket name (provider s3). |
destination.prefix |
str | influx-tier |
Key prefix for data and manifests. |
destination.endpoint_url |
str | — | Non-AWS S3 endpoint (R2, MinIO, B2). |
destination.region |
str | — | Region, or auto for R2. |
destination.path |
str | — | Directory (provider local). |
destination.addressing_style |
auto | path | virtual |
auto |
S3 addressing style. |
archive.compression |
zstd | snappy | gzip | none |
zstd |
Parquet codec. |
archive.compression_level |
int 1–22 | — | Codec level. |
archive.chunk |
duration | 1h |
Source time per Parquet object. This is the memory knob. |
archive.timestamp_precision |
ns | us | ms | s |
ns |
Precision recorded and used on rehydrate. |
archive.row_group_size |
int | 100000 |
Parquet row-group size. |
policies[].name |
str | — | Policy name, used by --policy. |
policies[].bucket |
str | — | Source bucket. |
policies[].measurements |
list | — | Measurements the policy covers. At least one. |
policies[].min_age |
duration | 90d |
Only archive data older than this. |
policies[].window |
duration | 1d |
Size of one archived window. |
policies[].max_lookback |
duration | 365d |
Oldest data the policy will consider. |
policies[].retain_days |
int | 7 |
Default safety margin for expire. |
catalog.path |
str | ./tiering.db |
SQLite index location. |
cost.hot_usd_per_gb_month |
float | 0.23 |
Hot storage price for stats. |
cost.cold_usd_per_gb_month |
float | 0.015 |
Cold storage price for stats. |
Durations are InfluxDB-style literals built from <int><unit> pairs where unit is s, m, h, d or w: 90d, 12h, 1h30m, 2w.
<prefix>/measurement=<name>/year=YYYY/month=MM/day=DD/<archive-id>-<seq>.parquet
<prefix>/_manifests/<archive-id>.json
Archive ids are deterministic: <bucket>.<measurement>.<start>_<end>, for example iot_hot.environment.20251101T000000Z_20251102T000000Z. The same window always yields the same id, which is what makes re-runs idempotent.
The Hive-style partitioning is not decoration. Point DuckDB at the prefix and the date filter prunes partitions with no extra catalog:
SELECT site, avg(temperature_c)
FROM read_parquet('s3://telemetry-archive/influx-tier/measurement=environment/**/*.parquet',
hive_partitioning = true)
WHERE year = '2025' AND month = '11'
GROUP BY site;| Field | Meaning |
|---|---|
archive_id |
Deterministic identifier for the window. |
source_bucket, measurement |
Where the data came from. |
window_start, window_end |
Half-open [start, end) range, RFC 3339 UTC. |
chunk_seconds |
Time per Parquet object at archive time. |
compression, timestamp_precision |
How the objects were written. |
row_count, byte_size |
Rows archived and compressed bytes stored. |
uncompressed_bytes |
In-memory Arrow footprint; the "hot" side of the cost model. |
checksum |
SHA-256 over the sorted key:sha256 pairs of every object. |
objects[] |
Per-object key, chunk range, rows, bytes and SHA-256. |
schema[] |
Every column with its Arrow type and its role (time, tag, field). |
created_at |
When the archive was written. |
verified_at, verified_depth |
When it last passed verification, and how deeply. |
rehydrated_at, rehydrated_into |
Restore bookkeeping; the basis of idempotency. |
expired_at, notes |
Set once the source window has been deleted. |
Because the archive checksum is derived from the per-object checksums rather than from concatenated bodies, it can be recomputed from the manifest alone — and any object being added, removed, renamed or altered changes it.
| Guardrail | Refuses when |
|---|---|
confirm |
--confirm was not passed. |
not-already-expired |
The manifest already records an expiry. |
has-objects |
The archive holds no Parquet objects at all. |
verified |
The archive has never passed verify. |
retain-days |
The window is newer than the retention safety margin. |
checksum |
Re-reading the objects does not reproduce the manifest checksum. |
row-drift |
InfluxDB now holds a different number of points in the window than were archived. Relaxable with --allow-drift. |
All seven are evaluated on every run, even after the first failure, so one command shows you the complete picture. Only row-drift can be relaxed, and only with an explicit flag. The verified and checksum guardrails re-read the manifest from the object store, never from the local SQLite index, so a stale or tampered local catalog cannot authorise a delete.
┌──────────────┐
tiering.yaml ────>│ plan │──> due windows (JSON or table)
+ catalog └──────────────┘
┌──────────────┐ chunk by time slice
InfluxDB ────────>│ archive │──> Arrow table ──> Parquet ──> object store
└──────┬───────┘ │
│ manifest (JSON) ────────────────────────────>│
└────────────> SQLite catalog (index) │
┌──────────────┐ │
│ verify │<─────────────────────────────────────┘
└──────┬───────┘ checksums, sizes, rows, schema
│ stamps verified_at
┌──────▼───────┐
│ expire │ 7 guardrails ──> delete from InfluxDB
└──────────────┘
┌──────────────┐
│ rehydrate │ object store ──> Arrow ──> InfluxDB
└──────────────┘
Two client protocols, four methods each. InfluxClient is count / read / write / delete; ObjectStore is put / get / exists / size / list / delete. Everything above those interfaces is backend-agnostic, which is why the entire test suite runs with an in-memory InfluxDB and an in-memory bucket and makes zero network calls — and why archiving to a local directory is a supported mode rather than a test fixture.
Chunking is the memory story. archive splits [start, end) into fixed slices and processes one at a time: query the slice, build an Arrow table, write Parquet, upload, drop it, move on. The first slice that returns rows fixes the schema for the whole archive; later slices are written against that schema, so every object in an archive is mutually compatible and a reader can concatenate them without reconciling types. A slice that cannot fit the schema raises SchemaDriftError rather than silently writing an incompatible file.
Two stores, one authoritative. The manifest JSON in the bucket is the source of truth — if you lose the machine running this tool, everything needed to verify and rehydrate is still in the bucket, and reindex rebuilds the local index from it. SQLite exists so list, show, stats and plan are instant and work offline. Writes go to the object store first, so a crash between the two leaves a recoverable archive rather than a dangling index row.
Role metadata travels with the data. Parquet does not know which columns were tags and which were fields, so the role map is written into the Arrow schema's key-value metadata. Every archived object is self-describing: you can rehydrate from the Parquet alone. Foreign Parquet with no such metadata is still readable via a positional fallback.
.venv/bin/pip install -r requirements-dev.txt
.venv/bin/python -m pytest --cov=influx_tier_mover --cov-report=term-missing
.venv/bin/ruff check . && .venv/bin/ruff format --check .434 tests, 98% branch coverage, no network. The suite asserts the properties that matter rather than the implementation: that a chunked archive of a multi-chunk window round-trips to exactly the points that went in, that every field type and microsecond timestamp survives, that verify catches a corrupted object, a resized object, a missing object, an unreadable body, a row-count mismatch and a schema change, that a second rehydrate writes nothing, and — one test per guardrail — that expire refuses in each unsafe condition.
CI runs the suite on Python 3.11, 3.12 and 3.13 as well as 3.14. That matrix is not ceremony: argparse's handling of option values beginning with - changed in 3.14, and a relative offset like --start -30d parses on the newest interpreter while failing on every older one. Test on the oldest version you support.
Background on the lifecycle patterns this tool implements:
- Exporting cold partitions to S3 before expiry — the archive-then-expire ordering that the guardrails here enforce mechanically.
- Cold-storage archival to object store — why object storage is the right destination for aged time-series, and what you give up.
- Storage tier migration automation — designing the hot/warm/cold policy that a
tiering.yamlencodes. - Bucket architecture and tiering boundaries — choosing where one tier ends and the next begins before you automate the move.
- Automating bucket expiration with scheduled purge tasks — how scheduled deletion goes wrong, and what to check first.
The operational runbook covers first archive, scheduling, disaster recovery and cost tuning.
See CONTRIBUTING.md. Bug reports with a failing test case are the most useful thing you can send.
MIT — see LICENSE.
