Purpose: This document defines the complete expected NextSQL product after the development roadmap in
TODO.mdis implemented and production-gated.
TODO.mdis the execution tracker and source of implementation status.
TODO.mdalso defines sequencing, dependencies, and phase gates.
SKILLS.mddefines agent engineering behavior.PROJECT.mddefines the finished product contract.A feature described here is part of the intended final product unless explicitly marked Deferred, Optional, Research, or Experimental. Its presence in this document does not mean it is already implemented; implementation truth remains in
TODO.md.
NextSQL is a high-performance, secure-by-default, encrypted-by-default, durable, native multimodel database platform written primarily in Go.
NextSQL is not a compatibility layer for:
- PostgreSQL
- MySQL
- MariaDB
- MongoDB
- Elasticsearch
- Redis
- external vector databases
NextSQL has its own:
- storage format
- SQL dialect
- wire protocol
- drivers
- optimizer
- transaction engine
- encryption architecture
- catalog
- multimodel execution engine
- administration interfaces
- development tooling
The final product combines:
Relational SQL
+ Native Binary JSON
+ Full-Text Search
+ Vector Search
+ Geospatial
+ Hybrid Retrieval
+ WORKFLOW / TRIGGER / SCHEDULE / TASK
+ CDC / Change Streams
+ Native Partitioning
+ Read Scaling
under:
one optimizer
one transaction model
one WAL/recovery model
one security model
one native protocol
The core architectural principle is:
One engine, one optimizer, one transaction model, one durability model, and one security model across every data modality.
The completed NextSQL product family consists of:
NextSQL Engine
Native database runtime:
SQL + JSON + FTS + Vector + Hybrid + Geo + Workflow + CDC
NextSQL CLI
Headless administration, automation, migrations, backup,
restore, maintenance, diagnostics, cluster and security operations
NextSQL Bench
Correctness-aware official benchmark and SLO measurement suite
NextSQL Drivers
Go / Node.js / TypeScript / Bun / PHP / Python / Ruby
plus future officially supported SDKs
NextSQL Admin
One application, three modes:
Setup — install / initialize / upgrade / repair / uninstall
Operations — server / cluster / security / backup / maintenance UI
Studio — native NextSQL database development IDE
All products must use official NextSQL interfaces and server truth.
No GUI or AI layer may bypass:
- authentication
- RBAC
- realm/database isolation
- catalog authority
- transaction rules
- server validation
- encryption rules
Always optimize in this order:
1. Correctness
2. Durability
3. Security
4. Data integrity
5. Availability
6. Predictable latency
7. Throughput
8. Resource efficiency
9. Developer experience
10. Additional features
Reject any optimization that weakens:
- ACID correctness
- WAL durability
- fsync guarantees
- encryption
- authentication
- authorization
- checksums/integrity authentication
- replication safety
- crash recovery
- realm/database isolation
- vector recall without disclosure
NextSQL must be fast with production safety enabled.
Official benchmarks keep enabled:
fsync
WAL
encryption
checksums/authentication
MVCC
authentication
authorization
durability
unless explicitly labeled experimental and excluded from official SLO claims.
NextSQL Clients
│
┌─────────────────────┼─────────────────────┐
│ │ │
Drivers CLI NextSQL Admin
│ │ │
└─────────────────────┼─────────────────────┘
│
Native NSQL Wire Protocol
│
TLS / mTLS
│
Authentication / Identity
│
RBAC / Realm/Database Policy
│
SQL Parser
│
Binder / Catalog
│
Logical Query Planner
│
Adaptive Cost-Based Optimizer
│
Vectorized / Parallel Executor
│
┌───────────────┬─────────────┼──────────────┬───────────────┐
│ │ │ │ │
▼ ▼ ▼ ▼ ▼
Relational JSON Full-Text Vector Geo
Engine Engine Engine Engine Engine
│ │ │ │ │
B+Tree Binary JSON Inverted Flat / Spatial
/ indexes Path indexes BM25/FTS HNSW/IVF Indexes
│ │ │ │ │
└───────────────┴─────────────┴──────────────┴───────┬───────┘
│
Hybrid Optimizer
│
Unified Transactions
MVCC + Locks
│
┌─────────────────┴─────────────┐
│ │
UNDO WAL
│ │
└─────────────────┬─────────────┘
│
Buffer / Storage
│
Authenticated Encryption
│
SSD / NVMe
Automation and change processing integrate with the same engine:
WORKFLOW
├── manual RUN
├── TRIGGER
└── SCHEDULE
│
▼
TASK
Committed WAL
│
▼
CDC / Change Streams
Distributed architecture:
NextSQL Endpoint
│
┌───────────┴───────────┐
│ │
Leader Followers
│ │
└───────────┬───────────┘
│
Raft Quorum
│
Synchronous Durability
Final intended distributed direction:
single Raft leader for writes
+ synchronous quorum durability
+ followers
+ optional follower reads
+ native local partitioning
Automatic distributed sharding is not part of the core product contract. Multi-primary writes are an explicitly requested future extension, but remain disabled until their versioned replicated conflict model, recovery semantics, and independent failure gates are implemented. The current production write path remains single-leader Raft.
NextSQL is primarily implemented in Go.
Representative repository structure:
nextsql/
├── cmd/
│ ├── nextsqld/
│ ├── nextsql/
│ └── nextsql-bench/
│
├── internal/
│ ├── protocol/
│ ├── auth/
│ ├── security/
│ ├── crypto/
│ ├── sql/
│ │ ├── lexer/
│ │ ├── parser/
│ │ ├── ast/
│ │ ├── binder/
│ │ ├── planner/
│ │ └── optimizer/
│ ├── catalog/
│ ├── executor/
│ ├── storage/
│ ├── txn/
│ ├── wal/
│ ├── recovery/
│ ├── json/
│ ├── fulltext/
│ ├── vector/
│ ├── geo/
│ ├── scheduler/
│ ├── workflow/
│ ├── task/
│ ├── cdc/
│ ├── partition/
│ ├── backup/
│ ├── replication/
│ ├── maintenance/
│ ├── metrics/
│ └── config/
│
├── drivers/
│ ├── go/
│ ├── node/
│ ├── bun/
│ └── php/
│
├── admin/
│ ├── setup/
│ ├── ops/
│ └── studio/
├── intelligence/
├── tests/
├── docs/
└── go.mod
Package boundaries must remain narrow and explicit.
Avoid cyclic dependencies.
Do not serialize raw Go structs directly to disk.
All persistent formats and wire formats must be deterministic and versioned.
A single table may contain relational, JSON, vector, temporal, and geospatial data.
Example:
CREATE TABLE products (
id UUID PRIMARY KEY DEFAULT UUID(),
account_id UUID NOT NULL,
name STRING NOT NULL,
description TEXT,
price DECIMAL(12,2),
metadata JSON,
embedding VECTOR<F32,1536>,
location POINT,
created_at TIMESTAMPTZ DEFAULT NOW()
);Indexes can span different modalities:
CREATE INDEX ix_category
ON products(metadata.category);
CREATE FULLTEXT INDEX ix_description
ON products(description);
CREATE VECTOR INDEX ix_embedding
ON products(embedding)
USING HNSW;
CREATE SPATIAL INDEX ix_location
ON products(location);Example relational query:
SELECT id, name, price
FROM products
WHERE price BETWEEN 1000 AND 5000;Example JSON query:
SELECT *
FROM products
WHERE metadata.category = 'electronics';Example full-text query:
SELECT *
FROM products
SEARCH description FOR 'wireless noise cancelling'
LIMIT 20;Example vector query:
SELECT id, name
FROM products
NEAREST embedding TO $query
LIMIT 20;Example hybrid query:
SELECT id, name, price
FROM products
WHERE metadata.category = 'headphones'
AND price <= 15000
SEARCH description FOR 'wireless noise cancelling'
NEAREST embedding TO $query
LIMIT 20;The optimizer must treat the hybrid query as one physical planning problem, not independent database calls.
NextSQL uses its own SQL dialect.
Its standards baseline is ISO/IEC 9075:2023, including SQL/CLI, SQL/PSM,
SQL/MED, SQL/Schemata, SQL/MDA, and SQL/PGQ as applicable design references.
ISO/IEC 9579:2000 RDA guides remote database protocol principles; remote
production transport uses TCP with TLS 1.3, and text uses Unicode/UTF-8. This
baseline does not itself claim formal conformance or replace the native NSQL
protocol. Planned standard areas are not shipped until TODO.md, capability
metadata, tests, and matching-version documentation say so. See
docs/standards.md.
The final SQL surface includes conventional relational operations while remaining NextSQL-native rather than compatibility-driven.
CREATE TABLE
ALTER TABLE
DROP TABLE
CREATE INDEX
DROP INDEX
REBUILD INDEX
INSERT
UPDATE
DELETE
SELECT
BEGIN
COMMIT
ROLLBACK
UPSERT
DDL and maintenance evolve according to the native catalog/storage model.
REBUILD INDEX ... ONLINE is only considered part of the final production surface when concurrent-write correctness is fully proven.
Final SQL includes:
DISTINCTHAVINGCASE- joins
- aggregations
- ordering
- limit/offset
- subqueries
- correlated subqueries
- derived tables
- CTEs
- recursive CTEs
- set operations
- window functions
RETURNING
Set operations:
UNION
UNION ALL
INTERSECT
EXCEPT
Window functions include:
ROW_NUMBER
RANK
DENSE_RANK
LAG
LEAD
FIRST_VALUE
LAST_VALUE
COUNT OVER
SUM OVER
AVG OVER
MIN OVER
MAX OVER
Native atomic UPSERT must support unique-index conflict semantics.
Support:
INSERT ... RETURNING ...
UPDATE ... RETURNING ...
DELETE ... RETURNING ...Results stream over NSQL rather than requiring full materialization.
Final standard function surface includes:
LOWERUPPERLENGTHSUBSTRINGTRIMLTRIMRTRIMREPLACECONCATSTARTS_WITHENDS_WITHCONTAINS
ABSROUNDCEILFLOORPOWERSQRTMOD
COALESCENULLIFGREATESTLEAST
EXTRACTor canonical NextSQL equivalentDATE_TRUNCor canonical equivalentDATE_ADDDATE_DIFF
JSON_GETJSON_SETJSON_REMOVEJSON_CONTAINSJSON_ARRAY_LENGTHJSON_TYPE
Final index capabilities include:
- clustered primary B+Tree
- secondary B+Tree
- UNIQUE
- covering indexes /
INCLUDE - partial indexes
- expression indexes
- JSON path indexes
- full-text indexes
- vector indexes
- spatial indexes
Index-only scans should be used where valid.
The optimizer is deterministic and cost-based.
Pipeline:
SQL
↓
AST
↓
Binding
↓
Logical Plan
↓
Rewrite
↓
Physical Alternatives
↓
Cost Model
↓
Physical Plan
↓
Vectorized / Parallel Execution
It includes:
- predicate pushdown
- projection pushdown
- constant folding
- limit pushdown
- index selection
- join simplification
- join reordering
- column pruning
- partition pruning
- segment pruning
- top-N optimization
- covering/index-only selection
- partial-index implication
- expression-index matching
- subquery flattening
- subquery decorrelation
- CTE inline/materialize decisions
- hybrid structured/FTS/vector plan selection
Statistics include:
- row count
- NULL ratio
- NDV
- min/max
- histograms
- most common values
- correlation
- index selectivity
- segment statistics
- vector statistics
- runtime estimated-vs-actual feedback
The optimizer must not depend on an LLM.
Primary execution is batch/vector-oriented for suitable operations.
Representative batch sizes:
1024
2048
4096
with benchmarks deciding actual defaults.
Support:
- vector filters
- vector projection
- batch decoding
- hash aggregation
- hash join
- merge join
- index scan
- parallel scan
- parallel aggregation
- parallel joins
- parallel index construction
- parallel vector distance computation
All work goes through explicit bounded schedulers.
Never spawn unbounded goroutines per query.
Every query has bounded:
- CPU workers
- memory
- disk spill
- I/O
- execution time
- result size
Large query results must stream.
Never:
materialize multi-GB result entirely in memory
→ then send
Instead:
executor batch
→ protocol
→ client ACK/backpressure
→ next batch
Slow clients must not grow server memory without bound.
Cancellation must propagate through execution.
Default logical page size:
16 KiB
Persistent structures use deterministic versioned binary encoding.
Primary row storage uses clustered B+Tree organization.
Features include:
- slotted pages
- variable-length rows
- page validation
- page allocation
- buffer management
- B+Tree insert
- lookup
- delete
- range scan
- split
- merge
- rebalance
- root collapse
- page reclamation
- durable freelist
- orphan detection
- restart-safe reuse
- storage integrity checking
Known corrupted records must never be silently returned.
Final engine includes native maintenance rather than permanent storage leakage.
Support:
DROP INDEX ...
REBUILD INDEX ...
MAINTAIN DATABASE;
MAINTAIN TABLE table_name;
MAINTAIN INDEX index_name;Maintenance covers:
- dead row/version cleanup
- MVCC garbage eligibility
- UNDO cleanup/compaction
- B+Tree tombstone cleanup
- page reclamation
- freelist reuse
- full-text posting cleanup
- HNSW tombstone diagnostics/rebuild policy
- WAL retention respecting PITR
- statistics refresh
Maintenance is bounded by:
- CPU
- memory
- I/O
- concurrency
- admission control
At most bounded background/coordinated work may execute; no independent unbounded maintenance goroutine model.
NextSQL uses ACID transactions with undo-oriented MVCC.
Concept:
Current Row
│
▼
Undo Record
│
▼
Previous Version
│
▼
Older Version
Support:
- transaction IDs
- snapshots
- MVCC version chains
- locks
- rollback
- deadlock detection
READ COMMITTEDSNAPSHOTSERIALIZABLE
Serializable may only be advertised while anomaly tests prove the implemented guarantee.
Readers must not see uncommitted writes.
Rollback must restore prior state.
Write-ahead logging is mandatory.
Commit invariant:
WAL representing a durable committed modification must reach the configured durability boundary before COMMIT is acknowledged.
Commit path:
transaction
↓
WAL
↓
group commit
↓
fsync / quorum durability
↓
COMMIT acknowledgement
WAL includes:
- LSNs
- authenticated checksums
- encryption
- segments
- rotation
- group commit
- checkpoints
- archival hooks
- redo recovery
NextSQL must survive covered failures including:
- process crash
- SIGKILL
- machine restart
- checkpoint interruption
- partial WAL tail
- partial data write
- crash during B+Tree changes
- crash during index lifecycle
- crash during backup/maintenance metadata operations
After recovery:
- committed state remains;
- uncommitted state does not become committed;
- storage/index invariants remain valid.
Encryption is mandatory by default in production mode.
Persistent user data must not be stored in readable plaintext under normal production configuration.
Protect:
- table pages
- B+Tree structures
- secondary indexes
- JSON
- JSON indexes
- vectors
- ANN structures
- full-text indexes
- UNDO
- WAL
- temp files
- query spills
- snapshots
- backups
- archived WAL
- partition metadata/data
- workflow/task durable state
- CDC durable state where persisted
Security property:
Stolen database files, disks, snapshots, WAL archives, backups, vector files, or full-text structures remain unreadable without separately authorized cryptographic key material.
NextSQL may have a native encryption architecture, but must not invent proprietary cryptographic primitives.
Do not invent:
- cipher
- hash
- MAC
- KDF
- AEAD
Use reviewed algorithms and established libraries.
Initial authenticated encryption:
AES-256-GCM
Cipher suites and persistent envelopes are versioned.
Use envelope encryption.
External / Client Root Authority
│
▼
Root Unlock Key
│
▼
Key Encryption Key
│
▼
Database Master Key
│
┌─────────┼──────────┬──────────┬───────────┐
▼ ▼ ▼ ▼ ▼
Page DEK WAL DEK Backup DEK Vector DEK FTS/Other DEKs
Additional separated domains can include:
- UNDO
- temp/spill
- replication
- task/workflow metadata
Do not use one permanent key for all purposes.
Support:
REQUIRE CLIENT KEY
The critical root/unlock key need not live permanently on the server.
Conceptual flow:
Application
│
KeyProvider
│
NextSQL Driver
│
TLS
│
nextsqld
│
Authenticated unlock exchange
│
temporary cryptographic context
│
query execution
Persistent host files may contain:
- ciphertext
- wrapped DEKs
- key IDs
- crypto metadata
but not the raw external root key.
Never put encryption keys in connection URLs.
Support online key rotation.
old key version
→ generate new version
→ new writes use new key
→ bounded/background re-encryption
→ retire old key
Encrypted objects identify key version.
Support:
- session termination
- credential revocation
- key-version revocation
- wrapped-key rotation
- audit
- optional high-privilege crypto-shredding
Crypto-shredding warning:
NO KEY = NO RECOVERY
Provide an optional stronger client-encryption mode.
Future/advanced syntax:
CREATE TABLE customers (
id UUID PRIMARY KEY,
email STRING ENCRYPTED CLIENT,
phone STRING ENCRYPTED CLIENT,
profile JSON ENCRYPTED CLIENT
);For strongly client-encrypted fields:
- plaintext remains client-side;
- official drivers handle encryption/decryption;
- server-side arbitrary SQL over plaintext is unavailable unless explicitly supported by a documented leakage model;
- searchable encryption, if introduced, must document leakage.
Do not claim a live unlocked server can never expose plaintext from memory.
Production remote connections require secure transport.
Core:
TLS 1.3
Final Security 2.0 target includes:
- server-side mTLS
- client certificate validation
- certificate-to-service identity mapping
- rotation
- revocation
- audit identity source
Signed short-lived credentials/tokens support:
- expiration
- audience/database scope
- role scope
- realm/database scope
- signing-key rotation
- revocation
- audit
OIDC integration may provide:
- external login
- identity mapping
- group/role mapping
External identity must never bypass NextSQL RBAC.
Password hash records are versioned.
Migration to stronger algorithms such as Argon2id may occur while preserving compatibility and safe rehash behavior.
Support tamper-evident/signed audit mechanisms where production-gated.
Support:
- users
- roles
- grants
- revocation
Permission scopes include:
- cluster
- database
- schema
- table
- column
- function/workflow
- backup
- replication
- maintenance
- CDC
- administration
Least privilege is the default.
Realm and database isolation are enforced server-side.
Cross-realm/database leakage tolerance:
0 known leakage
Physical partitioning never replaces authorization.
The final managed-service architecture supports multiple subscription realms
and multiple databases per realm behind one bounded nextsqld deployment or
HA cluster. A realm is an account/authentication/resource boundary. Row-level
shared tenancy and SET TENANT are not part of the target architecture.
deployment
└── realm (principals, roles, plan, quotas, root/KMS boundary)
└── database (catalog, transactions, WAL, recovery, keys, backup, tasks)
Connections bind immutably to one realm and database. Routing never grants
access: authentication is realm-scoped and CONNECT is database-scoped. Each
database has independent identity, master/domain DEKs, WAL/recovery, backup,
CDC/task state, and resource accounting. Cross-realm and cross-database leakage
tolerance is zero.
Shared hosting is bounded by deployment, realm, database, and user/resource- group ceilings. Dedicated-instance and dedicated-cluster tiers remain valid stronger-isolation deployments. Billing/control-plane availability is never in the SQL correctness or commit path. Multi-database support does not imply cross-database transactions, joins, multi-primary writes, or automatic sharding.
Audit at minimum:
- authentication success/failure
- role changes
- permission changes
- DDL
- backup/restore
- key operations
- cluster membership
- security settings
- workflow lifecycle/run
- schedule lifecycle
- task control
- CDC subscription/security actions
- maintenance
- high-risk Manager/Studio actions
Never log:
- passwords
- encryption keys
- tokens
- private keys
- secrets
JSON is stored in compact binary form, not merely raw UTF-8 text.
Support:
- typed scalars
- objects
- arrays
- path traversal
- partial decoding
- indexed JSON paths
- mutation functions
- containment
- type inspection
JSON participates fully in:
- transactions
- WAL
- recovery
- backup
- encryption
- replication
- partitioning
- hybrid optimization
Core FTS includes:
- tokenizer
- normalization
- inverted index
- posting lists
- term/document frequency
- positions
- BM25-style scoring
- phrase search
SQL:
SELECT *
FROM articles
SEARCH body FOR 'database performance'
LIMIT 20;FTS is:
- transactional
- WAL-durable
- recoverable
- encrypted
- replicated
Final extended search target includes:
- stemming
- stop-word dictionaries
- versioned language analyzers
- synonyms
- prefix search
- fuzzy matching
- typo tolerance
- highlight/snippet generation
- multi-field search
- field weighting
- faceting/aggregation where architecturally appropriate
- analyzer/index options in DDL
Query expansion must have CPU/memory limits.
Analyzer metadata must participate in:
- transaction/WAL
- recovery
- replication
- encryption
- backup/restore
Analyzer behavior across replicas must remain deterministic.
Vectors are first-class values.
Core type:
VECTOR<F32,N>
Distances:
- COSINE
- L2
- INNER_PRODUCT
First-class value operations include dimension inspection, norm/normalize, element-wise add/subtract, scalar multiplication, dot product, cosine distance, and L1/Manhattan distance. All operations are dimension-strict, finite-only, and bounded by the vector dimension limit.
Core search:
- exact flat search
- HNSW
Example:
CREATE VECTOR INDEX ix_embedding
ON documents(embedding)
USING HNSW;Large vectors are stored outside ordinary row pages by reference to avoid page bloat.
Final advanced vector target includes production-gated support for appropriate subsets of:
VECTOR<F16,N>VECTOR<I8,N>BITVECTOR<N>- quantized HNSW
- IVF
- IVF-PQ
- quantization
- sparse retrieval
- dense+sparse+BM25 fusion
Every new vector representation has versioned encoding.
Every ANN structure is:
- encrypted
- crash recoverable
- transaction-aware
- delete-aware
- rebuildable
- Raft compatible
Every ANN configuration must report:
- recall@10
- recall@100
- p50/p95/p99
- QPS
- RAM
- index size
- build time
- database size
Never silently lower recall to improve latency.
Portable Go remains the correctness baseline.
SIMD/unsafe/architecture-specific code is introduced only after profiling, tests, fuzzing, and measured improvement.
Native geospatial support has two families that sit side by side:
Fixed WGS84 shapes (docs/geo.md) — POINT/LOCATION, BOX,
LINESTRING, POLYGON with WKT coercion, coordinate validation, LON/
LAT, DISTANCE/DISTANCE_SPHEROID, DWITHIN/WITHIN/COVERS,
line-length, pairwise INTERSECTS/DISJOINT, polygon area/perimeter,
centroid/envelope, and geometry inspection.
General OGC types — GEOMETRY (planar) and
GEOGRAPHY (geodetic) with an explicit per-column SRID + subtype
declaration (GEOMETRY(Point, 3857)), the OGC Simple Features common
subset of ST_* functions, EWKB/EWKT/WKB/GeoJSON serialization, and a
closed SRID set {0, 4326, 3857} (ST_Transform covers 4326 ↔ 3857
with no external PROJ dependency). These cover MultiPoint/
MultiLineString/MultiPolygon/GeometryCollection. The two families
coerce into one another (POINT ⇄ GEOMETRY(Point)); the fixed shapes
are unchanged. This is a native subsystem targeting the common OGC subset,
not a PostGIS clone — 3D/M coordinates, curve geometries, arbitrary datum
transforms, and topology remain out of scope.
- spatial indexes over both families
Residual predicates remain exact.
Geo participates in:
- WAL/recovery
- MVCC
- encryption
- optimizer costing
- replication
- schema lifecycle
The optimizer understands combinations of:
relational predicates
+ JSON paths
+ full-text
+ vectors
+ geospatial predicates
Possible plan:
100M rows
│
structured indexes
▼
3M
│
additional filters
▼
250K
│
ANN candidate generation
▼
1K
│
BM25/vector rerank
▼
20
But the optimizer must also consider:
ANN first
→ structured filtering
when cheaper.
Operator order is cost-based, not hard-coded.
NextSQL uses a coherent programmable automation model instead of unrelated stored procedure/event subsystems.
WORKFLOW
├── manual invocation
├── trigger invocation
└── scheduled invocation
↓
TASK
Native workflow surface includes:
CREATE WORKFLOW ...
ALTER WORKFLOW ...
DROP WORKFLOW ...
RUN WORKFLOW ...Workflows support:
- typed parameters
- bounded multi-statement bodies
- explicit transaction semantics
- database-isolation semantics
- RBAC
- audit
- dependency tracking
- recursion/resource limits
A workflow can run manually without a trigger.
Native triggers execute workflows.
Events include:
BEFORE INSERT
AFTER INSERT
BEFORE UPDATE
AFTER UPDATE
BEFORE DELETE
AFTER DELETE
Example conceptual syntax:
CREATE TRIGGER ...
AFTER INSERT
ON orders
RUN WORKFLOW process_order(...);Trigger/workflow execution must enforce:
- trigger recursion depth
- workflow recursion depth
- cycle detection
- statement count
- time
- memory
- generated task count
- deterministic replication behavior
Native scheduler includes:
CREATE SCHEDULE ...
ALTER SCHEDULE ...
DROP SCHEDULE ...Initial scheduling primitives:
EVERY duration
AT timestamp
Schedules are stored durably.
In HA mode:
- one authoritative Raft-aware dispatcher exists;
- leader failover behavior is documented;
- duplicate execution is prevented within the documented guarantee;
- clock-skew behavior is explicit.
Cron-style syntax may be added only after the core scheduler is proven.
Scheduled/asynchronous execution is represented by durable tasks.
Task states:
PENDING
RUNNING
SUCCEEDED
FAILED
CANCELLED
RETRYING
Task metadata includes:
- task ID
- workflow/source
- trigger source
- database identity
- attempts
- error details
- timeout
- retry count
- retry backoff
- idempotency key
- final/dead-letter semantics
- concurrency policy
- retention
Expose through a canonical machine-readable surface such as:
SELECT * FROM system.tasks;and/or SHOW TASKS.
Tasks are cancellable and run in bounded worker pools.
NextSQL provides native committed change streaming sourced from WAL.
CDC emits committed transactions only.
Events include:
- INSERT
- UPDATE
- DELETE
- database/table identity
- primary-key identity
- transaction/commit identity where safe
- LSN/resume token
- relevant before/after metadata according to configured mode
CDC requirements:
- stable ordering semantics
- resume/restart behavior
- durable resume positions where applicable
- backpressure
- bounded buffers
- lag metrics
- cancellation
- RBAC
- stable realm/database filtering
- secure transport
- failover semantics
A slow consumer must never create unbounded server memory growth.
CDC must not expose uncommitted data.
NextSQL supports native physical table partitioning.
Partitioning modes target:
RANGE
HASH
LIST
for the production-gated subset.
Partitioning includes:
- catalog metadata
- partition routing
- optimizer pruning
- partition-aware statistics
- indexes
- transaction correctness
- WAL/recovery
- Raft replication
- backup/restore
- PITR
- maintenance
- schema lifecycle
Optimizer must expose pruning in EXPLAIN.
After physical partitioning exists, enable:
- partition-wise aggregation
- partition-wise joins
Older TENANT partition descriptors remain decodable for recovery and explicit offline migration only. New SQL cannot create or extend them. Subscription isolation uses immutable realm/database connection identity rather than row partitioning.
Physical partitioning is never an authorization mechanism. No cross-realm or cross-database result may be returned.
Automatic distributed sharding remains deferred beyond the core roadmap.
Writes remain single-leader.
NextSQL supports explicit read consistency modes.
Required concepts:
Strong reads use:
- leader execution
- or a valid Raft read barrier/proven equivalent
and satisfy the documented consistency guarantee.
Optional follower-read modes may include:
- eventual/stale
- bounded staleness
MAX STALENESSsemantics if adopted
Routing considers:
- read-only eligibility
- follower health
- replica lag
- requested consistency
- transaction context
- read-after-write expectations
Writes never silently route to stale followers.
Stale reads must never be labeled strong.
Drivers expose supported routing/consistency metadata.
NextSQL uses a versioned native wire protocol.
Support:
- protocol negotiation/versioning
- TLS 1.3
- authentication
- typed parameters
- prepared statements
- streaming results
- backpressure
- query cancellation
- packet-size limits
- SQL-size limits
- result limits
- runtime/memory/worker limits
- capability negotiation
- bounded SELECT result-cache semantics with explicit invalidation
- durable database-user-scoped mutation idempotency keys
Official drivers use the same protocol.
Network input is untrusted.
Never allocate directly from unchecked attacker-controlled lengths.
Official drivers target:
- Go
- Node.js
- TypeScript
- Bun
- PHP
- Python
- Ruby
Drivers must support applicable server features such as:
- TLS
- mTLS when enabled
- secure credential handling
KeyProvider- typed parameters
- prepared statements
- streaming
- cancellation
- transactions
- consistency profiles
- realm/database context when negotiated
- workflow/task APIs
- CDC subscriptions
- client-side encrypted fields when production-gated
- capability/version negotiation
- idempotent mutation execution where supported
Keys never belong in connection URLs.
NextSQL HA uses proven Raft consensus.
Minimum recommended voting topology:
3 voting nodes
Support:
- replication
- leader election
- failover
- replica repair
- rolling maintenance
- quorum-loss handling
- synchronous quorum commit
- leader health
- cluster membership
If a safe leader cannot be identified:
reject writes
rather than risk split brain.
No custom consensus algorithm.
Do not advertise guaranteed 100% uptime.
HA design objective:
>= 99.999% availability SLO
for properly configured supported clusters.
Failover engineering targets:
leader election < 3 s
service recovery < 5 s
For acknowledged synchronous quorum commits under supported failures:
RPO = 0
No acknowledged commit may be reported successful before the selected durability policy is satisfied.
CLI surface includes:
nextsql backup
nextsql restore
nextsql export
nextsql importBackups remain encrypted.
A live server also exposes backup over SQL for NextSQL Admin's Operations mode: BACKUP DATABASE and VERIFY BACKUP 'name' (both BACKUP-privilege / cluster
ADMIN gated) operate on the server's configured backup_dir via a
hot-backup path that reuses the running engine (backup.CreateFromEngine) —
the same checkpoint + fuzzy-copy + WAL-replay model, never a second engine
open. Restore and PITR stay offline-CLI-only: a running server cannot restore
into itself.
Required backup flow:
backup
↓
manifest
↓
integrity verification
↓
storage
↓
verification
↓
periodic restore test
A successful upload is not proof of a valid backup.
PITR uses:
base backup
+ archived WAL
= point-in-time recovery
Restore targets include:
- timestamp according to documented semantics
- LSN
Backup/restore must understand all persistent production-gated structures, including future workflows, tasks, partitions, security metadata, and index formats.
Known silent corruption tolerance:
0
Use:
- authenticated page integrity
- WAL authentication/checksums
- backup verification
- version/format validation
- LSN validation
- index consistency checks
- structural invariants
Corruption handling:
detect
→ isolate
→ fail safely
→ recover/repair where supported
Never return a known-corrupted record as valid.
Every query/task/maintenance operation must be resource-bounded.
Controls include:
- admission control
- bounded queueing
- throttling
- cancellation
- memory budgets
- CPU/worker budgets
- I/O budgets
- temp/spill budgets
- execution timeout
- result-size limits
- recursion limits
- query-complexity limits
Final workload-governance target includes:
- connection/session limits
- graceful shutdown
- connection draining
- query cancellation
- session termination
- resource groups/classes
- concurrency quotas
- workload prioritization
- operational diagnostics
Overload should cause:
queue
throttle
reject
cancel
spill
not:
unbounded goroutines
unbounded allocations
OOM
NextSQL provides a machine-queryable virtual system schema as the canonical introspection interface.
It should expose authorized information about applicable objects such as:
- server/version
- capabilities
- databases
- schemas
- tables
- columns
- constraints
- indexes
- index status
- statistics
- sessions
- active queries
- transactions
- locks
- users
- roles
- grants
- realms/databases
- replication
- cluster nodes
- leader/followers
- replica lag
- backups
- WAL/PITR
- maintenance
- workflows
- triggers
- schedules
- tasks
- CDC subscriptions/status
- partitions
- full-text indexes/analyzers
- vector indexes
- security configuration
- audit status
- operational metrics
Important final capability object:
system.capabilities
It is authoritative for:
- supported features
- experimental features
- deprecated features
- unsupported features
- feature/version metadata
NextSQL Admin (all modes) and drivers must negotiate against server capabilities rather than assuming feature availability.
All system views obey RBAC and realm/database boundaries.
NextSQL exposes metrics and diagnostics for:
- queries
- latency
- QPS/TPS
- storage
- memory
- CPU
- WAL
- checkpoints
- cache/buffer behavior
- spills
- worker utilization
- replication
- failover
- replica lag
- backup
- maintenance
- index rebuild
- workflows
- tasks
- CDC
- partition pruning
- vector index behavior
- FTS behavior
- admission control
- security/authentication events
Diagnostics must not leak secrets.
Support:
EXPLAIN ...
EXPLAIN ANALYZE ...Expose where applicable:
- operator
- estimated rows
- actual rows
- execution time
- CPU
- memory
- disk reads
- cache hits
- spill
- workers
- index
- partition pruning
- FTS candidate generation
- vector candidates
- hybrid reranking
- join strategy
Plans must be suitable for both CLI output and Studio visualization.
NextSQL provides native migration/version workflow.
Migration tooling must understand native NextSQL DDL rather than translate through another database dialect.
Schema lifecycle includes safe handling of:
- tables
- indexes
- constraints
- workflows
- schedules
- partitions
- future persistent database objects
Migration validation must use parser/binder/catalog truth.
NextSQL Admin is the official single application for installing, operating, and
developing against NextSQL. It ships as one binary (nextsql-admin) with one frontend
shell and three modes:
- Setup mode (this section) — first-run install/upgrade/repair/uninstall lifecycle.
- Operations mode (§47) — day-to-day server/cluster/security/backup administration.
- Studio mode (§48–§55) — the database development IDE. Intelligence / RAG is not in the product (§56).
The three modes share one process, one visual/accessibility baseline, and one product
identity; they differ in what they connect to and what credentials they require, not in
product identity. Setup mode runs before a database exists and needs no login (a
single-operator local trust boundary); Operations and Studio modes both connect to a
running nextsqld using real NSQL credentials. A cross-mode contract applies to all
three — see §73.
The official installer provides a professional installation lifecycle.
Functions include:
- platform prerequisites
- install
- initialize
- data directory selection
- key/bootstrap configuration
- service registration
- start/stop validation
- repair
- upgrade
- rollback strategy where supported
- uninstall
- diagnostics
UX principles:
- clear and user-friendly
- safe defaults
- explicit destructive-action warnings
- accessible
- keyboard-friendly
- reliable progress/error reporting
- no plaintext secret storage
- production/readiness checks
Installer must not imply unsupported OS/platform combinations are production-ready.
Native Windows is not a product platform. On a Windows machine, NextSQL's server, CLI and Admin run inside WSL 2 from the Linux packages; the official client drivers remain usable from native Windows applications.
Operations mode is NextSQL Admin's official operational administration surface, complementing Setup mode (§46) and Studio mode (§48).
Primary responsibilities:
- server overview
- health
- configuration
- databases
- storage
- backups
- restore/PITR
- maintenance
- users
- roles
- grants
- encryption/key status
- audit
- HA/replication
- node/leader status
- replica lag
- workload controls
- logs/metrics/diagnostics
- upgrades
- workflow/task operational status
- CDC operational status
- partition status
Operations mode uses:
- official NSQL/API interfaces
systemschema- server capability negotiation
Operations mode must never:
- read raw database pages directly
- read WAL files as a shortcut
- require the raw root unlock key
- bypass server RBAC
- display fake cluster/security state derived only from local UI assumptions
Server truth is authoritative.
Studio mode is NextSQL Admin's official database development IDE experience, complementing Setup mode (§46) and Operations mode (§47).
Target users:
- developers
- DBAs
- data engineers
- backend engineers
- system architects
Core rule:
Do not build a generic SQL client with a NextSQL logo.
Studio understands native:
- NextSQL SQL
- JSON
- FTS
- vectors
- hybrid search
- geospatial
- realms/databases
- workflows/tasks
- CDC
- partitions
- migrations
- query plans
- capabilities
Studio communicates only through official supported interfaces.
It never directly reads pages/WAL/catalog files.
Studio includes:
- professional IDE layout
- connection explorer
- editor workspace
- results panel
- plan panel
- messages panel
- statistics panel
- inspector
- command palette
- recent connections/projects
- light/dark/system themes
- layout persistence without secrets
- keyboard-first navigation
- accessibility baseline
- high-DPI support
- unsaved editor crash recovery
Connection profiles support:
- name
- host
- port
- database
- user
- TLS
- CA
- client certificate where supported
- realm/database profile
- read-consistency profile
- secure saved credentials
- environment profile
Environment categories can include:
development
test
staging
production
Production connections must be visually obvious.
Production safety features include:
- optional read-only default
- destructive DDL warnings
- parsed-AST warning for UPDATE/DELETE without WHERE
- capability-aware validation
Never store raw credentials in plaintext configuration files.
Studio provides lazy-loaded exploration of:
- databases
- schemas
- tables
- columns
- primary keys
- foreign keys
- constraints
- indexes
- statistics
- DDL
- dependencies
- workflows
- triggers
- schedules
- tasks
- CDC
- partitions
Design tools include:
- table designer
- index designer
- native DDL preview
- ER diagram from actual foreign-key metadata
- global object search
All metadata comes from authorized server introspection.
SQL editor includes:
- NextSQL-native syntax highlighting
- line numbers
- bracket matching
- indentation
- tabs
- execute statement
- execute selection
- execute script
- cancel query
- formatting
- find/replace
- live-catalog IntelliSense
- JSON-path completion
- vector-aware completion
- inline parser/binder diagnostics
- safe identifier suggestions
- query history with privacy controls
- saved queries
- folders/tags
- Git-friendly workspace artifacts
Editor diagnostics should use actual NextSQL parser/binder semantics where possible.
Results must support:
- streaming
- virtualization/paging
- typed rendering
- NULL distinction
- JSON inspection
- vector-aware display
- geospatial display where useful
- copy
- export
- bounded memory
Editable data grids must preserve:
- transaction safety
- primary-key identity
- concurrency behavior
- server validation
- RBAC
Studio must never simulate a successful edit when the server rejects it.
Studio visualizes EXPLAIN / EXPLAIN ANALYZE.
Show:
- plan tree
- estimates vs actuals
- timings
- CPU
- memory
- disk
- cache
- spills
- workers
- indexes
- partition pruning
- FTS candidates/ranking
- vector candidates
- hybrid reranking
The plan UI must reflect server output rather than reconstruct an imagined plan client-side.
Studio includes dedicated experiences for:
- structured viewer
- path inspection
- JSON index awareness
- JSON function tooling
- analyzer/index inspection
- query testing
- ranking information
- highlights/snippets where supported
- vector dimension/index inspection
- HNSW/ANN configuration
- recall/latency test support where appropriate
- index lifecycle visibility
- structured + FTS + vector query construction/testing
- candidate/ranking inspection
- point/shape inspection
- distance/filter testing
- index status
- workflow editor
- trigger/schedule inspection
- run history
- task state/error/attempt view
- cancellation controls as authorized
- subscription configuration/inspection
- resume token/lag visibility
- event viewer for authorized streams
These were the former P30 phase (NextSQL Intelligence + built-in RAG,
including a Studio assistant, RAG Playground, and research toward
CREATE RETRIEVER). They are removed from the product. They are not
deferred, not planned, and not a later-version commitment.
Do not implement:
- a Studio AI assistant, chat panel, or “Ask NextSQL Intelligence” action
- a RAG Playground product surface
CREATE RETRIEVER/RETRIEVEas an Intelligence/RAG object- an Intelligence knowledge base, provider layer, or tool layer
- Intelligence-specific encryption domains, RBAC privileges, or capabilities
Studio is a native database development IDE without an AI assistant. Full-text search, vector search, and hybrid SQL remain ordinary engine features; they are not an Intelligence product.
LLM-inside-the-optimizer and LLM-required-for-correctness remain rejected (see §77).
nextsql-bench measures at minimum:
- point SELECT
- range SELECT
- INSERT
- UPDATE
- DELETE
- transactions
- joins
- aggregations
- JSON
- full-text
- vector
- hybrid
- partition-aware workloads
- follower-read scaling where applicable
Report:
- QPS
- TPS
- p50
- p95
- p99
- p99.9
- CPU
- RAM
- allocations
- disk
- WAL
- database size
- index size
- encryption overhead
- hardware
- OS
- filesystem
- cache condition
- concurrency
Vector benchmarks also report recall.
Performance numbers are engineering targets and measured results, not universal guarantees.
p50 < 0.5 ms
p95 < 1 ms
p99 < 3 ms
p50 < 1 ms
p95 < 3 ms
p99 < 5 ms
Target:
p50 < 2 ms
p95 < 5 ms
p99 < 10 ms
subject to storage durability latency.
Target classes:
25K rows < 1 s
100K rows < 1 s for suitable optimized processing
1M rows < 1 s for suitable scans/aggregations
10M rows < 5 s for suitable optimized aggregation
100M rows < 30–60 s for appropriate analytical workloads
Top-10
p50 < 10 ms
p95 < 25 ms
p99 < 50 ms
must be reported with:
- recall@10
- recall@100
- QPS
- RAM
- index size
Initial target on appropriate indexed data/hardware:
p50 < 50 ms
p95 < 100 ms
p99 < 250 ms
Continuously fuzz applicable untrusted inputs:
- SQL parser
- wire protocol
- authentication protocol
- page decoder
- WAL decoder
- backup parser
- export/import parser
- JSON parser
- vector metadata
- full-text structures
- replication command decoder
- workflow syntax/runtime edges
- CDC/resume metadata
- partition metadata
- new persistent format decoders
Malformed input must produce controlled errors.
Every feature must receive applicable coverage across:
unit
integration
transaction
concurrency
restart
crash injection
WAL/recovery
Raft/failover
backup/restore
PITR
RBAC
realm/database isolation
prepared statements
wire protocol
official drivers
race detector
fuzz/property tests
resource-limit tests
benchmarks
documentation
A phase is not complete because code merely exists.
It is complete only after implementation, tests, docs, and its exit gate are green.
Applies to NextSQL Admin (Setup, Operations, and Studio modes).
All user-facing products should provide:
- professional consistent design
- accessibility baseline
- clear errors
- actionable diagnostics
- keyboard-friendly interaction
- confirmation for destructive operations
- production environment visibility
- secret-safe storage
- server-authoritative status
- capability negotiation
- no fake success states
- installable as a PWA where the product is a web UI (NextSQL Admin): a manifest, a
service worker, and offline app-shell loading — but a service worker must never cache
live session/auth/query traffic (
/api/*-shaped surfaces) or a one-time auth token URL; only the static shell is ever cached - no silent privilege escalation
Never advertise:
100% secure
unhackable
guaranteed zero downtime
fastest database in the world
impossible to lose data
Use:
- engineering target
- measured benchmark
- design objective
- SLO
- supported failure model
- documented threat model
instead.
Every new persistent structure must document:
- format version
- encryption domain
- key version
- integrity/authentication
- backup behavior
- restore behavior
- PITR behavior where relevant
- rotation behavior
- replication behavior
- upgrade/migration behavior
- corruption behavior
No unversioned persistent format is allowed.
The complete NextSQL product succeeds when the production-gated roadmap provides:
Native SQL relational engine
Native binary JSON
Native full-text search
Native vector search
Native geospatial
Unified hybrid optimizer
ACID MVCC transactions
WAL + crash recovery
Mandatory production encryption
Client-held key support
RBAC + realm/database isolation
Backup / restore / PITR
Raft HA
Follower read scaling
Native table partitioning
Schema lifecycle + maintenance
Modern SQL completeness
WORKFLOW / TRIGGER / SCHEDULE / TASK
CDC / change streams
Canonical system introspection
Operational workload governance
Official drivers
NextSQL Admin (Setup / Operations / Studio modes)
with long-term quality objectives:
Persistent plaintext:
0 by default in production mode
Known silent corruption tolerance:
0
Known critical unresolved production vulnerabilities:
0 at a production release gate
Cross-realm/database leakage tolerance:
0
Lost acknowledged synchronous quorum commits:
0 within supported failure assumptions
HA availability design SLO:
>= 99.999%
Leader election target:
< 3 seconds
HA service recovery target:
< 5 seconds
Cached point lookup:
p50 < 0.5 ms target
Indexed query:
p95 < 3 ms target
25K processed rows:
< 1 second target
1M suitable optimized aggregation:
< 1 second target
10M suitable optimized aggregation:
< 5 seconds target
100M analytical workload:
< 30–60 seconds target
1M-vector HNSW Top-10:
p95 < 25 ms target with recall reported
Encryption:
mandatory in production mode
The following are not required to consider the product family complete:
- NextSQL Intelligence
- RAG Playground
CREATE RETRIEVER/ Intelligence-layer retrieval objects
- automatic distributed sharding
- autonomous cross-node shard placement/rebalancing
- LLM inside the deterministic query optimizer
- LLM required for database correctness
- LLM required for transactions/WAL/recovery/security
- mandatory cloud account for local NextSQL
- hidden PostgreSQL/MySQL compatibility engine
The preferred distributed model remains:
single Raft write leader
+ synchronous quorum durability
+ followers
+ optional follower reads
+ local native partitioning
+ later explicit shard-placement research if justified
Use the project files as follows:
PROJECT.md
→ What the finished NextSQL product is expected to be.
TODO.md
→ What is implemented, open, blocked, deferred, or production-gated.
SKILLS.md
→ How an AI coding agent must behave while working on NextSQL.
docs/*
→ Detailed technical specifications and measured implementation truth.
If PROJECT.md and TODO.md differ on implementation status:
TODO.md wins for status.
PROJECT.md wins for intended end-state.
If a feature is only designed but not implemented/tested:
- keep it unchecked in
TODO.md; - do not claim it is shipped;
- it may still appear in
PROJECT.mdas part of the intended final product.
NextSQL must be fast without weakening correctness, encrypted without pretending custom cryptography is safer, highly available without risking split brain, multimodel without becoming several loosely coupled databases, and professional without allowing tooling to bypass the engine's security model.
The finished system is not merely a database executable.
It is a complete native database platform:
Engine
+ Protocol
+ Drivers
+ CLI
+ Bench
+ Automation
+ CDC
+ Partitioning
+ HA / Read Scaling
+ Security
+ Operations
+ Installer
+ Manager
+ Studio
all built around one authoritative NextSQL engine.