veristamp’s cover photo
veristamp

veristamp

IT Services and IT Consulting

kolkata, west bengal 3 followers

Independent lab building sovereign AI infrastructure.

About us

Veristamp builds sovereign AI infrastructure: search, agents, identity, and proof you host yourself. No mandatory cloud.

Website
veristamp.in
Industry
IT Services and IT Consulting
Company size
1 employee
Headquarters
kolkata, west bengal
Type
Self-Owned

Locations

Employees at veristamp

Updates

  • veristamp reposted this

    🦖 I taught a 32M text classifier to play Chrome T-Rex. No game AI. No PPO. It just reads emit a single decision. Every frame the board becomes one English sentence: “Dino y=93, vy=0, spd=6, ground. Obs1: CACTUS_SMALL, dist=120px, time=20f, win=[1,140]px. Action:” GLiClass scores RUN / JUMP / DUCK. Highest number wins. ~11ms a move. That’s the engine. Snake was 20k SFT and it worked. T-Rex needed more because the timing window is unforgiving. So this run was SFT → DAgger → DPO. No PPO. didn’t need it. reverse-engineered Chromium’s official `runner.js`, built a verified oracle, and cloned the expert. The journey: ❌ Zero-shot: dead in 10–20 steps. Confident. Always ducks. 📊 v1 (SFT + DAgger): 2000/2000 steps, ~30 obstacles. Then DPO on easy pairs. Loss 0.037 → 0.0017. Policy went RUN 100%. 87 steps. 0 obstacles. 20/20 deaths. It learned “don’t jump.” ✅ v2 (balanced boundary pairs + a closed-loop gate): 20/20 games. 6000 steps. 0 deaths. ~97 obstacles cleared. Same failure mode as Snake when the objective is sloppy: it takes the safe shortcut and dies. Same tech stack.. 1. Free Kaggle single T4. Same 32M GliClass edge model from Knowledgator. took 33 min to train e2e. 2. Real Chrome physics from original source. 3. A perfect expert beat PPO. Model stays local for now. Write-up and codes coming. #MachineLearning #AI #NLP #PyTorch #Kaggle #BuildInPublic #GliClass #KnowledgeGator

  • veristamp reposted this

    QQL 0.4.0 is live today across PyPI, npm, and crates.io. QQL is the typed SQL query layer for Qdrant. One declarative syntax across Python, Node, Rust, and in-process Edge runtimes. If it parses, it compiles into validated HTTP or gRPC. In 0.3.0, we bound parameters on raw text. In 0.4.0, parameter binding moves directly to the AST. We boxed parameter spans to drop `size_of::<Value>()` from 56 bytes down to 32 bytes using Rust's null-pointer niche optimization. AST traversal is 3% to 11% faster, and Python `pyqql.parse` hits 1.50M ops/sec via PyO3. Here is what changed in 0.4.0 along with bunch of other fixes: 1. AST Parameter Binding & Bulk Ingestion No string concatenation. You prepare `UPSERT INTO docs VALUES :rows` once, and `client.upsert_many()` streams chunks through a zero-clone splice path. Packed `Float32Array` with a single memory copy. 2. First-Class BATCH Blocks in One Roundtrip Sequential HTTP calls kill agent latency. With `BATCH { QUERY ...; QUERY ...; }`, multiple queries or mutations execute inside a single wire RPC call. Envelope headers like `WAIT true` and `PARAMS (consistency = 'majority')` propagate to REST, gRPC, and Edge automatically. 3. `qql convert` translates 25 Qdrant OpenAPI endpoints into canonical, re-parseable QQL. And `qql record` acts as a transparent reverse proxy: point your app at it, your traffic forwards untouched, and clean `.qql` replay scripts are logged in the background. 4. Zero-Snapshot Cluster Migration (`qql migrate`) Native binary snapshots break across minor versions, can't change shard topologies, and carry index fragmentation. `qql migrate` streams schema and points between clusters. It automatically discovers shard keys (`--shard-key-field`), suppresses `indexing_threshold` to 2GB during ingestion to prevent CPU starvation from HNSW rebuilds, verifies exact COUNT parity, and resumes from atomic checkpoints. 5. Closed Typed Response Model & DB-API 2.0 We eliminated raw JSON passthrough. All backend responses normalize into a closed `ExecData` enum. Python and Node drivers now return native PyO3 and N-API `ExecutionReport` and `ScoredPoint` classes with memory-bounded scroll cursors. Python gets a full DB-API 2.0 interface (`pyqql.connect()`), and CI gets `qql lint --fix` and `qql doctor`. Zero friction to test locally. The quickstart pipeline runs completely offline without Docker or a running server: pip install pyqql==0.4.0 npm install @veristamp/nqql@0.4.0 cargo install qql-cli --locked Agents shouldn't assemble nested JSON dictionaries. Queries should be typed, deterministic, and auditable. Release notes: https://lnkd.in/gMAQr8BK Deep dive: https://lnkd.in/gKqYuArA Docs: https://lnkd.in/dK2e3EJm #Qdrant #VectorSearch #RAG #AIEngineering #AIAgents #OpenSource #Rust #Python #Database #ZeroTrust

  • veristamp reposted this

    Most location search architectures are built around a hard boolean: `distance <= 1.5 km` AND `price <= €95` 1.49 km is in. 1.51 km is gone. A circle cannot prefer quality; it can only include or exclude. If a 4.9★ Superhost apartment sits 50 meters past the line for €90, a hard filter deletes it. I tested this on 8,317 real Berlin Airbnb listings across 10 traveler personas to see what that circle actually throws away: It drops an average of 37.1% (~114 listings per query) of high-quality stays sitting right in the adjacent 1.0–1.25× buffer band. If you swing to the other extreme pure vector search with no geo, you get the matching vibe, but 9.3 km away on the other side of the city. To solve this without gluing together PostGIS, an elastic cluster, and a separate reranker pipeline, we split the job between two modern tools: 1. Upstream Enrichment (@DuckDB Spatial, run once): • Evaluates `ST_Contains` against 139 Berlin LOR district polygons in memory. • Computes `h3_latlng_to_cell` (res 8) for client-side heatmap binning (since vector DBs don't do hex indexing). • Computes ellipsoidal geodesic distances (`ST_Distance_Spheroid`) as ground truth for evaluation. 2. Real-Time Ranking (Qdrant 1.19, single database round-trip): • Hybrid prefetch: Dense (1-bit BQ) + BM25 fused with DBSF so scores sit on the same 0–1 scale. • `FormulaQuery`: continuous `GaussDecay` on distance and `ExpDecay` on price. Proximity and budget cost points instead of acting as a guillotine. • Native Polygon Pushdown: a 387-vertex Alexanderplatz boundary evaluates directly inside Qdrant. No PostGIS hop. • Live Map Facets: `client.facet` aggregates district and room types during live map pans. The Benchmark Across 10 Traveler Personas: ❌ Pure Semantic (BQ): Spatially blind. 4,631 m avg distance | €70.7 off budget | 4.01 ★ | 2.3 ms ⚠️ Hard Circle Cliff: 1,013 m avg distance | €57.7 off budget | 44% Superhost | 37.1% ring loss | 2.5 ms ✅ Multi-Decay Formula: 1,945 m avg distance | €34.0 off budget | 72% Superhost | 0% ring loss | 8.1 ms 🗺️ Viewport + In-DB Facets: 1,068 m avg distance | €14.7 off budget | 94% Superhost | 6.8 ms The core trade-off: If the boundary is a legal city limit or delivery zone, use `geo_radius`. If the boundary is human intent ("near Alexanderplatz, around €95"), make distance and budget continuous scoring curves. On the map overlay: 🟢 Green dots = listings inside the hard circle. 🟠 Amber dots = high-quality stays in the perimeter ring that the hard circle killed, but continuous decay recovered. DuckDB does the spatial preparation. Qdrant handles the vector retrieval, hybrid fusion, continuous decay, and polygon filtering in under 9 ms. 💻 Repo: https://lnkd.in/gJR7beyy 🔗 Blog: https://lnkd.in/gFNQvrRP #Qdrant #DuckDB #VectorSearch #Geospatial #SearchEngine #Databases #OpenSource #DataEngineering #InformationRetrieval

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • veristamp reposted this

    India renamed the criminal law in 2023. Seventy years of Supreme Court case law did not get the memo. A lawyer now searches Section 482 BNSS. Anticipatory bail. No time limit. The governing case is still written under Section 438 CrPC. Same rule. Different number. Dense search treats them as strangers. I ran 22 labeled eval queries on 2,034 real Supreme Court criminal precedents (cleaned from 4,960 NyayaRAG judgments). Four ranking modes. Same collection in Qdrant. No extra reranker service. 1. Pure Dense (BGE-small BQ): 7/22 Top-1 (31.8%) 2.4 ms p50. Fast, and wrong on two out of three. On that bail question, dense put Siddharam Satlingappa Mhetre (2010) first. Sushila Aggarwal, the 2020 Constitution Bench, was sitting at number two. Right neighbourhood. Wrong lead case. 2. Hybrid (Dense + BM25): 10/22 Top-1 (45.5%) Catches exact section strings, but keyword overlap still cannot cross the statutory gap. 3. The Concordance Map: 16/22 Top-1 (72.7%) The actual jump is a 75-section map. Every old holding carries both old and new citations on the exact same record (438 and 482). A scoring formula inside Qdrant boosts the bridge. Section queries: 8/8. Queries written only in the new codes: 6/6. ⚠️ Then the twist: On plain-English questions with no section numbers, that same map over-helps. Natural language Top-1 drops 3/8 → 2/8. The statutory boost amplifies false positives. 4. Native Late Interaction (ColBERT): 19/22 Top-1 (86.4%) An in-database ColBERT MaxSim rescore pulls those paraphrases back to 5/8 without an external cross-encoder. Full stack: 19/22 Top-1 (86.4%). 21/22 in the top 3. 13.2 ms p50. One query still misses the top 3 (N8: custody disclosure, old Evidence Act 27). The source dataset simply does not have it. Ranking did not fail. I am not gonna dress that up. Dataset is NyayaRAG from Shubham Kumar Nigam. When the law gets renamed, the archive needs a map. Not a bigger model. Same rule. Dense search doesn’t know 438 = 482. Blog Link (full tables, formula weights, the miss): 🔗 https://lnkd.in/gdFunaBa Code, concordance map, and 22-query eval: 🔗 https://lnkd.in/gu4bvBMV #Qdrant #VectorSearch #RAG #LegalTech #AIAgents #ColBERT #BM25 #OpenSource #HybridSearch #IndianLaw

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image