Skip to content

Repository files navigation

AI Language — Emergent Communication Research

This project investigates whether independently trained AI agents can develop a communication protocol that is more compact, less ambiguous and more portable than human language for a defined task.

The first version is called ECP-0: Emergent Communication Protocol, experiment 0. ECP-0 does not use words or pre-assigned symbol meanings. A sender and receiver must reconstruct a meaning from an artificial world via four discrete symbols.

Research question

Can independent AI agents themselves develop a communication protocol that is efficient, generalizable, translatable and portable given the same task performance?

Within this research, a protocol only counts as a serious new form of communication when it:

  1. is not semantically pre-programmed by humans;
  2. can express unknown combinations;
  3. can be accessed by an independent translator;
  4. can be learned by new agents;
  5. is demonstrably efficient compared to fixed baselines;
  6. only uses the allowed and fully logged channel.

Current status

ECP-6 remains the latest confirmed result and successful scale replication of the fully efficient protocol. All five seeds reconstruct known and completely new factor combinations for 100%. The universal translator and the worst agent pair also achieve 100%. The predetermined classification is strong evidence.

The protocol uses no words or alphabet, but four meaning-free local symbols. For 16 × 16 × 8 × 8 = 16,384 uniform meanings it uses exactly 4+4+3+3=14 bits: the information-theoretic lower bound. For this defined task, this is on average 23.9 times more compact than the Dutch text template.

See docs/results-ecp6.md for the conclusion, docs/protocol-specification-ecp6.md for the wire format and evidence/ecp6/report.md for the compact confirmatory evidence.

ECP-7 development tested how much of that result depends on the explicit one-factor-per-slot architecture. All 29 sealed batches are valid development results, but none passes the full gate. Batch 15 established the strong position-aware base: at 30,000 optimization steps it reaches 83.46% train exactness, 82.59% validation and 83.37% translator validation while using 12,585–13,200 hard messages per sender. Validation and translator thresholds pass together, but the registered train and injectivity thresholds do not. Batch 16 added late worst-factor pressure and regressed to 76.46% validation. Batch 17 replayed globally mined training collisions and modestly improved code use, but still regressed to 77.09% validation. Batch 18 reduced replay to task-loss scale, recovering 80.71% validation and a new-best 84.06% translator score, but train exactness and injectivity still failed. Batch 19 bounded replay to a pulse and reached 82.04% validation plus a new-best 80.57% worst-link validation, but again failed train exactness and injectivity. The ECP-6 positive controls remain perfect. Batch 20 replayed population-hard training meanings and established new-best 83.45% mean and 82.13% worst-link validation, but again failed train exactness and injectivity. Batch 21 restricted replay to errors shared by all 16 links and improved validation again to 83.63% mean and 82.62% worst-link, but its target error pool grew and train exactness still failed. The ECP-7 confirmatory test remains sealed. Batch 22 routed shared-error replay only into senders; this reduced shared errors but increased cross-link fragmentation and regressed validation to 83.51%. Batch 23 restored joint replay after the sender-only warmup. Validation remained 83.50%; total fragmented errors fell, but the shared-error pool immediately regrew and injectivity still failed. Batch 24 made the second phase receiver-only. It established new-best means of 84.69% train, 84.48% validation and 84.55% translator validation, but the worst-link and injectivity gates still failed. Batch 25 extended that catch-up to 45,000 steps and improved again to 85.26% train, 85.24% validation and 85.13% translator validation. Receiver performance now nearly reaches the mathematical ceiling of the non-injective sender codebooks, so another horizon-only extension is not justified. Batch 26 added one sender/receiver route cycle and reached 85.38% validation, but did not add a single net unique code: shared errors became link-specific errors instead. Batch 27 directly penalized globally mined sender collisions after step 30,000. It reduced collision multiplicity and reached 85.38% validation, but unique code counts again remained unchanged and worst-link validation regressed to 82.13%. Both replay routing and the existing collision-loss schedule are now closed. Batch 28 added an identity-initialized generic residual sender block, but it collapsed to 9.65% validation and roughly 8,000 unique messages. Generic depth active from the start promotes memorization rather than compositional generalization. Batch 29 activated that branch only after the exact B25 trajectory; it still reduced validation to 84.74% and removed occupied codes. ECP-7 development is therefore closed after 29 variants. No confirmatory ECP-7 configuration was frozen and its confirmatory split was never opened. See docs/development-log-ecp7.md. The compact synthesis is in docs/ecp7-development-conclusion.md.

ECP-8 Batch 1 is a valid negative sealed-development result. It tested whether the strongest generic ECP-7 architecture becomes injective when all four positions may use the complete 16-symbol vocabulary: 16 fixed bits instead of the 14-bit lower bound. Compared with the paired 14-bit control, mean train exactness rose from 79.55% to 98.76%, worst-link validation from 64.84% to 73.34%, and the minimum codebook from 12,048 to 15,095 messages. The arm still had 225–265 collisions per sender and failed both 80% validation thresholds. The ECP-8 confirmatory split remains sealed. See docs/development-log-ecp8.md.

Build on this work

A new developer or fresh AI agent can start with no previous chat context. First read AGENTS.md for the binding research rules and then follow docs/AI_AGENT_START.md for installation, architecture overview, reproduction and proper setup of ECP-7.

Contributions are welcome via CONTRIBUTING.md. Large generated runs stay local; compact verifiable results come under evidence/.

Language and frozen provenance

All reader-facing documentation, schemas, command-line output, and source commentary are in English. Two reproducibility artifacts intentionally retain original Dutch literals:

  • the Dutch UTF-8 baseline sentence in src/ai_taal/baselines.py, because its byte length is a measured comparison;
  • descriptive metadata and comments inside existing ECP-0 through ECP-6 YAML configurations, because their exact bytes are preregistered and referenced by published SHA-256 hashes.

Do not translate those artifacts in place. New experiment configurations and all new public documentation must be written in English.

Files

ECP-0 in one minute

  • The world contains 8 × 8 × 4 × 4 = 1024 possible meanings.
  • Each meaning consists of four independent factors: color, shape, size and texture.
  • The agents only see numerical categories, not human labels.
  • The channel has 16 possible symbols and exactly 4 positions: 16 bits per message.
  • The sender and receiver do not share weights, embeddings or state.
  • 128 meanings are completely excluded from training as unknown color-shape combinations.
  • After training, the protocol is frozen and translated by a third model.
  • Five independent training runs prevent conclusions based on one random code.

The theoretical lower limit for exactly distinguishing 1024 uniform meanings is 10 bits. ECP-0 deliberately uses 16 bits: more than enough to first test learnability, compositionality and translatability. Compression to 12 bits and then towards the 10-bit lower limit is part of step 2.

Simulator

The reproducible simulator includes:

  1. deterministic dataset generation;
  2. separate sender, receiver and translation models;
  3. discrete evaluation without hidden side channel;
  4. automatic baselines, checks, measurements and run reports.

Execute

Required: Python 3.12.

python3.12 -m venv .venv
.venv/bin/pip install -e '.[dev]'
.venv/bin/ecp0 validate
.venv/bin/pytest

The three output modes deliberately access the data differently:

# 25 steps: verify that the complete technical pipeline works
.venv/bin/ecp0 smoke --seed 11

# Full training on train and validation; the test split remains sealed
.venv/bin/ecp0 develop --seed 11

# Confirmatory run: all preregistered seeds and one-time test unsealing
.venv/bin/ecp0 experiment --unseal-test

# Verify hashes and compare protocol structure with random assignments
.venv/bin/ecp0 analyze runs/<run-id> --permutations 100

# ECP-1 uses the same interface with the population configuration
.venv/bin/ecp1 --config config/ecp1.yaml validate
.venv/bin/ecp1 --config config/ecp1.yaml analyze runs/<ecp1-run-id> --permutations 100

# ECP-3: injective atomic codes and learned protocol consensus
.venv/bin/ecp3 --config config/ecp3.yaml validate
.venv/bin/ecp3 --config config/ecp3.yaml analyze runs/<ecp3-run-id> --permutations 100

# ECP-5: theoretically minimal 10-bit code with a calibrated reader
.venv/bin/ecp5 --config config/ecp5.yaml validate
.venv/bin/ecp5 --config config/ecp5.yaml analyze runs/<ecp5-run-id> --permutations 100

# ECP-6: 16,384 meanings at the exact 14-bit lower bound
.venv/bin/ecp6 --config config/ecp6.yaml validate
.venv/bin/ecp6 --config config/ecp6.yaml analyze runs/<ecp6-run-id> --permutations 100

# ECP-7 sealed development; never add --unseal-test to these configs
.venv/bin/ecp7 --config config/ecp7-development.yaml validate
.venv/bin/ecp7 --config config/ecp7-development.yaml develop --seed 11
.venv/bin/ecp7 --config config/ecp7-development.yaml analyze runs/<ecp7-development-run-id> --permutations 100

Each run writes configuration, environment, checkpoints, raw messages, checks, hashes, metrics, and a report to a new folder under runs/.

About

Reproducible emergent AI communication experiments at the information-theoretic limit

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages