Using gh-ost v1.1.11 (2ed192cd39d8a5173b6ef7c4beaa746a2d3bf1aa) on Aurora MySQL 8.0 with binlog_format=ROW, without --gtid.
After running with --checkpoint --execute, stopping the process, and restarting with --resume, binlog streaming repeatedly fails with:
StreamEvents encountered unexpected error:
... invalid table id, no corresponding table map event
The streamer repeatedly reconnects at the same position. Row copying can reach 100%, but binlog processing remains stuck.
Evidence
The following is an actual checkpoint and surrounding binlog events, with database/table names anonymized and irrelevant columns omitted. Numeric positions are unchanged.
mysql> SELECT gh_ost_chk_id, gh_ost_chk_coords, gh_ost_is_cutover
FROM app_db._table_a_ghk;
gh_ost_chk_id: 1
gh_ost_chk_coords: mysql-bin-changelog.001644:98771662
gh_ost_is_cutover: 0
mysql> SHOW BINLOG EVENTS
IN 'mysql-bin-changelog.001644' FROM 98746801 LIMIT 10;
Pos Event_type End_log_pos Info
98746801 Anonymous_Gtid 98746886 SET @@SESSION.GTID_NEXT='ANONYMOUS'
98746886 Query 98746991 BEGIN
98746991 Table_map 98747104 table_id: 5131 (app_db._other_table_gho; another concurrent migration)
98747104 Write_rows 98755290 table_id: 5131
98755290 Write_rows 98763476 table_id: 5131
98763476 Write_rows 98771662 table_id: 5131
98771662 Write_rows 98779848 table_id: 5131 <-- saved checkpoint
98779848 Write_rows 98786279 table_id: 5131 flags: STMT_END_F
98786279 Xid 98786310 COMMIT
98786310 Anonymous_Gtid 98786395 SET @@SESSION.GTID_NEXT='ANONYMOUS'
The checkpoint points to a Write_rows event after its required Table_map, rather than a transaction boundary.
Suspected cause
Checkpoint() saves GetCurrentBinlogCoordinates() as LastTrxCoords.
|
// Checkpoint attempts to write a checkpoint of the Migrator's current state. |
|
// It gets the binlog coordinates of the last received trx and waits until the |
|
// applier reaches that trx. At that point it's safe to resume from these coordinates. |
|
func (mgtr *Migrator) Checkpoint(ctx context.Context) (*Checkpoint, error) { |
|
coords := mgtr.eventsStreamer.GetCurrentBinlogCoordinates() |
|
mgtr.applier.LastIterationRangeMutex.Lock() |
|
if mgtr.applier.LastIterationRangeMaxValues == nil || mgtr.applier.LastIterationRangeMinValues == nil { |
|
mgtr.applier.LastIterationRangeMutex.Unlock() |
|
return nil, errors.New("iteration range is empty, not checkpointing") |
|
} |
|
chk := &Checkpoint{ |
|
Iteration: mgtr.migrationContext.GetIteration(), |
|
IterationRangeMin: mgtr.applier.LastIterationRangeMinValues.Clone(), |
|
IterationRangeMax: mgtr.applier.LastIterationRangeMaxValues.Clone(), |
|
LastTrxCoords: coords, |
|
RowsCopied: atomic.LoadInt64(&mgtr.migrationContext.TotalRowsCopied), |
|
DMLApplied: atomic.LoadInt64(&mgtr.migrationContext.TotalDMLEventsApplied), |
|
} |
|
mgtr.applier.LastIterationRangeMutex.Unlock() |
|
|
|
for { |
|
if err := ctx.Err(); err != nil { |
|
return nil, err |
|
} |
|
mgtr.applier.CurrentCoordinatesMutex.Lock() |
|
if coords.SmallerThanOrEquals(mgtr.applier.CurrentCoordinates) { |
|
id, err := mgtr.applier.WriteCheckpoint(chk) |
- In file-position mode,
currentCoordinates advances on every event, including intermediate row events.
|
// Update binlog coords if using file-based coords. |
|
// GTID coordinates are updated on receiving GTID events. |
|
if !gmr.migrationContext.UseGTIDs { |
|
gmr.currentCoordinatesMutex.Lock() |
|
coords := gmr.currentCoordinates.(*mysql.FileBinlogCoordinates) |
|
prevCoords := coords.Clone().(*mysql.FileBinlogCoordinates) |
|
coords.LogPos = int64(ev.Header.LogPos) |
|
coords.EventSize = int64(ev.Header.EventSize) |
|
if coords.IsLogPosOverflowBeyond4Bytes(prevCoords) { |
|
gmr.currentCoordinatesMutex.Unlock() |
|
return fmt.Errorf("unexpected rows event at %+v, the binlog end_log_pos is overflow 4 bytes", coords) |
|
} |
|
gmr.currentCoordinatesMutex.Unlock() |
|
} |
- Resume loads the saved coordinates and passes that position directly to
StartSync(). The new reader therefore lacks the preceding table map.
|
if mgtr.migrationContext.Resume { |
|
lastCheckpoint, err := mgtr.applier.ReadLastCheckpoint() |
|
if err != nil { |
|
return mgtr.migrationContext.Log.Errorf("no checkpoint found, unable to resume: %+v", err) |
|
} |
|
mgtr.migrationContext.Log.Infof("Resuming from checkpoint coords=%+v range_min=%+v range_max=%+v iteration=%d", |
|
lastCheckpoint.LastTrxCoords, lastCheckpoint.IterationRangeMin.String(), lastCheckpoint.IterationRangeMax.String(), lastCheckpoint.Iteration) |
|
|
|
mgtr.migrationContext.MigrationIterationRangeMinValues = lastCheckpoint.IterationRangeMin |
|
mgtr.migrationContext.MigrationIterationRangeMaxValues = lastCheckpoint.IterationRangeMax |
|
mgtr.migrationContext.Iteration = lastCheckpoint.Iteration |
|
mgtr.migrationContext.TotalRowsCopied = lastCheckpoint.RowsCopied |
|
mgtr.migrationContext.TotalDMLEventsApplied = lastCheckpoint.DMLApplied |
|
mgtr.migrationContext.InitialStreamerCoords = lastCheckpoint.LastTrxCoords |
|
if err := mgtr.initiateStreaming(); err != nil { |
|
return err |
|
// Start sync with specified GTID set or binlog file and position |
|
if gmr.migrationContext.UseGTIDs { |
|
coords := coordinates.(*mysql.GTIDBinlogCoordinates) |
|
gmr.binlogStreamer, err = gmr.binlogSyncer.StartSyncGTID(coords.GTIDSet.Clone()) |
|
} else { |
|
coords := gmr.currentCoordinates.(*mysql.FileBinlogCoordinates) |
|
gmr.binlogStreamer, err = gmr.binlogSyncer.StartSync(gomysql.Position{ |
|
Name: coords.LogFile, |
|
Pos: uint32(coords.LogPos)}, |
|
) |
|
} |
The reader separately tracks completed transactions in LastTrxCoords on XID events, but checkpoint creation does not use that value.
|
case *replication.XIDEvent: |
|
if gmr.migrationContext.UseGTIDs { |
|
// event.GSet is the full executed GTID set maintained by the |
|
// syncer (MysqlGTIDSet or MariadbGTIDSet depending on flavor). |
|
if event.GSet != nil { |
|
gmr.LastTrxCoords = &mysql.GTIDBinlogCoordinates{GTIDSet: event.GSet} |
|
} |
|
} else { |
|
gmr.LastTrxCoords = gmr.currentCoordinates.Clone() |
|
} |
Expected behavior
File-position checkpoints should store a restartable transaction boundary, while ensuring all preceding DML has been applied, so --resume can continue.
Using gh-ost v1.1.11 (
2ed192cd39d8a5173b6ef7c4beaa746a2d3bf1aa) on Aurora MySQL 8.0 withbinlog_format=ROW, without--gtid.After running with
--checkpoint --execute, stopping the process, and restarting with--resume, binlog streaming repeatedly fails with:The streamer repeatedly reconnects at the same position. Row copying can reach 100%, but binlog processing remains stuck.
Evidence
The following is an actual checkpoint and surrounding binlog events, with database/table names anonymized and irrelevant columns omitted. Numeric positions are unchanged.
The checkpoint points to a
Write_rowsevent after its requiredTable_map, rather than a transaction boundary.Suspected cause
Checkpoint()savesGetCurrentBinlogCoordinates()asLastTrxCoords.gh-ost/go/logic/migrator.go
Lines 1812 to 1838 in 2ed192c
currentCoordinatesadvances on every event, including intermediate row events.gh-ost/go/binlog/gomysql_reader.go
Lines 163 to 176 in 2ed192c
StartSync(). The new reader therefore lacks the preceding table map.gh-ost/go/logic/migrator.go
Lines 580 to 595 in 2ed192c
gh-ost/go/binlog/gomysql_reader.go
Lines 92 to 102 in 2ed192c
The reader separately tracks completed transactions in
LastTrxCoordson XID events, but checkpoint creation does not use that value.gh-ost/go/binlog/gomysql_reader.go
Lines 216 to 225 in 2ed192c
Expected behavior
File-position checkpoints should store a restartable transaction boundary, while ensuring all preceding DML has been applied, so
--resumecan continue.