fix: persist WAL tail truncation per design protocol (C5+H8)

Recovery tail-truncation had two compounding bugs in wal/recover.go:

C5: truncateSegment only called os.Truncate. Missing per design §3.2
    line 787-794:
      - Step 2: fsync the truncated segment
      - Step 3: delete empty trailing segments
      - Step 4: fsync WAL directory
    And all errors were swallowed into result.TruncateError with recovery
    still returning success, violating design line 799: "若 ftruncate、
    segment fsync、空 segment 删除或 WAL directory fsync 任一步失败,
    recovery 必须报错,DB 不得进入可写状态".

H8: findValidOffset only checked physical record CRCs, ignoring the
    FragmentCollector state machine. For a tail of First + Middle*
    without Last, it returned the offset AFTER the last Middle fragment
    instead of the last COMPLETE batch end. Result: residual half-batch
    fragments caused repeated tail-corruption reports on every restart.

Changes:
- wal/recover.go:
  - Add findLastCompleteBatchEnd: batch-aware offset finder using
    FragmentCollector state machine. Handles block-boundary padding
    correctly (continue across full-block padding, return on short-block).
  - Add truncateAndPersist: 4-step protocol (ftruncate + fsync segment +
    delete empty trailing + fsync dir). Any step failure is fatal.
  - Add segmentFsyncFn (package-level var for test injection, same
    pattern as C6's dirFsyncFn).
  - Refactor Recover failure path: use new functions, hard-error on
    truncation persist failure (was: swallow to TruncateError).
  - TruncateError field semantics: informational only ("tail corruption
    was detected and repair attempted"). Persist failures return error.
  - Delete findValidOffset and truncateSegment (replaced).
- wal/recover_offset_test.go (new): 8 unit tests for
  findLastCompleteBatchEnd covering clean/partial-tail/no-batch/
  physical-corruption/partial-only/block-boundary-padding/non-zero-tail/
  zero-tail cases. 5 unit tests for truncateAndPersist covering success/
  ftruncate-fail/dir-fsync-fail/segment-fsync-fail/retry-after-failure.
- wal/recover_test.go: add TestRecoverPartialFragmentTailIdempotent
  (H8 e2e regression: truncation point must be at last complete batch),
  TestRecoverTruncationFailureFailsRecovery (C5 e2e regression: any
  step failure fails Recover), TestRecoverInvalidBatchNotTruncatable
  (design line 778-781: invalid batch content hard-fails, NOT truncatable).

Injection note: segmentFsyncFn and dirFsyncFn (from C6) are package-level
vars; tests that override either must not use t.Parallel().

Verified: each new test fails on pre-fix code by logical analysis and
passes after the fix. Full suite green including go test -race ./... .

Phase 1 simplification: emptyTrailingSegments is always nil in Phase 1
(truncated segment is always segments[last]). The parameter is kept in
truncateAndPersist's signature for forward compatibility with the C4 fix.

Audit context: docs/audit-3.2.md C5 and H8 (H8 Oracle-verified bg_ef425776).
This commit is contained in:
dailz
2026-06-15 15:15:49 +08:00
parent 0739966e55
commit 273229ac9b
4 changed files with 1202 additions and 72 deletions
+129 -72
View File
@@ -13,12 +13,15 @@ type RecoveryResult struct {
NextSegmentID uint64
ReplayedEntries int
Truncated bool
TruncateError error // non-nil if tail corruption was found
// TruncateError is informational only: non-nil means "tail corruption
// was found and repair was attempted". It does NOT report persistence
// failures — those cause Recover to return an error instead.
TruncateError error
}
// Recover performs a full WAL recovery: reads the recovery checkpoint from
// MANIFEST (or CURRENT), scans segments, replays entries, and handles tail
// truncation. On success the MANIFEST is updated with the new recovery state.
// MANIFEST, scans segments, replays entries, and persists tail truncation
// per design §3.2 line 787-800.
func Recover(dir string, replayer BatchReplayer) (*RecoveryResult, error) {
if replayer == nil {
return nil, fmt.Errorf("wal: recover: replayer is nil")
@@ -37,15 +40,14 @@ func Recover(dir string, replayer BatchReplayer) (*RecoveryResult, error) {
return nil, fmt.Errorf("wal: recover: %w", err)
}
// Step 3: Tail corruption — truncate the last segment and accept
// partial data loss for Phase 1.
// Step 3: Tail corruption — truncate the last segment per design
// §3.2 line 787-800.
result := &RecoveryResult{
NextSequence: nextSequence,
Truncated: true,
TruncateError: err,
TruncateError: err, // informational: tail corruption was detected
}
// Determine nextSegmentID from scanned segments.
segments, scanErr := ScanSegments(dir, recoverySegmentID)
if scanErr != nil {
return nil, fmt.Errorf("wal: recover: scan after tail corruption: %w", scanErr)
@@ -56,27 +58,31 @@ func Recover(dir string, replayer BatchReplayer) (*RecoveryResult, error) {
result.NextSegmentID = recoverySegmentID
}
// Truncate the last segment file to remove corrupted tail.
if len(segments) > 0 {
lastSeg := segments[len(segments)-1]
validOffset, truncErr := findValidOffset(lastSeg.FilePath)
if truncErr != nil {
// Best-effort: record the truncation error but don't fail recovery.
result.TruncateError = fmt.Errorf("%w (find valid offset: %v)", err, truncErr)
} else if truncErr := truncateSegment(lastSeg.FilePath, validOffset); truncErr != nil {
result.TruncateError = fmt.Errorf("%w (truncate: %v)", err, truncErr)
lastCompleteBatchEnd, findErr := findLastCompleteBatchEnd(lastSeg.FilePath)
if findErr != nil {
return nil, fmt.Errorf("wal: recover: find truncation offset: %w", findErr)
}
// Phase 1: truncated segment is always segments[last], no
// trailing empty segments to clean up. C4 fix will need to
// identify trailing empties based on the actually-corrupted
// segment's index, which is not the same as segments[last].
var emptyTrailing []string
if err := truncateAndPersist(lastSeg.FilePath, lastCompleteBatchEnd, dir, emptyTrailing); err != nil {
// Per design §3.2 line 799: DB must NOT enter writable state.
return nil, fmt.Errorf("wal: recover: persist tail truncation: %w", err)
}
}
// Count replayed entries by re-scanning the replayer state.
// For Phase 1 we accept that ReplayedEntries may be approximate;
// the replayer interface doesn't expose a count.
result.ReplayedEntries = 0 // caller can inspect replayer directly
result.ReplayedEntries = 0 // Phase 1: replayer interface doesn't expose count
// Per design §3.2 line 280, recovery repair must NOT update MANIFEST.
// The truncated WAL state is persisted via ftruncate (see C5 for the
// remaining fsync gaps). MANIFEST can only advance via checkpoint
// (MemTable flush) in future phases.
// The truncated WAL state is persisted via ftruncate + fsync segment
// + fsync dir (see truncateAndPersist). MANIFEST can only advance via
// checkpoint (MemTable flush) in future phases.
return result, nil
}
@@ -114,23 +120,76 @@ func resolveRecoverySegmentID(dir string) (uint64, error) {
return mf.RecoverySegmentID, nil
}
// truncateSegment truncates the file at filePath to validOffset bytes,
// removing any corrupted data after that point.
func truncateSegment(filePath string, validOffset int64) error {
if validOffset < 0 {
return fmt.Errorf("wal: truncate: invalid offset %d", validOffset)
// segmentFsyncFn is the package-level indirection for fsyncing a truncated
// segment file. Tests that override this must not use t.Parallel().
// Same pattern as dirFsyncFn (see wal/dir_fsync.go).
var segmentFsyncFn = segmentFsync
func segmentFsync(filePath string) error {
f, err := os.OpenFile(filePath, os.O_WRONLY, 0o644)
if err != nil {
return fmt.Errorf("open for fsync: %w", err)
}
return os.Truncate(filePath, validOffset)
defer f.Close()
if err := f.Sync(); err != nil {
return fmt.Errorf("fsync: %w", err)
}
return nil
}
// findValidOffset parses a segment file and returns the byte offset of the
// last valid record boundary. The offset includes the file header size.
func findValidOffset(filePath string) (int64, error) {
// Re-parse the file to find where valid records end.
// We need to track the byte offset as we parse.
// truncateAndPersist executes the 4-step tail-truncation protocol per
// design §3.2 line 787-794. Any step failure is fatal: per line 799,
// DB must NOT enter writable state if truncation cannot be persisted.
func truncateAndPersist(
filePath string,
lastCompleteBatchEnd int64,
dir string,
emptyTrailingSegments []string,
) error {
// Step 1: ftruncate
if lastCompleteBatchEnd < 0 {
return fmt.Errorf("wal: invalid truncate offset %d", lastCompleteBatchEnd)
}
if err := os.Truncate(filePath, lastCompleteBatchEnd); err != nil {
return fmt.Errorf("ftruncate %s to %d: %w", filePath, lastCompleteBatchEnd, err)
}
// Step 2: fsync the truncated segment
if err := segmentFsyncFn(filePath); err != nil {
return fmt.Errorf("fsync truncated segment: %w", err)
}
// Step 3: delete empty trailing segments
for _, segPath := range emptyTrailingSegments {
if err := os.Remove(segPath); err != nil {
return fmt.Errorf("remove empty segment %s: %w", segPath, err)
}
}
// Step 4: fsync WAL directory (reuses C6's dirFsyncFn)
if err := dirFsyncFn(dir); err != nil {
return fmt.Errorf("fsync WAL dir after truncation: %w", err)
}
return nil
}
// findLastCompleteBatchEnd walks the segment file, runs physical records
// through the FragmentCollector state machine, and returns the byte offset
// of the END of the last complete WAL Batch.
//
// This is the correct truncation target per design §3.2 line 786. A previous
// version (findValidOffset) only checked physical record CRCs, missing the
// case where a First + Middle* fragment chain has no Last (H8 bug): physical
// CRCs pass but no complete batch exists at that offset.
//
// Block-boundary handling: WAL format allows a full block to end with zero
// padding when the next record doesn't fit (see BlockWriter.paddingNeeded).
// This function CONTINUES to the next block on padding in a full block, and
// only RETURNS on padding in a short (final) block or actual corruption.
func findLastCompleteBatchEnd(filePath string) (int64, error) {
f, err := os.Open(filePath)
if err != nil {
return 0, fmt.Errorf("open for offset scan: %w", err)
return 0, fmt.Errorf("open %s: %w", filePath, err)
}
defer f.Close()
@@ -138,58 +197,56 @@ func findValidOffset(filePath string) (int64, error) {
return 0, fmt.Errorf("seek past header: %w", err)
}
validOffset := int64(WalFileHeaderSize)
collector := NewFragmentCollector()
lastCompleteEnd := int64(WalFileHeaderSize)
blockStartOffset := int64(WalFileHeaderSize)
buf := make([]byte, WalBlockSize)
for {
n, readErr := f.Read(buf)
if readErr != nil {
break
}
if n == 0 {
break
}
if n > 0 {
blockData := buf[:n]
isFullBlock := n == WalBlockSize && readErr == nil
pos := 0
for pos < len(blockData) {
remaining := len(blockData) - pos
blockData := buf[:n]
blockStartOffset := validOffset
pos := 0
for pos < len(blockData) {
remaining := len(blockData) - pos
if remaining < PhysicalRecordHeaderSize {
// Check if remaining bytes are zero-padding.
if isAllZeros(blockData[pos:]) {
// Valid padding — update offset to end of last valid record.
validOffset = blockStartOffset + int64(pos)
if remaining < PhysicalRecordHeaderSize {
if isFullBlock {
break // padding in full block, continue to next block
}
return lastCompleteEnd, nil // tail padding in short block
}
// Either way, we're done with this block.
break
}
if isAllZeros(blockData[pos : pos+PhysicalRecordHeaderSize]) {
if isAllZeros(blockData[pos:]) {
validOffset = blockStartOffset + int64(pos)
if isAllZeros(blockData[pos : pos+PhysicalRecordHeaderSize]) {
if isFullBlock {
break // zero-led padding in full block, continue
}
return lastCompleteEnd, nil // tail padding
}
break
}
_, consumed, err := DecodePhysicalRecord(blockData[pos:])
if err != nil {
// Corruption starts here — offset is up to last valid record.
validOffset = blockStartOffset + int64(pos)
return validOffset, nil
}
rec, consumed, err := DecodePhysicalRecord(blockData[pos:])
if err != nil {
return lastCompleteEnd, nil // physical corruption
}
if err := collector.Append(rec.Type, rec.Payload); err != nil {
return lastCompleteEnd, nil // fragment state machine rejected
}
// Valid record found.
validOffset = blockStartOffset + int64(pos+consumed)
pos += consumed
pos += consumed
recordEndAbsolute := blockStartOffset + int64(pos)
if collector.IsComplete() {
lastCompleteEnd = recordEndAbsolute
collector.Reset()
}
}
blockStartOffset += int64(n)
}
if n < WalBlockSize {
if readErr != nil || n < WalBlockSize {
break
}
}
return validOffset, nil
return lastCompleteEnd, nil
}
+363
View File
@@ -0,0 +1,363 @@
package wal
import (
"errors"
"os"
"path/filepath"
"strings"
"testing"
)
// writeRawSegmentHeader writes a 32-byte WAL header to filePath. Used by
// findLastCompleteBatchEnd tests to construct minimal segment files.
func writeRawSegmentHeader(t *testing.T, filePath string) {
t.Helper()
hdr := &WalFileHeader{
BlockSize: 32 * 1024,
SegmentID: 0,
StartSequence: 0,
}
encoded := EncodeWalHeader(hdr)
if err := os.WriteFile(filePath, encoded[:], 0o644); err != nil {
t.Fatalf("WriteFile header: %v", err)
}
}
// appendRawBytes appends arbitrary bytes to filePath.
func appendRawBytes(t *testing.T, filePath string, data []byte) {
t.Helper()
f, err := os.OpenFile(filePath, os.O_WRONLY|os.O_APPEND, 0o644)
if err != nil {
t.Fatalf("OpenFile append: %v", err)
}
defer f.Close()
if _, err := f.Write(data); err != nil {
t.Fatalf("Write: %v", err)
}
}
// appendFullRecord encodes and appends a complete WAL batch as a Full record.
func appendFullRecord(t *testing.T, filePath string, batch []byte) {
t.Helper()
rec := EncodePhysicalRecord(RecFull, batch)
appendRawBytes(t, filePath, rec)
}
// appendFragment encodes and appends a single fragment record.
func appendFragment(t *testing.T, filePath string, recType uint8, payload []byte) {
t.Helper()
rec := EncodePhysicalRecord(recType, payload)
appendRawBytes(t, filePath, rec)
}
func TestFindLastCompleteBatchEnd_CleanSegment(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
batchA := []byte("batch-A-content")
batchB := []byte("batch-B-content")
appendFullRecord(t, filePath, batchA)
endOfA, _ := fileSize(filePath)
appendFullRecord(t, filePath, batchB)
endOfB, _ := fileSize(filePath)
got, err := findLastCompleteBatchEnd(filePath)
if err != nil {
t.Fatalf("findLastCompleteBatchEnd: %v", err)
}
if got != endOfB {
t.Errorf("got %d, want %d (end of Batch B)", got, endOfB)
}
if got == endOfA {
t.Errorf("got end of Batch A, should be end of Batch B")
}
}
// Regression guard for H8: partial fragment tail must return end of last
// COMPLETE batch, not end of last physical record.
func TestFindLastCompleteBatchEnd_PartialTailFragment(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
batchA := []byte("batch-A-content")
appendFullRecord(t, filePath, batchA)
endOfA, _ := fileSize(filePath)
// Append First + Middle fragments (no Last) — H8 case.
appendFragment(t, filePath, RecFirst, []byte("first-fragment-data"))
appendFragment(t, filePath, RecMiddle, []byte("middle-fragment-data"))
got, err := findLastCompleteBatchEnd(filePath)
if err != nil {
t.Fatalf("findLastCompleteBatchEnd: %v", err)
}
if got != endOfA {
t.Errorf("got %d, want %d (end of Batch A, NOT end of Middle fragment)", got, endOfA)
}
}
func TestFindLastCompleteBatchEnd_NoBatches(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
got, err := findLastCompleteBatchEnd(filePath)
if err != nil {
t.Fatalf("findLastCompleteBatchEnd: %v", err)
}
if got != int64(WalFileHeaderSize) {
t.Errorf("got %d, want %d (WalFileHeaderSize)", got, WalFileHeaderSize)
}
}
func TestFindLastCompleteBatchEnd_PhysicalCorruption(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
batchA := []byte("batch-A-content")
appendFullRecord(t, filePath, batchA)
endOfA, _ := fileSize(filePath)
// Append corrupted bytes (will fail CRC).
appendRawBytes(t, filePath, []byte{0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF})
got, err := findLastCompleteBatchEnd(filePath)
if err != nil {
t.Fatalf("findLastCompleteBatchEnd: %v", err)
}
if got != endOfA {
t.Errorf("got %d, want %d (end of Batch A, before corruption)", got, endOfA)
}
}
func TestFindLastCompleteBatchEnd_PartialBatchOnly(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
// Only First + Middle fragments, no complete batch ever.
appendFragment(t, filePath, RecFirst, []byte("first-data"))
appendFragment(t, filePath, RecMiddle, []byte("middle-data"))
got, err := findLastCompleteBatchEnd(filePath)
if err != nil {
t.Fatalf("findLastCompleteBatchEnd: %v", err)
}
if got != int64(WalFileHeaderSize) {
t.Errorf("got %d, want %d (no complete batch, stay at header)", got, WalFileHeaderSize)
}
}
// Oracle BLOCKING test: must continue past padding in a full block to read
// the next block. Original code returned on first padding, truncating all
// later batches.
func TestFindLastCompleteBatchEnd_BlockBoundaryPadding(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
// Batch A: large enough to nearly fill block 1.
// Block 1 layout: [32-byte header][Batch A (Full record)] [padding to end]
// We need: 32 + len(Full record of A) + padding == 32 + WalBlockSize
// Full record = 7 (header) + len(payload). So:
// len(A) such that 7 + len(A) leaves < 7 bytes before block end
// Then writer pads to end of block and Batch B goes in block 2.
// Available in block 1 after file header = WalBlockSize - 32 = 32736
// We want Batch A record size to be 32736 - 6 = 32730 (leaving 6 bytes, < 7, padding)
// So payload size = 32730 - 7 = 32723
// Batch A = batch header (18) + entry bytes. Entry: Put with key + value.
// Simplest: use raw bytes (we're testing physical layout, not batch validity).
bigPayload := make([]byte, 32723)
for i := range bigPayload {
bigPayload[i] = byte('A')
}
recA := EncodePhysicalRecord(RecFull, bigPayload)
appendRawBytes(t, filePath, recA)
endOfBlock1 := int64(WalFileHeaderSize + WalBlockSize) // 32 + 32768
currentSize, _ := fileSize(filePath)
// Pad to end of block 1 with zeros.
padLen := int(endOfBlock1 - currentSize)
if padLen > 0 {
appendRawBytes(t, filePath, make([]byte, padLen))
}
// Batch B in block 2 (normal small batch after padding boundary).
batchB := []byte("batch-B")
recB := EncodePhysicalRecord(RecFull, batchB)
appendRawBytes(t, filePath, recB)
endOfB, _ := fileSize(filePath)
got, err := findLastCompleteBatchEnd(filePath)
if err != nil {
t.Fatalf("findLastCompleteBatchEnd: %v", err)
}
if got != endOfB {
t.Errorf("got %d, want %d (end of Batch B in block 2, after padding)", got, endOfB)
}
// Specifically: must NOT be at end of Batch A (would be ~32755, before padding).
if got < endOfBlock1 {
t.Errorf("got %d < %d (returned at end of Batch A, missed block 2)", got, endOfBlock1)
}
}
func TestFindLastCompleteBatchEnd_NonZeroTailPadding(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
batchA := []byte("batch-A")
appendFullRecord(t, filePath, batchA)
endOfA, _ := fileSize(filePath)
// Append 3 non-zero bytes (< PhysicalRecordHeaderSize=7). This is
// technically invalid padding (per design: padding must be zeros), but
// findLastCompleteBatchEnd should still return end of last complete
// batch — this is "tail corruption" classification territory.
appendRawBytes(t, filePath, []byte{0xFF, 0xFF, 0xFF})
got, err := findLastCompleteBatchEnd(filePath)
if err != nil {
t.Fatalf("findLastCompleteBatchEnd: %v", err)
}
if got != endOfA {
t.Errorf("got %d, want %d (end of Batch A)", got, endOfA)
}
}
func TestFindLastCompleteBatchEnd_ZeroTailPadding(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
batchA := []byte("batch-A")
appendFullRecord(t, filePath, batchA)
endOfA, _ := fileSize(filePath)
// Append 3 zero bytes (valid tail padding in a short final block).
appendRawBytes(t, filePath, []byte{0, 0, 0})
got, err := findLastCompleteBatchEnd(filePath)
if err != nil {
t.Fatalf("findLastCompleteBatchEnd: %v", err)
}
if got != endOfA {
t.Errorf("got %d, want %d (end of Batch A)", got, endOfA)
}
}
func fileSize(filePath string) (int64, error) {
fi, err := os.Stat(filePath)
if err != nil {
return 0, err
}
return fi.Size(), nil
}
// -------- truncateAndPersist tests --------
func TestTruncateAndPersist_Success(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
appendRawBytes(t, filePath, []byte("batch-A"))
appendRawBytes(t, filePath, []byte("extra-bytes-to-be-truncated"))
endOfA, _ := fileSize(filePath)
truncateOffset := int64(WalFileHeaderSize) + 7 // just past header, before "batch-A"
if err := truncateAndPersist(filePath, truncateOffset, dir, nil); err != nil {
t.Fatalf("truncateAndPersist: %v", err)
}
gotSize, _ := fileSize(filePath)
if gotSize != truncateOffset {
t.Errorf("file size = %d, want %d (truncated)", gotSize, truncateOffset)
}
_ = endOfA // not used; verify only that size matches truncate point
}
func TestTruncateAndPersist_FtruncateFailure(t *testing.T) {
dir := t.TempDir()
nonExistent := filepath.Join(dir, "no-such-file.wal")
err := truncateAndPersist(nonExistent, 100, dir, nil)
if err == nil {
t.Fatal("expected error on non-existent file")
}
if !strings.Contains(err.Error(), "ftruncate") {
t.Errorf("error should mention 'ftruncate', got: %v", err)
}
}
func TestTruncateAndPersist_DirFsyncFailure(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
appendRawBytes(t, filePath, []byte("extra"))
orig := dirFsyncFn
dirFsyncFn = func(string) error { return errors.New("simulated dir fsync failure") }
t.Cleanup(func() { dirFsyncFn = orig })
err := truncateAndPersist(filePath, int64(WalFileHeaderSize), dir, nil)
if err == nil {
t.Fatal("expected error on dir fsync failure")
}
if !strings.Contains(err.Error(), "fsync WAL dir") {
t.Errorf("error should mention 'fsync WAL dir', got: %v", err)
}
}
func TestTruncateAndPersist_SegmentFsyncFailure(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
appendRawBytes(t, filePath, []byte("extra"))
orig := segmentFsyncFn
segmentFsyncFn = func(string) error { return errors.New("simulated segment fsync failure") }
t.Cleanup(func() { segmentFsyncFn = orig })
err := truncateAndPersist(filePath, int64(WalFileHeaderSize), dir, nil)
if err == nil {
t.Fatal("expected error on segment fsync failure")
}
if !strings.Contains(err.Error(), "fsync truncated segment") {
t.Errorf("error should mention 'fsync truncated segment', got: %v", err)
}
}
// Oracle nice-to-have: verify retry after dir-fsync failure. The first
// call fails after ftruncate succeeded; the second call (with fsync
// restored) must succeed and the file must end up correctly truncated.
func TestTruncateAndPersist_RetryAfterDirFsyncFailure(t *testing.T) {
dir := t.TempDir()
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
appendRawBytes(t, filePath, []byte("extra-bytes"))
truncateOffset := int64(WalFileHeaderSize)
// First call: inject dir fsync failure.
orig := dirFsyncFn
dirFsyncFn = func(string) error { return errors.New("simulated") }
if err := truncateAndPersist(filePath, truncateOffset, dir, nil); err == nil {
t.Fatal("first call should fail")
}
dirFsyncFn = orig
// Second call: must succeed (ftruncate is idempotent).
if err := truncateAndPersist(filePath, truncateOffset, dir, nil); err != nil {
t.Fatalf("retry truncateAndPersist: %v", err)
}
gotSize, _ := fileSize(filePath)
if gotSize != truncateOffset {
t.Errorf("after retry, file size = %d, want %d", gotSize, truncateOffset)
}
}
+118
View File
@@ -2,9 +2,11 @@ package wal
import (
"bytes"
"errors"
"os"
"path/filepath"
"reflect"
"strings"
"testing"
"github.com/dailz/go-kv/manifest"
@@ -333,3 +335,119 @@ func fileExists(t *testing.T, path string) bool {
t.Fatalf("stat %s: %v", path, err)
return false
}
// Regression guard for C5+H8: end-to-end recovery with partial fragment
// tail must persist truncation at the last COMPLETE batch boundary
// (H8), and the truncation must be persisted with all 4 steps (C5).
// After repair, second recovery must not see corruption.
func TestRecoverPartialFragmentTailIdempotent(t *testing.T) {
dir := t.TempDir()
batchA := []*WalEntry{makePutEntry("key-A", "val-A")}
batchB := []*WalEntry{makePutEntry("key-B", "val-B")}
filePath := writeTestSegment(t, dir, 0, 0, [][]*WalEntry{batchA, batchB})
// Compute exact byte offset where Batch B's Full record ends.
// Layout: [header][Batch A Full record][Batch B Full record][padding to 32KB]
encA, _ := EncodeWalBatch(0, batchA)
encB, _ := EncodeWalBatch(1, batchB)
endOfBatchB := int64(WalFileHeaderSize) +
int64(PhysicalRecordHeaderSize+len(encA)) +
int64(PhysicalRecordHeaderSize+len(encB))
fiBefore, _ := os.Stat(filePath)
// Append First + Middle* (no Last) to simulate partial fragment tail.
appendFileBytes(t, filePath, EncodePhysicalRecord(RecFirst, []byte("first-fragment-payload")))
appendFileBytes(t, filePath, EncodePhysicalRecord(RecMiddle, []byte("middle-fragment-payload")))
replayer1 := &mockReplayer{}
result1, err := Recover(dir, replayer1)
if err != nil {
t.Fatalf("1st Recover: %v", err)
}
if !result1.Truncated {
t.Fatal("1st Recover: Truncated = false, want true")
}
fiAfter, _ := os.Stat(filePath)
if fiAfter.Size() != endOfBatchB {
t.Errorf("file size after truncation = %d, want %d (end of Batch B, H8)",
fiAfter.Size(), endOfBatchB)
}
if fiAfter.Size() >= fiBefore.Size() {
t.Errorf("file should shrink after truncation: before=%d after=%d",
fiBefore.Size(), fiAfter.Size())
}
replayer2 := &mockReplayer{}
result2, err := Recover(dir, replayer2)
if err != nil {
t.Fatalf("2nd Recover: %v", err)
}
if result2.Truncated {
t.Error("2nd Recover: Truncated = true, want false (truncation should be persisted)")
}
if result1.NextSequence != result2.NextSequence {
t.Errorf("NextSequence differs: %d vs %d", result1.NextSequence, result2.NextSequence)
}
}
// Regression guard for C5: any truncation persist step failure must
// fail Recover, causing DB.Open to fail. No swallowing allowed.
func TestRecoverTruncationFailureFailsRecovery(t *testing.T) {
dir := t.TempDir()
filePath := writeTestSegment(t, dir, 0, 0, [][]*WalEntry{
{makePutEntry("key-A", "val-A")},
})
// Append corruption to trigger tail corruption path.
appendFileBytes(t, filePath, []byte{0xDE, 0xAD, 0xBE, 0xEF})
// Inject dir fsync failure (Step 4 of truncateAndPersist).
orig := dirFsyncFn
dirFsyncFn = func(string) error { return errors.New("simulated dir fsync failure") }
t.Cleanup(func() { dirFsyncFn = orig })
_, err := Recover(dir, &mockReplayer{})
if err == nil {
t.Fatal("expected Recover to fail when truncation persist fails")
}
if !strings.Contains(err.Error(), "persist tail truncation") {
t.Errorf("error should mention 'persist tail truncation', got: %v", err)
}
}
// Regression guard for design line 778-781: CRC-valid but batch-content-
// invalid must hard-fail through DecodeWalBatch, NOT enter truncation path.
func TestRecoverInvalidBatchNotTruncatable(t *testing.T) {
dir := t.TempDir()
// Build segment with: physical records CRC-valid, but assembled batch
// has invalid header (entryCount=0).
filePath := filepath.Join(dir, "segment-0.wal")
writeRawSegmentHeader(t, filePath)
// Construct an "invalid batch": WalBatchHeaderSize=18 bytes, with
// entryCount=0 (invalid per ReplayBatch check at recovery.go).
invalidBatch := make([]byte, WalBatchHeaderSize)
// flags(2) + baseSequence(8) + entryCount(4)=0 + entriesSize(4)=0
// All zeros, except entryCount=0 is invalid by itself.
// Encode as Full physical record (CRC-valid).
rec := EncodePhysicalRecord(RecFull, invalidBatch)
appendFileBytes(t, filePath, rec)
_, err := Recover(dir, &mockReplayer{})
if err == nil {
t.Fatal("expected Recover to fail on invalid batch content")
}
// Should NOT mention truncation — must be a different error path.
if strings.Contains(err.Error(), "truncat") {
t.Errorf("error should not be about truncation; got: %v", err)
}
// File must NOT have been truncated (size unchanged).
fi, _ := os.Stat(filePath)
if fi.Size() != int64(WalFileHeaderSize)+int64(len(rec)) {
t.Errorf("file was truncated; size = %d, want %d",
fi.Size(), int64(WalFileHeaderSize)+int64(len(rec)))
}
}