Files

8.4 KiB

Lattice Roadmap

Milestone 1: Single-Node Append-Only Log

Goal: A single node can create, sign, and persist entries to its own log. No networking yet.

Deliverables

  • HLC timestamps
  • Node identity (Ed25519 keypair, save/load)
  • Entry signing & verification
  • Log file I/O (append, read, hash verification)
  • SigChain (validate entries before appending)
  • Store (redb) — kv + meta tables, log replay
  • Interactive CLI: init, put, get, delete, status, quit

Success Criteria

  • Can create a new identity
  • Can append entries to local log
  • Can replay log to reconstruct KV state
  • All operations survive restart

Multi-KV Refactoring (before M2) ✓

  • DataDir → stores/{uuid}/ subdirectories
  • Store → per-store state.db
  • Log paths → stores/{uuid}/logs/{author}.log
  • Proto: Entry has store_id (UUID)
  • CLI → init, create-store, list-stores, use
  • meta.db stores table (MetaStore)
  • SigChain → validate entry.store_id

Milestone 1.5: DAG Conflict Resolution

Goal: Upgrade store from simple LWW to DAG-based conflict resolution per architecture.md.

Deliverables

  • Proto: Add repeated bytes parent_hashes to Entry (for DAG causality)
  • Proto: Add HeadInfo message for multi-head storage
  • Store: KV table schema → Vec<u8> → Vec<HeadInfo>
  • Store: apply_entry → track multiple heads, merge parent tips
  • Store: get → deterministic winner (highest HLC, author tiebreaker)
  • Store: get_heads → inspect all heads for a key
  • EntryBuilder: .parent_hashes(...) method for DAG ancestry
  • CLI: Show conflict indicator when multiple heads

Success Criteria

  • Concurrent writes to same key create multiple heads
  • Reads return deterministic winner
  • Next write citing both heads merges fork to single tip
  • All existing tests still pass (71 tests)

Milestone 1.9: Async Refactor

Goal: Prepare codebase for concurrent CLI + network operation.

Deliverables

Phase 1: Store Actor (sync)

  • Store actor pattern: dedicated thread owns Store, receives commands via std::sync::mpsc
  • StoreHandle wraps channel sender, keeps current API
  • Validate: CLI works as before with actor

Phase 2: Async Runtime

  • Add tokio runtime (#[tokio::main])
  • Migrate std::sync::mpsctokio::sync::mpsc
  • Async CLI using block_in_place for sync handlers

Success Criteria

  • CLI still works as before
  • Store operations serialized (no data races)
  • Ready for concurrent network tasks

Milestone 2: Two-Node Sync

Goal: Two nodes can sync their logs over the network.

Deliverables

Phase 1: Sync Logic (no network)

  • SyncState with AuthorInfo (seq + hash) for hash-based log resumption
  • Store::sync_state() → author-to-seq+hash map from AUTHOR_TABLE
  • SyncState::diff()Vec<MissingRange> with from_hash for read_entries_after
  • Multi-store sync test: compute diff, fetch entries, apply, verify same state

Phase 2: Iroh Integration

Completed:

  • Node info in root store on init: /nodes/{pubkey}/info + /status
  • CLI: invite <pubkey> to authorize peers
  • CLI: peers to list known nodes (with name/added_at info, sorted)
  • CLI: remove <pubkey> to remove a peer
  • Iroh endpoint on startup (same Ed25519 key, mDNS + DNS discovery)
  • CLI: join <nodeid> - connects to peer, verifies invited
  • Peer verification via /nodes/{pubkey}/status check

Join Protocol (new→existing):

  • Proto: JoinRequest / JoinResponse with store UUID
  • Accept handler sends root store UUID in response
  • Join command creates empty store with received UUID (no writes until sync)

Sync Protocol (bidirectional):

  • Proto: PeerMessage wrapper with oneof for message type discrimination
  • framing.rs with MessageSink/MessageStream using LengthDelimitedCodec
  • Proto: SyncRequest/SyncResponse using SyncState
  • Store::read_entries_after(hash) to fetch log chunks
  • Accept handler: receive SyncState, compute diff, send missing entries
  • Sync command: receive entries, apply to store via apply_entry
  • CLI: sync [nodeid] command (syncs with all active peers if no nodeid)
  • After sync: node updates own /nodes/{pubkey}/info with hostname

Cleanup:

  • Move core logic from cmd_join and cmd_sync out of commands.rs (now in sync.rs)
  • Add 'invited' state: invite sets 'invited', peer sets 'active' after sync

Regressions:

  • Entry ordering: Per-author streaming is correct (hash chain per author, HLC for cross-author).
  • Multi-head sync fixed: SyncState now tracks HashSet of head hashes per author.
  • Sync entry ordering: Entries sent in HLC order (merge-sort across authors) to ensure causal order.
  • join_mesh doesn't populate node.root_store: Fixed with complete_join method.

Success Criteria

  • Node A writes, Node B syncs, both have same state
  • Works offline-first (sync when connected)

Post-M2 Refactoring:

  • Unify node.rs from lattice-cli and lattice-core
  • Move network code to lattice-net

Milestone 3: Multi-Node Mesh

Goal: N nodes form a gossip mesh for real-time sync.

Deliverables

Phase 1: LatticeServer Refactor

  • LatticeServer struct in lattice-net wrapping Arc<Node> + Endpoint
  • Move join_mesh, sync_with_peer, sync_all to LatticeServer methods
  • Encapsulate accept loop inside LatticeServer (via Router + ProtocolHandler)
  • CLI uses LatticeServer instead of raw Node + Endpoint
  • Integration test: invite → join → sync end-to-end
  • Periodic background sync with known peers
  • Track last sync time per peer

Phase 2: Gossip Protocol ✓ (iroh-gossip)

  • Router handles both lattice-sync/1 and /iroh-gossip/1 ALPNs
  • NodeEvent::RootStoreActivated emitted when root store opens
  • Auto-join gossip topic on root store activation
  • Broadcast local entries to gossip topic on commit
  • Receive gossip entries and apply to store
  • Topic ID via blake3::hash("lattice/{store_id}")
  • Gossip bootstrap peers from /peers/ (needs Prefix Watch)

Next: Prefix Watch (reactive store updates)

  • store.watch_prefix(prefix) -> Receiver<WatchEvent>
  • WatchEvent::Put { key, value } / WatchEvent::Delete { key }
  • StoreActor tracks watchers per prefix, emits on matching put/delete
  • LatticeServer uses /peers/ watch to update gossip bootstrap peers dynamically
  • Enables reactive patterns: config changes, presence, app-level subscriptions

Technical Debt

Logging

  • Replace println!/eprintln! with tracing crate (tracing::info!, tracing::error!)
  • Standard in Rust async ecosystem, used by Iroh internally

Lifecycle Management (Zombie Tasks)

  • Spawned infinite loops (spawn_node_event_listener, spawn_entry_forward_loop, gossip receive loop) keep running if LatticeServer is dropped
  • Use tokio_util::sync::CancellationToken or keep JoinHandles for graceful shutdown

Error Handling

  • Replace Result<..., String> with anyhow::Result or define LatticeNetError enum
  • String errors make it hard to handle specific failure cases

Future

  • offline nodes should not delay sync
  • sync command should transitive sync all peers
  • Gossip:
    • gossip new entries to peers
    • backfill missing entries from peers (how do peers notice missing entries?)
    • snapshots for kv store
    • prune using consensus watermark
  • remove_peer should be a transactional operation on store
  • Watermark tracking & log pruning
    • Track minimum confirmed seq per author across all peers
    • Log pruning: remove entries below watermark
  • Multi-KV-Store sync
  • Optimized sync on join. Only transfer current watermark state, then sync missing entries. This would allow pruning. Might need snapshot support in KV store.
  • Mobile (iOS/Android) clients
  • Key rotation
  • Secure storage (Keychain, TPM)
  • Snapshots for fast bootstrap
  • FUSE filesystem mount
    • Note: FUSE requires u64 inode numbers → maintain BiMap<u64, Hash> in redb
  • Merkle-ized State
    • state.db as Merkle tree with signed root hash
    • O(1) sync checks (compare root), efficient binary-search diffing
    • Light clients: fetch value + Merkle proof, verify without full state
    • Trade-off: write amplification, requires deterministic tree (Patricia Trie / Merkle Search Tree)