feat(molt): add autonomous recovery via clawdbot wake

After a successful rollback, molt now triggers the clawdbot agent to
diagnose and fix the issue. The agent receives crash logs and context,
attempts common fixes, and reports to the user if manual intervention
is needed.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
Corey H 2026-01-25 13:37:13 +00:00
parent 7a900e90f5
commit 3c0ee018f4
2 changed files with 196 additions and 50 deletions

View File

@ -93,7 +93,7 @@ clawdbot molt init
``` ```
┌─────────────────────────────────────────────────────────────────┐ ┌─────────────────────────────────────────────────────────────────┐
│ MOLT AGENT │ MOLT UPDATE FLOW
├─────────────────────────────────────────────────────────────────┤ ├─────────────────────────────────────────────────────────────────┤
│ │ │ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
@ -104,18 +104,53 @@ clawdbot molt init
│ ▼ ▼ ▼ ▼ │ │ ▼ ▼ ▼ ▼ │
│ Acquire lock git pull Health check Slack/Log │ │ Acquire lock git pull Health check Slack/Log │
│ Check remote pnpm install Module checks Changelog │ │ Check remote pnpm install Module checks Changelog │
│ Save state Restart Stability wait Fix attempt │ Save state Restart Stability wait Update history
│ │ │ │
│ │ │ │ │ │
│ ▼ (on failure) │ │ ▼ (on failure) │
│ ┌──────────────┐ │ │ ┌──────────────┐ │
│ │ Recovery │ │ │ │ Rollback │ │
│ │ (agentic) │ │ │ └──────┬───────┘ │
│ └──────────────┘ │ │ │ │
│ ┌──────────┴──────────┐ │
│ ▼ ▼ │
│ ┌─────────────┐ ┌─────────────┐ │
│ │ Rollback │ │ Rollback │ │
│ │ SUCCEEDS │ │ FAILS │ │
│ └──────┬──────┘ └──────┬──────┘ │
│ │ │ │
│ ▼ ▼ │
│ ┌─────────────┐ ┌─────────────┐ │
│ │ AUTONOMOUS │ │ RECOVERY │ │
│ │ RECOVERY │ │ .md │ │
│ │ │ │ (manual) │ │
│ │ Gateway UP │ └─────────────┘ │
│ │ Agent runs │ │
│ │ Diagnose │ │
│ │ Fix & retry │ │
│ └─────────────┘ │
│ │ │ │
└─────────────────────────────────────────────────────────────────┘ └─────────────────────────────────────────────────────────────────┘
``` ```
**Autonomous Recovery Flow:**
```
Rollback succeeds
Gateway is UP (old version)
clawdbot wake --text "diagnose and fix"
Agent reads crash log
├──▶ Fixable? ──▶ Apply fix ──▶ Retry molt.sh ──▶ Success!
└──▶ Not fixable? ──▶ Report to user via Slack
```
## Phases ## Phases
### Phase 0: Preflight ### Phase 0: Preflight
@ -262,65 +297,97 @@ case $OUTCOME in
esac esac
``` ```
### Recovery (Agentic) ### Recovery (Agentic & Autonomous)
When verification fails, Molt doesn't just blindly rollback. It: When verification fails, Molt doesn't just rollback and give up. It leverages the fact that **after a successful rollback, Clawdbot is running again** — so it can use Clawdbot's own agent system to diagnose and fix the issue.
1. **Captures context** — Gateway logs, error messages, what failed **The key insight:** Rollback restores the old (working) version → Gateway is up → Agent can run → Agent diagnoses and fixes → Retry update.
2. **Attempts simple rollback**`git checkout $OLD_HEAD && pnpm install && restart`
3. **If rollback fails, escalates** — Provides context for the AI or human to fix ```
Update fails → Rollback succeeds → Gateway UP → Trigger agent → Diagnose & fix → Retry
```
**Recovery flow:**
1. **Capture context** — Crash logs, error messages, failed commit
2. **Rollback**`git checkout $OLD_HEAD && pnpm install && restart`
3. **If rollback succeeds** — Trigger autonomous recovery agent via `clawdbot wake`
4. **If rollback fails** — Write RECOVERY.md for manual intervention
**Autonomous recovery agent prompt:**
```bash ```bash
recover() { # After successful rollback, wake the agent to diagnose and fix
# Capture what went wrong clawdbot wake --mode now --text "$(cat <<'EOF'
journalctl --user -u clawdbot-gateway -n 100 > ~/.clawdbot/molt/crash-log.txt 🦞 MOLT AUTONOMOUS RECOVERY
# Try simple rollback The nightly update failed, but rollback succeeded. I'm running on the old version.
git checkout $(cat ~/.clawdbot/molt/pre-update-head)
pnpm install --frozen-lockfile --prefer-offline
clawdbot daemon restart
# Verify rollback worked ## Your Mission
sleep 10 1. Diagnose what went wrong
if clawdbot ping --timeout 10; then 2. If fixable, fix it and retry the update
notify "Rolled back successfully. See crash log for details." 3. If not fixable, report findings to the user
return 0
fi
# Rollback failed - this needs human/AI intervention ## Context
notify "CRITICAL: Rollback failed. Gateway is down. Manual fix required." - Old HEAD (current): ${CURRENT_HEAD}
notify "Crash log: ~/.clawdbot/molt/crash-log.txt" - Failed HEAD: ${NEW_HEAD}
notify "Pre-update HEAD: $(cat ~/.clawdbot/molt/pre-update-head)" - Crash log: ~/.clawdbot/molt/crash-log.txt
# Write instructions for the next agent/human ## Steps
cat > ~/.clawdbot/molt/RECOVERY.md << 'EOF' 1. Read the crash log: cat ~/.clawdbot/molt/crash-log.txt
2. Identify the error (common causes below)
3. If you can fix it:
- Apply the fix
- Run: ~/.clawdbot/molt/molt.sh
- If it succeeds, we're done!
4. If you can't fix it:
- Explain what went wrong
- Message the user via Slack with your findings
## Common Fixable Issues
- "Cannot find module X" → Try: cd ~/clawd && pnpm install --force
- "ENOSPC" (disk full) → Try: pnpm store prune && pnpm cache clean
- Network timeout during install → Just retry: ~/.clawdbot/molt/molt.sh
- Lockfile conflict → Try: rm pnpm-lock.yaml && git checkout pnpm-lock.yaml && pnpm install
## Important
- You have 1 retry attempt. If molt.sh fails again, report to the user.
- Don't get stuck in a loop - if unsure, ask for help.
EOF
)"
```
**Why this works:**
- Uses existing Clawdbot infrastructure (no new agent framework)
- Agent has full access to bash, file reading, etc.
- Agent can iterate: try fix → retry molt → verify
- Falls back to human if truly stuck
**Fallback (rollback fails):**
If rollback itself fails, Clawdbot is down and can't help. In this case, Molt writes a `RECOVERY.md` file with:
- What happened
- Crash log location
- Manual recovery steps
- Context for an external AI (like Claude Code via SSH) to fix
```markdown
# Molt Recovery Required # Molt Recovery Required
The nightly update failed and automatic rollback also failed. The nightly update failed and automatic rollback also failed.
## What happened ## What happened
- Update started at: $TIMESTAMP - Old HEAD: abc123
- Old HEAD: $OLD_HEAD - Failed HEAD: def456
- New HEAD: $NEW_HEAD (attempted) - Error: Gateway didn't start
- Error: $ERROR
## Crash log ## Crash log
See: ~/.clawdbot/molt/crash-log.txt ~/.clawdbot/molt/crash-log.txt
## Manual recovery steps ## Manual recovery
1. Check the crash log for the root cause 1. SSH into the server
2. Try: `cd ~/clawd && git checkout $OLD_HEAD && pnpm install && clawdbot daemon restart` 2. cd ~/clawd && git checkout abc123 && pnpm install && pnpm build
3. If that fails, see CLAUDE.md for nuclear options 3. systemctl --user restart clawdbot-gateway
## Context for AI recovery
The gateway failed to start after update. Common causes:
- Missing dependency (check for "Cannot find module" in logs)
- Syntax error in new code (check for "SyntaxError" in logs)
- Config incompatibility (check for "Invalid config" in logs)
EOF
return 1
}
``` ```
## Configuration ## Configuration

View File

@ -211,6 +211,9 @@ if ! $gateway_up; then
if systemctl --user is-active --quiet clawdbot-gateway.service; then if systemctl --user is-active --quiet clawdbot-gateway.service; then
log "Rollback successful - gateway is running on ${CURRENT_HEAD:0:8}" log "Rollback successful - gateway is running on ${CURRENT_HEAD:0:8}"
# Save the failed HEAD for the agent to analyze
echo "$NEW_HEAD" > "${MOLT_DIR}/attempted-head"
# Update changelog to reflect rollback # Update changelog to reflect rollback
cat >> "${WORKSPACE_DIR}/update-changelog.md" << EOF cat >> "${WORKSPACE_DIR}/update-changelog.md" << EOF
@ -220,7 +223,54 @@ Gateway didn't come up after update.
See crash log: ~/.clawdbot/molt/crash-log.txt See crash log: ~/.clawdbot/molt/crash-log.txt
EOF EOF
die "Update failed, rolled back successfully" # Log rollback to history
echo "{\"timestamp\":\"$(timestamp)\",\"from\":\"${CURRENT_HEAD}\",\"to\":\"${NEW_HEAD}\",\"commits\":${COMMIT_COUNT},\"status\":\"rollback\",\"reason\":\"health_check_failed\"}" >> "${MOLT_DIR}/history.jsonl"
# === AUTONOMOUS RECOVERY ===
# Gateway is back up on old version - trigger agent to diagnose and fix
log "Triggering autonomous recovery agent..."
cd "$CLAWDBOT_DIR"
node dist/entry.js wake --mode now --text "$(cat <<AGENT_PROMPT
🦞 MOLT AUTONOMOUS RECOVERY
The nightly update failed, but rollback succeeded. I'm running on the old version now.
## Your Mission
1. Diagnose what went wrong
2. If fixable, fix it and retry the update
3. If not fixable, report findings to Corey
## Context
- Old HEAD (current): ${CURRENT_HEAD:0:8}
- Failed HEAD: ${NEW_HEAD:0:8}
- Commits attempted: ${COMMIT_COUNT}
- Crash log: ~/.clawdbot/molt/crash-log.txt
## Steps
1. Read the crash log: cat ~/.clawdbot/molt/crash-log.txt
2. Identify the error (common causes below)
3. If you can fix it:
- Apply the fix
- Run: ~/.clawdbot/molt/molt.sh
- If it succeeds, we're done!
4. If you can't fix it:
- Explain what went wrong
- Message Corey via Slack with your findings
## Common Fixable Issues
- "Cannot find module X" → Try: cd ~/clawd && pnpm install --force
- "ENOSPC" (disk full) → Try: pnpm store prune && pnpm cache clean
- Network timeout during install → Just retry: ~/.clawdbot/molt/molt.sh
- Lockfile conflict → Try: cd ~/clawd && rm pnpm-lock.yaml && git checkout pnpm-lock.yaml && pnpm install
## Important
- You have 1 retry attempt. If molt.sh fails again, report to Corey.
- Don't get stuck in a loop - if unsure, ask for help.
AGENT_PROMPT
)" 2>&1 || log "Warning: Could not trigger recovery agent"
exit 1 # Exit with error so cron knows it failed
else else
log "CRITICAL: Rollback also failed!" log "CRITICAL: Rollback also failed!"
@ -270,6 +320,9 @@ if ! systemctl --user is-active --quiet clawdbot-gateway.service; then
pnpm build 2>&1 || true pnpm build 2>&1 || true
systemctl --user restart clawdbot-gateway.service systemctl --user restart clawdbot-gateway.service
# Save the failed HEAD
echo "$NEW_HEAD" > "${MOLT_DIR}/attempted-head"
cat >> "${WORKSPACE_DIR}/update-changelog.md" << EOF cat >> "${WORKSPACE_DIR}/update-changelog.md" << EOF
## ROLLBACK ## ROLLBACK
@ -277,7 +330,33 @@ Update failed - crashed during stability window. Rolled back to \`${CURRENT_HEAD
See crash log: ~/.clawdbot/molt/crash-log.txt See crash log: ~/.clawdbot/molt/crash-log.txt
EOF EOF
die "Gateway crashed during stability window, rolled back" # Log rollback
echo "{\"timestamp\":\"$(timestamp)\",\"from\":\"${CURRENT_HEAD}\",\"to\":\"${NEW_HEAD}\",\"commits\":${COMMIT_COUNT},\"status\":\"rollback\",\"reason\":\"stability_window_crash\"}" >> "${MOLT_DIR}/history.jsonl"
# Trigger autonomous recovery
log "Triggering autonomous recovery agent..."
sleep 5 # Give gateway a moment to stabilize
cd "$CLAWDBOT_DIR"
node dist/entry.js wake --mode now --text "$(cat <<AGENT_PROMPT
🦞 MOLT AUTONOMOUS RECOVERY
Update crashed during stability window. Rolled back successfully.
## Context
- Old HEAD (current): ${CURRENT_HEAD:0:8}
- Failed HEAD: ${NEW_HEAD:0:8}
- Crash log: ~/.clawdbot/molt/crash-log.txt
## Steps
1. Read crash log: cat ~/.clawdbot/molt/crash-log.txt
2. Diagnose the crash (likely a runtime error, not build error)
3. If fixable, fix and retry: ~/.clawdbot/molt/molt.sh
4. If not, report to Corey via Slack
AGENT_PROMPT
)" 2>&1 || log "Warning: Could not trigger recovery agent"
exit 1
fi fi
# Final health check # Final health check