feat(molt): add autonomous recovery via clawdbot wake
After a successful rollback, molt now triggers the clawdbot agent to diagnose and fix the issue. The agent receives crash logs and context, attempts common fixes, and reports to the user if manual intervention is needed. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
parent
7a900e90f5
commit
3c0ee018f4
163
docs/molt.md
163
docs/molt.md
@ -93,7 +93,7 @@ clawdbot molt init
|
|||||||
|
|
||||||
```
|
```
|
||||||
┌─────────────────────────────────────────────────────────────────┐
|
┌─────────────────────────────────────────────────────────────────┐
|
||||||
│ MOLT AGENT │
|
│ MOLT UPDATE FLOW │
|
||||||
├─────────────────────────────────────────────────────────────────┤
|
├─────────────────────────────────────────────────────────────────┤
|
||||||
│ │
|
│ │
|
||||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||||
@ -104,18 +104,53 @@ clawdbot molt init
|
|||||||
│ ▼ ▼ ▼ ▼ │
|
│ ▼ ▼ ▼ ▼ │
|
||||||
│ Acquire lock git pull Health check Slack/Log │
|
│ Acquire lock git pull Health check Slack/Log │
|
||||||
│ Check remote pnpm install Module checks Changelog │
|
│ Check remote pnpm install Module checks Changelog │
|
||||||
│ Save state Restart Stability wait Fix attempt │
|
│ Save state Restart Stability wait Update history │
|
||||||
│ │
|
│ │
|
||||||
│ │ │
|
│ │ │
|
||||||
│ ▼ (on failure) │
|
│ ▼ (on failure) │
|
||||||
│ ┌──────────────┐ │
|
│ ┌──────────────┐ │
|
||||||
│ │ Recovery │ │
|
│ │ Rollback │ │
|
||||||
│ │ (agentic) │ │
|
│ └──────┬───────┘ │
|
||||||
│ └──────────────┘ │
|
│ │ │
|
||||||
|
│ ┌──────────┴──────────┐ │
|
||||||
|
│ ▼ ▼ │
|
||||||
|
│ ┌─────────────┐ ┌─────────────┐ │
|
||||||
|
│ │ Rollback │ │ Rollback │ │
|
||||||
|
│ │ SUCCEEDS │ │ FAILS │ │
|
||||||
|
│ └──────┬──────┘ └──────┬──────┘ │
|
||||||
|
│ │ │ │
|
||||||
|
│ ▼ ▼ │
|
||||||
|
│ ┌─────────────┐ ┌─────────────┐ │
|
||||||
|
│ │ AUTONOMOUS │ │ RECOVERY │ │
|
||||||
|
│ │ RECOVERY │ │ .md │ │
|
||||||
|
│ │ │ │ (manual) │ │
|
||||||
|
│ │ Gateway UP │ └─────────────┘ │
|
||||||
|
│ │ Agent runs │ │
|
||||||
|
│ │ Diagnose │ │
|
||||||
|
│ │ Fix & retry │ │
|
||||||
|
│ └─────────────┘ │
|
||||||
│ │
|
│ │
|
||||||
└─────────────────────────────────────────────────────────────────┘
|
└─────────────────────────────────────────────────────────────────┘
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**Autonomous Recovery Flow:**
|
||||||
|
```
|
||||||
|
Rollback succeeds
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
Gateway is UP (old version)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
clawdbot wake --text "diagnose and fix"
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
Agent reads crash log
|
||||||
|
│
|
||||||
|
├──▶ Fixable? ──▶ Apply fix ──▶ Retry molt.sh ──▶ Success!
|
||||||
|
│
|
||||||
|
└──▶ Not fixable? ──▶ Report to user via Slack
|
||||||
|
```
|
||||||
|
|
||||||
## Phases
|
## Phases
|
||||||
|
|
||||||
### Phase 0: Preflight
|
### Phase 0: Preflight
|
||||||
@ -262,65 +297,97 @@ case $OUTCOME in
|
|||||||
esac
|
esac
|
||||||
```
|
```
|
||||||
|
|
||||||
### Recovery (Agentic)
|
### Recovery (Agentic & Autonomous)
|
||||||
|
|
||||||
When verification fails, Molt doesn't just blindly rollback. It:
|
When verification fails, Molt doesn't just rollback and give up. It leverages the fact that **after a successful rollback, Clawdbot is running again** — so it can use Clawdbot's own agent system to diagnose and fix the issue.
|
||||||
|
|
||||||
1. **Captures context** — Gateway logs, error messages, what failed
|
**The key insight:** Rollback restores the old (working) version → Gateway is up → Agent can run → Agent diagnoses and fixes → Retry update.
|
||||||
2. **Attempts simple rollback** — `git checkout $OLD_HEAD && pnpm install && restart`
|
|
||||||
3. **If rollback fails, escalates** — Provides context for the AI or human to fix
|
```
|
||||||
|
Update fails → Rollback succeeds → Gateway UP → Trigger agent → Diagnose & fix → Retry
|
||||||
|
```
|
||||||
|
|
||||||
|
**Recovery flow:**
|
||||||
|
|
||||||
|
1. **Capture context** — Crash logs, error messages, failed commit
|
||||||
|
2. **Rollback** — `git checkout $OLD_HEAD && pnpm install && restart`
|
||||||
|
3. **If rollback succeeds** — Trigger autonomous recovery agent via `clawdbot wake`
|
||||||
|
4. **If rollback fails** — Write RECOVERY.md for manual intervention
|
||||||
|
|
||||||
|
**Autonomous recovery agent prompt:**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
recover() {
|
# After successful rollback, wake the agent to diagnose and fix
|
||||||
# Capture what went wrong
|
clawdbot wake --mode now --text "$(cat <<'EOF'
|
||||||
journalctl --user -u clawdbot-gateway -n 100 > ~/.clawdbot/molt/crash-log.txt
|
🦞 MOLT AUTONOMOUS RECOVERY
|
||||||
|
|
||||||
# Try simple rollback
|
The nightly update failed, but rollback succeeded. I'm running on the old version.
|
||||||
git checkout $(cat ~/.clawdbot/molt/pre-update-head)
|
|
||||||
pnpm install --frozen-lockfile --prefer-offline
|
|
||||||
clawdbot daemon restart
|
|
||||||
|
|
||||||
# Verify rollback worked
|
## Your Mission
|
||||||
sleep 10
|
1. Diagnose what went wrong
|
||||||
if clawdbot ping --timeout 10; then
|
2. If fixable, fix it and retry the update
|
||||||
notify "Rolled back successfully. See crash log for details."
|
3. If not fixable, report findings to the user
|
||||||
return 0
|
|
||||||
fi
|
|
||||||
|
|
||||||
# Rollback failed - this needs human/AI intervention
|
## Context
|
||||||
notify "CRITICAL: Rollback failed. Gateway is down. Manual fix required."
|
- Old HEAD (current): ${CURRENT_HEAD}
|
||||||
notify "Crash log: ~/.clawdbot/molt/crash-log.txt"
|
- Failed HEAD: ${NEW_HEAD}
|
||||||
notify "Pre-update HEAD: $(cat ~/.clawdbot/molt/pre-update-head)"
|
- Crash log: ~/.clawdbot/molt/crash-log.txt
|
||||||
|
|
||||||
# Write instructions for the next agent/human
|
## Steps
|
||||||
cat > ~/.clawdbot/molt/RECOVERY.md << 'EOF'
|
1. Read the crash log: cat ~/.clawdbot/molt/crash-log.txt
|
||||||
|
2. Identify the error (common causes below)
|
||||||
|
3. If you can fix it:
|
||||||
|
- Apply the fix
|
||||||
|
- Run: ~/.clawdbot/molt/molt.sh
|
||||||
|
- If it succeeds, we're done!
|
||||||
|
4. If you can't fix it:
|
||||||
|
- Explain what went wrong
|
||||||
|
- Message the user via Slack with your findings
|
||||||
|
|
||||||
|
## Common Fixable Issues
|
||||||
|
- "Cannot find module X" → Try: cd ~/clawd && pnpm install --force
|
||||||
|
- "ENOSPC" (disk full) → Try: pnpm store prune && pnpm cache clean
|
||||||
|
- Network timeout during install → Just retry: ~/.clawdbot/molt/molt.sh
|
||||||
|
- Lockfile conflict → Try: rm pnpm-lock.yaml && git checkout pnpm-lock.yaml && pnpm install
|
||||||
|
|
||||||
|
## Important
|
||||||
|
- You have 1 retry attempt. If molt.sh fails again, report to the user.
|
||||||
|
- Don't get stuck in a loop - if unsure, ask for help.
|
||||||
|
EOF
|
||||||
|
)"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why this works:**
|
||||||
|
- Uses existing Clawdbot infrastructure (no new agent framework)
|
||||||
|
- Agent has full access to bash, file reading, etc.
|
||||||
|
- Agent can iterate: try fix → retry molt → verify
|
||||||
|
- Falls back to human if truly stuck
|
||||||
|
|
||||||
|
**Fallback (rollback fails):**
|
||||||
|
|
||||||
|
If rollback itself fails, Clawdbot is down and can't help. In this case, Molt writes a `RECOVERY.md` file with:
|
||||||
|
- What happened
|
||||||
|
- Crash log location
|
||||||
|
- Manual recovery steps
|
||||||
|
- Context for an external AI (like Claude Code via SSH) to fix
|
||||||
|
|
||||||
|
```markdown
|
||||||
# Molt Recovery Required
|
# Molt Recovery Required
|
||||||
|
|
||||||
The nightly update failed and automatic rollback also failed.
|
The nightly update failed and automatic rollback also failed.
|
||||||
|
|
||||||
## What happened
|
## What happened
|
||||||
- Update started at: $TIMESTAMP
|
- Old HEAD: abc123
|
||||||
- Old HEAD: $OLD_HEAD
|
- Failed HEAD: def456
|
||||||
- New HEAD: $NEW_HEAD (attempted)
|
- Error: Gateway didn't start
|
||||||
- Error: $ERROR
|
|
||||||
|
|
||||||
## Crash log
|
## Crash log
|
||||||
See: ~/.clawdbot/molt/crash-log.txt
|
~/.clawdbot/molt/crash-log.txt
|
||||||
|
|
||||||
## Manual recovery steps
|
## Manual recovery
|
||||||
1. Check the crash log for the root cause
|
1. SSH into the server
|
||||||
2. Try: `cd ~/clawd && git checkout $OLD_HEAD && pnpm install && clawdbot daemon restart`
|
2. cd ~/clawd && git checkout abc123 && pnpm install && pnpm build
|
||||||
3. If that fails, see CLAUDE.md for nuclear options
|
3. systemctl --user restart clawdbot-gateway
|
||||||
|
|
||||||
## Context for AI recovery
|
|
||||||
The gateway failed to start after update. Common causes:
|
|
||||||
- Missing dependency (check for "Cannot find module" in logs)
|
|
||||||
- Syntax error in new code (check for "SyntaxError" in logs)
|
|
||||||
- Config incompatibility (check for "Invalid config" in logs)
|
|
||||||
EOF
|
|
||||||
|
|
||||||
return 1
|
|
||||||
}
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Configuration
|
## Configuration
|
||||||
|
|||||||
@ -211,6 +211,9 @@ if ! $gateway_up; then
|
|||||||
if systemctl --user is-active --quiet clawdbot-gateway.service; then
|
if systemctl --user is-active --quiet clawdbot-gateway.service; then
|
||||||
log "Rollback successful - gateway is running on ${CURRENT_HEAD:0:8}"
|
log "Rollback successful - gateway is running on ${CURRENT_HEAD:0:8}"
|
||||||
|
|
||||||
|
# Save the failed HEAD for the agent to analyze
|
||||||
|
echo "$NEW_HEAD" > "${MOLT_DIR}/attempted-head"
|
||||||
|
|
||||||
# Update changelog to reflect rollback
|
# Update changelog to reflect rollback
|
||||||
cat >> "${WORKSPACE_DIR}/update-changelog.md" << EOF
|
cat >> "${WORKSPACE_DIR}/update-changelog.md" << EOF
|
||||||
|
|
||||||
@ -220,7 +223,54 @@ Gateway didn't come up after update.
|
|||||||
See crash log: ~/.clawdbot/molt/crash-log.txt
|
See crash log: ~/.clawdbot/molt/crash-log.txt
|
||||||
EOF
|
EOF
|
||||||
|
|
||||||
die "Update failed, rolled back successfully"
|
# Log rollback to history
|
||||||
|
echo "{\"timestamp\":\"$(timestamp)\",\"from\":\"${CURRENT_HEAD}\",\"to\":\"${NEW_HEAD}\",\"commits\":${COMMIT_COUNT},\"status\":\"rollback\",\"reason\":\"health_check_failed\"}" >> "${MOLT_DIR}/history.jsonl"
|
||||||
|
|
||||||
|
# === AUTONOMOUS RECOVERY ===
|
||||||
|
# Gateway is back up on old version - trigger agent to diagnose and fix
|
||||||
|
log "Triggering autonomous recovery agent..."
|
||||||
|
|
||||||
|
cd "$CLAWDBOT_DIR"
|
||||||
|
node dist/entry.js wake --mode now --text "$(cat <<AGENT_PROMPT
|
||||||
|
🦞 MOLT AUTONOMOUS RECOVERY
|
||||||
|
|
||||||
|
The nightly update failed, but rollback succeeded. I'm running on the old version now.
|
||||||
|
|
||||||
|
## Your Mission
|
||||||
|
1. Diagnose what went wrong
|
||||||
|
2. If fixable, fix it and retry the update
|
||||||
|
3. If not fixable, report findings to Corey
|
||||||
|
|
||||||
|
## Context
|
||||||
|
- Old HEAD (current): ${CURRENT_HEAD:0:8}
|
||||||
|
- Failed HEAD: ${NEW_HEAD:0:8}
|
||||||
|
- Commits attempted: ${COMMIT_COUNT}
|
||||||
|
- Crash log: ~/.clawdbot/molt/crash-log.txt
|
||||||
|
|
||||||
|
## Steps
|
||||||
|
1. Read the crash log: cat ~/.clawdbot/molt/crash-log.txt
|
||||||
|
2. Identify the error (common causes below)
|
||||||
|
3. If you can fix it:
|
||||||
|
- Apply the fix
|
||||||
|
- Run: ~/.clawdbot/molt/molt.sh
|
||||||
|
- If it succeeds, we're done!
|
||||||
|
4. If you can't fix it:
|
||||||
|
- Explain what went wrong
|
||||||
|
- Message Corey via Slack with your findings
|
||||||
|
|
||||||
|
## Common Fixable Issues
|
||||||
|
- "Cannot find module X" → Try: cd ~/clawd && pnpm install --force
|
||||||
|
- "ENOSPC" (disk full) → Try: pnpm store prune && pnpm cache clean
|
||||||
|
- Network timeout during install → Just retry: ~/.clawdbot/molt/molt.sh
|
||||||
|
- Lockfile conflict → Try: cd ~/clawd && rm pnpm-lock.yaml && git checkout pnpm-lock.yaml && pnpm install
|
||||||
|
|
||||||
|
## Important
|
||||||
|
- You have 1 retry attempt. If molt.sh fails again, report to Corey.
|
||||||
|
- Don't get stuck in a loop - if unsure, ask for help.
|
||||||
|
AGENT_PROMPT
|
||||||
|
)" 2>&1 || log "Warning: Could not trigger recovery agent"
|
||||||
|
|
||||||
|
exit 1 # Exit with error so cron knows it failed
|
||||||
else
|
else
|
||||||
log "CRITICAL: Rollback also failed!"
|
log "CRITICAL: Rollback also failed!"
|
||||||
|
|
||||||
@ -270,6 +320,9 @@ if ! systemctl --user is-active --quiet clawdbot-gateway.service; then
|
|||||||
pnpm build 2>&1 || true
|
pnpm build 2>&1 || true
|
||||||
systemctl --user restart clawdbot-gateway.service
|
systemctl --user restart clawdbot-gateway.service
|
||||||
|
|
||||||
|
# Save the failed HEAD
|
||||||
|
echo "$NEW_HEAD" > "${MOLT_DIR}/attempted-head"
|
||||||
|
|
||||||
cat >> "${WORKSPACE_DIR}/update-changelog.md" << EOF
|
cat >> "${WORKSPACE_DIR}/update-changelog.md" << EOF
|
||||||
|
|
||||||
## ROLLBACK
|
## ROLLBACK
|
||||||
@ -277,7 +330,33 @@ Update failed - crashed during stability window. Rolled back to \`${CURRENT_HEAD
|
|||||||
See crash log: ~/.clawdbot/molt/crash-log.txt
|
See crash log: ~/.clawdbot/molt/crash-log.txt
|
||||||
EOF
|
EOF
|
||||||
|
|
||||||
die "Gateway crashed during stability window, rolled back"
|
# Log rollback
|
||||||
|
echo "{\"timestamp\":\"$(timestamp)\",\"from\":\"${CURRENT_HEAD}\",\"to\":\"${NEW_HEAD}\",\"commits\":${COMMIT_COUNT},\"status\":\"rollback\",\"reason\":\"stability_window_crash\"}" >> "${MOLT_DIR}/history.jsonl"
|
||||||
|
|
||||||
|
# Trigger autonomous recovery
|
||||||
|
log "Triggering autonomous recovery agent..."
|
||||||
|
sleep 5 # Give gateway a moment to stabilize
|
||||||
|
|
||||||
|
cd "$CLAWDBOT_DIR"
|
||||||
|
node dist/entry.js wake --mode now --text "$(cat <<AGENT_PROMPT
|
||||||
|
🦞 MOLT AUTONOMOUS RECOVERY
|
||||||
|
|
||||||
|
Update crashed during stability window. Rolled back successfully.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
- Old HEAD (current): ${CURRENT_HEAD:0:8}
|
||||||
|
- Failed HEAD: ${NEW_HEAD:0:8}
|
||||||
|
- Crash log: ~/.clawdbot/molt/crash-log.txt
|
||||||
|
|
||||||
|
## Steps
|
||||||
|
1. Read crash log: cat ~/.clawdbot/molt/crash-log.txt
|
||||||
|
2. Diagnose the crash (likely a runtime error, not build error)
|
||||||
|
3. If fixable, fix and retry: ~/.clawdbot/molt/molt.sh
|
||||||
|
4. If not, report to Corey via Slack
|
||||||
|
AGENT_PROMPT
|
||||||
|
)" 2>&1 || log "Warning: Could not trigger recovery agent"
|
||||||
|
|
||||||
|
exit 1
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# Final health check
|
# Final health check
|
||||||
|
|||||||
Loading…
Reference in New Issue
Block a user