I run AI models and coding agents on my own hardware and write down what I measure: what worked, what failed, and how to check it yourself.
- NousResearch/hermes-agent: three commits in
mainfix(tirith): half-open the circuit breaker instead of latching it openfix(cron): redact secrets from delivery content before sendingtest(cron): cover every outward cron delivery lane- The maintainers cherry-picked these changes into other pull requests that were merged into
main, keeping my authorship; my original pull requests were closed without merging.
- openclaw/openclaw: #116558 fix: gateway wedges on every startup when legacy runtime-state files conflict with SQLite
- garrytan/gbrain: #4536 fix(files): give the transcribable audio formats a MIME_TYPES entry
- axiom: checks an AI agent's "done" claims against evidence declared before the work.
- ryanai-evalbank: a method for scoring a model on your own work without publishing the questions.
- ryanai-lab-cookbooks: an index of 54 RyanAI Lab cookbooks on serving, training, evaluation and coding agents.
