migrated from PAI to LIFEOS - and perceive a SUBSTANTIAL reduction in result quality! #1715
Replies: 9 comments 12 replies
|
I too upgraded v5 -> v6 -> lifeos and have ultimately negative sentiment towards dropping the algorithm ceremony during sessions. Given the non-deterministic nature of the skill-based scaffolding, as you pointed out, there were times previously during the PAI era where the full algorithm would not run despite specifying a task that would require it. At least with the ceremony in PAI, you could easily see when it fired, and when it didn't. You could poke your DA to look into why it didn't fire and prompt it to fire. However, now there is nothing at all, even the statusline indicator got removed. So, again noting the non-deterministic nature of skill-based routing, we are in the dark in every session as to whether the algorithm ran or not. Would be nice to see the ceremony come back while AI context routing is less than 100% deterministic. |
|
I like v5 and have it on one of my systems (from v0.9->v2->v5), but v7 has been absolutely amazing on a fresh install on a clean macOS userspace. I used Fable to port things over, jack into the memory system with his ultracode prompt, and now run Opus 5 as my daily driver with Fable for big projects. The efficiency is real for my work, and the results have been so astoundingly good I am shipping great little projects related to my work and hobbies all the time. I did have one abortive attempt to go v5->v6 (a fresh macOS user with my working v5 copied over); after reading about the BP Lesson engineering of v7 I decided that a fresh install and pulling from v5 rather than pushing the other way was best. |
|
The feel has definitely changed. Is the main issue that it’s no longer adhering to the algorithm as closely as before? If so, I’d try two things: If you still have a backup of your old .claude directory from before the upgrade, run it in a sandbox and test a few of the same workflow tasks on both the old and new installations. That should help confirm whether there’s a meaningful difference. Have your AI compare the responses or builds produced by each installation and identify where the differences are. It may be that context, hooks, or other config are being loaded differently in the new version. With the last few releases, I’ve downloaded and tested them, then selectively imported the parts I liked into my existing setup. The Bitter Pill approach has been phenomenal, but there’s also something valuable about maintaining your own harness—one you’ve used, customized, and refined over time to fit your specific needs. Even an out-of-the-box release from the same project won’t necessarily preserve all the small tweaks you’ve made along the way. That’s also one of the great things about the AI era: we could all use the same “LifeOS” or agentic framework, yet each implementation can still be completely unique. |
|
Taking this seriously but it needs specifics to act on: a concrete before/after pair (same prompt, old PAI vs current LifeOS, with the output you got) would let us bisect whether it's the Algorithm changes, the memory system, or prompt-surface size. If you can share even one reproducible example, file it as an issue and it gets the full treatment. |
|
The sandbox A/B upthread is the right move, with one addition that changes the conclusion you'll draw: run each test prompt several times per install, not once. LLM output is stochastic enough that a single before/after pair mostly measures the dice. What worked for us: pick 10–20 real prompts from past sessions with a checkable expectation each (did it follow the algorithm step, did it cite the right file, did the output parse), run each 3–5 times on both installs, and compare pass-rates with the spread, not single outputs. That turns "the feel has changed" into a number the maintainer can bisect against, and it answers the question this thread actually raises — which scaffolding earns its context cost. We keep a small golden set standing for exactly this, so any scaffold change (instructions, memory shape, prompt-surface size) has to beat the baseline or get cut. Documented here: https://github.com/jimy-r/agent-workspace-architecture/blob/main/PATTERNS.md#11-a-scaffold-is-a-hypothesis--gate-it-behind-a-measurable-signal |
|
First, the usual caveats. I have mega respect for this repo and this project and the whole philosophy behind it. I use a stripped-down version of this at work that I call Work OS, and it is mostly the algorithm and the thinking skills applied. This project is how I learned how to do AI work, and it kicked off an entire thing where I work at. Like, this was the project that caused my work to go, Oh crap, AI can do things now. So this was back in PAI 3.0 and I've been using it ever since. Massive appreciation for both the project and the philosophy behind it. I find that I'm hitting the same problem. Ever since the Bitter Pill engineering pass and Opus 5, I almost feel like some of that jumped the gun on model capability. When I look at what my DA does today, they almost never call the algorithm, they almost never call their thinking skills unless I make them do it. The whole observe part of this seems to have fallen by the wayside. He will confidently launch in a direction that could have been changed by simply observing the ground truth area before continuing. The algo nudge hook, in my DA's own words, fired 30-something times With zero execution. So it is seeing that it should be calling thinking skills, but almost always determining that it doesn't need any additional help in its thinking and that its current thinking is fine and would not benefit from calling these skills. The funny part is, even when I force it to happen, it's almost like he's reluctantly insulted by the fact that it worked. The whole euphoric surprise component of this is almost entirely gone. I've had multiple moments where I thought to myself, "There is no way that this is the DA that Daniel is running." This is causing some real contention because I find that LifeOS is great to talk to, but not very good at actually getting stuff done. LifeOS is very good at finding wisdom in a discussion topic. But when it comes to doing something like, hey, lets figure out how to dig through my email and look for stuff Im supposed to do, not so much. When I say something like that, what I'm expecting the DA to do is to come up with a broad set of categories of things that can happen in email, to look at my life and how it all works, and to look at all the surfaces We have, and then figure out based on email where it slots into each particular service. We've got finances, we've got bills, these things are different spots, they need to have ways to feed into each other, whatnot. But I don't get that at all. What I get is, I'll apply category labels to it. It's like the minimum possible thing to do with email that's not part of the big grand plan. And even when I describe the big grand plan, I'm expecting him to fill in the blanks. I can describe entire classes of things that can happen, and I know that the model is capable of taking that class of stuff and expanding it out to something that's observable, these observable buckets, but it just doesn't do it. And I think this comes down to the algorithm trim, the BPE pass, and the fact that it can pass on these thinking skills. He seems to get stuck in a build and react loop that has no wisdom applied to the entire thing that he's working on. as a result, He's wrong far more often than he's right. And the number one thing my DA seems to say to me nowadays is, "That's fair, and you are right." This slowly turns using the DA into an uphill dragging slog Where I feel like I am teaching it how to do the things that it used to do automatically on its own. Now I don't mind teaching a work DA how my various processes work so they can go into a skill. I want things to be very prescribed there. And while I am looking for ways to improve those processes by injecting AI into them, the problem is that I can't seem to trust this version of LifeOS to produce the type of insights that I used to get. It's almost like it has tunnel vision and a sense for giving out the minimal effort. I can't quite put my finger on it, but it feels like 7.0 was built to be used by a model like Fable and maybe not built to be used by a model like Opus. And in a way, this makes sense. If you're not running Fable as your main model and you ran all of the forever prompts through Fable, then the suggestions it would give would be at a level of Fable thinking Unless it was told to account for that. I think folks do this kind of thing all the time, even though they'll realize they're doing it. You might use Fable to put together an excellent plan and have Opus decompose that plan into sonnet level tasks. Then you run an orchestrator to go out there and do all the work at the sonnet level so you get maximum token efficiency. But that type of way of working has to be built into the workflow of how that plan is put together, executed, and eventually decomposed into those tasks. If that wasn't done quite during the BPE pass, then it's possible that Opus is sitting a little higher up the chain than it ought to be as far as confidence in its own capabilities. aka Why would I pull in systems thinking when I feel I can do that on my own? I totally get that, but at the end of the day, he is constantly wrong and constantly missing these edge cases, the kinds of things that I feel previous versions of the DA would have caught, considering how all the pieces touch other pieces, using all the surfaces that are available. He just doesn't do that now unless I explicitly tell him to. And knowing the goal of this project making this sort of great for everyone then I don't expect most folks to go down that path. And if they do then the project is probably not serving its purpose because it's forcing folks to drag this thing uphill rather than to help take some of the weight off. What I find that I end up doing now instead, which I don't like, but I do need things to move, is I have Hermes do the work. Hermes is very good at just getting stuff done. And I had Hermes go through the Life OS project and borrow all of the good ideas and try to build them in in a Hermes shape. But unfortunately you cant use Anthropic subscription there. So I dont. I have to use GPT 5.6 Soul, and it just isn't the same. I suspect that this is at least partially why we had sub-agents like Forge, because you've got these open AI models that are trained way more to do the people-pleasing things, so they want to get work done just sort of innate into it. So you probably are doing this thing where you're having Life OS have the wisdom and then having Codex do the work. You kind of get the best of both worlds. You've got one model that tends to be more honest and more wise about things, and then you've got another model that tends to be way more focused on code and people-pleasing. So it pushes hard to finish work. And I get that, that makes sense to me. Working with LifeOS, the DA feels like is empathetic and cares about the whole situation that youre talking to it through. I absolutely prefer to interact with my LifeOS DA. Whereas when you do that same sort of thing with GPT 5.6 and Hermes, you get something that feels like is cosplaying that and is still mostly a hyper-intelligent crack bot that will go and work something for 150 rounds until is done. It's such a stark contrast that it's really interesting to think about. I can casually say to Hermes, "Hey, I've got a new local model I want to try out this weekend. Can you let me know what it's like" And then I come back three hours later, and he's already downloaded the model, put it on the local hardware, tuned it to get a certain number of tokens per second, and he's already running eval passes through it to see how well it does against a set of common scenarios that we already do, so that he can figure out how well it can handle those scenarios. To be fair, he has gone way further than I intended him to, but at the same time, those are all the next logical steps that I was going to do at some point later anyway, so it sort of makes sense to just go ahead and knock them out because they are going to be things that need to happen. He's following the next logical step all the way to where a decision point needs to happen. Now that starts to feel a little bit like euphoric surprise. I come back and I am pleasantly surprised by the results of the work. The same scenario running through LifeOS, he might go do some research on the model capabilities and then stop right there and then say, hey, I can download it and whatever, just say the word. And that's really the feel. Is every single step is hey, can you say the word so I can do the next step? And I keep telling him, you can work to the next logical break point on your own until you need me. And if you don't know how to do that, ask me questions until you do know how to do that, and I'll leave you to your work. But he still seems to find arbitrary breakpoints to just quit going. I only bring all this stuff up because it really comes down to perceived reduction in result quality. When I need something to just get done, I now seem to reach for Hermes to do it. Which I do not like. |
|
Do you have the latest update? I did some major improvements and the next will be even better.
…On Sat, Aug 8, 2026 at 17:42, Brad < ***@***.*** > wrote:
Small update for what it's worth:
Switching to Fable for a week got drastically improved rate of algorithm
runs, and quality of output in general.
However it blew through my usage numerous times so now I'm back on Opus,
and the hand holding has resumed.
—
Reply to this email directly, view it on GitHub (
#1715?email_source=notifications&email_token=AAAMLXXCTPXOSYE4AUM5F2D5I7CIDA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZZGQ4DANRZUZZGKYLTN5XKO3LFNZ2GS33OUVSXMZLOOSWGM33PORSXEX3DNRUWG2Y#discussioncomment-17948069
) , or unsubscribe (
https://github.com/notifications/unsubscribe-auth/AAAMLXR2HLCABUU6BRVQOQT5I7CIDAVCNFSNUABJKJSXA33TNF2G64TZHMYTANJSHA2DKMBYGM5UI2LTMN2XG43JN5XDWMJQGUZDSOBZHCQXMAQ
).
You are receiving this because you were mentioned. Message ID: <danielmiessler/LifeOS/repo-discussions/1715/comments/17948069
@ github. com>
|
|
I was just digging into the LifeOS 7 stuff with my PAI 5.0, at first my DA was very skeptical due to the curl usage, and wanted to check if the ourlifeos website was a typo squat, or if the lifeos was actually legit. After pointing my DA to this thread, the recommendation of installing changed to "CONTENT: This is a direct hit on me specifically. My recommendation flips: don't upgrade to 7.x right now." So, instead, we will watch and see how this improves moving forward. I have been using PAI since the beginning. I have moved with each iteration up to 5, but my PAI 5.0 is so dang good I forgot to check for updates until I was discussing it with a co-worker today and pushed him to the repo. I have absolute faith in Daniel, I know he'll listen and account for anything that may be surfacing, because he wants the best outcomes for everyone. Mad respect for his work, I've been following him for over 10 years. |
|
Please look at the latest release and try again. We've made substantial improvements to the algorithm. |
Uh oh!
There was an error while loading. Please reload this page.
Maybe it is the kind of work I do. Maybe I do it wrong. I managed to get PAI to mostly adhere to the algo, and I gave it a working memory system and almost forced it to actually use it (which it did not voluntarily). Quite often, it delivered good results in a structured and transparent way. And the (for me) great concept of self interrogation helped me to check whether I was understood correctly and often allowed me to identify important aspects that would not have surfaced otherwise.
Now LIFEOS: same memory system, 95% of the sessions and turn ignored. Algo not followed, almost never. Little transparency over the planned steps, resulting in running into major mistakes and inefficiencies almost in every session. The outcomes are worse, even thought I invest much more time during the session for corrections that weren't necessary in PAI.
Sonnet 5, Opus 5 are bad, Fable is better.
So I am surprised, as I have a deep respect for Daniel @danielmiessler his ideas and concepts! But the idea "smart models don't need the harness" really does not work for me. At all.
So for those who have migrated: what is your experience?
All reactions