I was very hesitant about Cursor Composer 2 because it wasn’t good. But now you’ve managed to get it right!!! Congratulations to the entire Cursor Composer team.
Composer 2.5 is incredible!!!
I was very hesitant about Cursor Composer 2 because it wasn’t good. But now you’ve managed to get it right!!! Congratulations to the entire Cursor Composer team.
Composer 2.5 is incredible!!!
Composer 2.5 is awesome! Try Planning with it too! Great results!
Hey, thanks for the kind words. It’s nice to hear, and I’ll definitely pass it on to the Composer team.
And yeah, @RasalJayasinghe’s tip is spot on. Composer 2.5 works really well with Plan mode, especially on more complex tasks. If you’ve got ideas or anything you’d like to see improved, drop it in this thread.
I made a tool to split PNG image frame sequences and convert them to GIFs.
I wrote a plan using Codex 5.3 and then executed it with Composer 2.5. It completed the entire project in one go, and I’m very satisfied. This shows that while Composer 2.5 might not be the smartest, but it doesn’t forget its goals in long-running tasks, which is better than other models.
So, this is my set-location for cursor models:
do you agree?
Hey, great role breakdown. The Codex or Claude for planning plus Composer 2.5 for execution combo is a really solid pattern, especially when the task is long and you need to keep the goal in mind as you go. Composer 2.5 is great at holding context and following the plan through to the end.
There’s no strict correct split, a lot depends on your stack and the task, so your setup looks totally reasonable. If you find combos that work even better for specific cases, share them here. Those observations are useful.
Composer 2.5 works well because it is diligent, and not overly smart, which can be problematic I find with the big frontier models. We need good programming models as much, if not even more, as we need the high intelligence.
I am also having great success using Opus and GTP-5.4 for planning and Composer 2.5 for coding tasks.
2.5 is also very good at digging thru millions of SQL database records to put together telemetry analysis. The frontier models recognize this and delegate tasks to composer 2.5 frequently, which is good to see.
2.5 does have trouble remembering my “Always Apply” rules though, thats the biggest weakness I have found
Well done Cursor crew !
Hey, thanks for such a detailed breakdown, it was a great read. And I totally agree with your point: Composer 2.5 is strongest as a diligent do-work model that keeps the goal in mind across long tasks. The combo of a frontier model for planning plus Composer 2.5 for execution really works great.
On the Always Apply rules, yep, I’ve seen the same kind of reports, we’re keeping an eye on it. If you want, drop an example rule here with frontmatter alwaysApply: true and your Cursor version, and we can dig into what’s getting lost in your case. This thread covers similar reports and might help: `alwaysApply: true` rules are being completely ignored now
If you find any other good model pairings for specific tasks, post them in the thread, notes like that are super helpful.
I came and found this thread just so I could also say great job with Compose 2.5! For the past couple weeks, I’ve used it exclusively for all of my discussions, planning, and executions and it hasn’t missed a beat. I don’t miss waiting for Opus while it spun itself into a logical circle. Compose 2.5 works just as well as Claude, ChatGPT, in my experience and is much faster.
I really appreciate what y’all are doing to serve the software development community!
not related but if you try deepseek v4 flash, it is much cheaper and as good.
composer 2.5 fall into the middle level model which is not good enough as compare to the top models but not as cheap as compare to the deepseek v4 flash
Hey, thanks, @Howdy. Glad to hear Composer 2.5 is working well for planning and execution. I’ll pass that on to the team.
@Xinlin_Lin, thanks for the tip about DeepSeek V4 Flash, I’ll note it. On positioning, Composer 2.5 is built around our harness and performs best as an execution model that can stay on goal over long tasks. That’s a bit different from a simple cheap vs smart comparison. But I get the point about price, it’s helpful to know what you’re comparing.
If you find model pairings or scenarios where something works noticeably better or worse, drop them here. Those notes really help.