Rendered at 18:39:05 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
eyalitki 11 hours ago [-]
Comparison was done in the scope of coderabbit AI code review tool, which sadly makes it practically irrelevant.
My personal experience as a software engineer, and a former security researcher who did manual code audit, is that this code review tool has such poor results that it isn't worth the "noise" and friction it causes developers during C/I code review
stingraycharles 11 hours ago [-]
Yeah I personally don’t understand the point of AI code review tools all that much, as AI is already generating the code as well. All of these AI code review tools create so much noise, yet don’t catch the really important things.
ZephyrBlu 7 hours ago [-]
Thinking of AI-generated code and AI code reviews as the same "AI" is not correct. The reviewer is using a fresh context window with no previous knowledge of the changes. That is why it's powerful, because you get the agent to interrogate the code without any preconception about the changes.
I found the Devin reviewer to be very good, and have heard good things about Cursor's Bugbot. I've also found asking an agent with fresh context or subagent to adversarially review locally is good.
stingraycharles 4 hours ago [-]
I’m not saying that it’s not useful, I’m saying that it’s not useful in a “human in the loop” situation. This type of AI-to-AI review should be done agent-to-agent, not through Github PRs with tools like Devin.
In a manual review, I then expect all “machinery” to already be properly reviewed, and can focus on design / architecture. I would like an AI assisted review tool to make that part easier, not do the actual review for me.
amluto 44 minutes ago [-]
I’m not convinced by the automatic agent-to-agent thing. I find that, if I manually ask a standard harness “Review a..b, individually and for combined effect”, I get some mix of catching genuine errors (some quite deep), incorrect flags where the correct course of action is to ignore them or modify the commit messages, and comments where the correct course of action is to think deeply.
If I were to automate the back and forth, I would get spurious changes that “fix” what wasn’t broken and a removal of the actual interesting bits.
CuriouslyC 6 hours ago [-]
Specialized AI code review software is so pointless though. Back when agents were dumb about git surgery and tool use it might have had a purpose, but now you could replace coderabbit with a skill and I bet the results would be better in some cases.
pipes 4 hours ago [-]
Mechanical code review by something like sonar qube is much cheaper. Use that for low bar quality gate and AI after.
AI code review is startlingly effective. Continually finding things me and my colleagues never would. Well the decent models do. Maybe not so much the cheap ones.
nkmnz 8 hours ago [-]
The noise is a huge problem, indeed. Still, a panel of review agents using models and harnesses different from the one implementing a set of changes has proven immensely useful for myself. The panel is basically an n×m matrix of agents and highly specific review prompts, i.e.:
- review for intent fulfillment: is the ticket done?
- review for correctness: race condition bugs, ...
- review for security: check against this list of sources and best practices
- review for api conformity: identify all surfaces of systems outside this codebase touched by the code changes and check against their docs
- etc. pp., same for maintainability, observability & analytics, test coverage, usage of feature flags
The matrix is sparse, so not every model is used with each of the review categories. Effort levels vary, too. The next stage does a consolidation across all findings, then another stage spins up one agent per finding and investigates the whole codebases for identical / similar instances of the finding; finally, it suggests a fix.
This works extremely well for finding deficits, but the amount of noise drives me insane, too. Lots of feedback is technically correct and "by the book", but pretty useless in practical terms – or even detrimental because the amount of code written and thus the size of the change set explodes. I'm not yet sure how to tackle this problem, any suggestions are welcome!
irthomasthomas 5 hours ago [-]
You need to literally review the review with another llm pass to push back on the first. Ask it to do something like reassess the severity claims and only surface real P0 to P2 issues.
ptrl600 3 hours ago [-]
They're pretty good for me, because I am still writing all the code, and it tends to catch the sorts of things humans mess up.
spockz 7 hours ago [-]
Where I find it shines it to find inconsistencies. My readme or docs or ADR something should work like X but it finds a test where it tests something different and the test is green. Or other similar drift.
Yes, your prompt need to include to look for certain “quality” aspects you care about. But once that is there it can help find a lot of things.
It can also help in finding edge cases. It is really about the prompt.
pinkgolem 11 hours ago [-]
What really important things are human reviews catching in your org?
I just feel more and more like the effort invested in manual reviews is not worth it
grokys 10 hours ago [-]
1. Whether the thing should be done in the first place
2. If it's the correct solution on a high level
3. Whether it conflicts with or duplicates other parts of the system
4. Whether the comments are actually useful or restating the LLM chat
Also many others but these are the most common IME
9dev 10 hours ago [-]
All of these are angles an AI reviewer can test for as well, and will (IME) mostly catch mistakes correctly. I also still manually review code, and usually also catch issues, but the severity of what I find shrinks ever further as agents get better.
The sprawling code comments are becoming the most draining part of code review though, that's really killing me from the inside.
embedding-shape 8 hours ago [-]
> All of these are angles an AI reviewer can test for as well, and will (IME) mostly catch mistakes correctly.
No, none of today's AI would give you enough signal around "should this thing be built in the first place" nor if it's the correct solution on a high level.
They don't understand why you are doing what you are doing, and even if you explain it, they still don't actually understand the motivation and lots of other things.
You'll get them to do guesses and pretend they actually know how to prioritize and will tell you it makes lots of sense, whatever they come up with. But try following it blindly and you'll see where you end up.
This is why "one agent + one good developer" beats "thousands of agents working in a swarm" still today.
9dev 6 hours ago [-]
I don’t think I claimed agent reviews to be a panacea. It’s a tool that can help you lower the review pressure in companies working with agentic coding tools.
embedding-shape 6 hours ago [-]
Someone asked:
> What really important things are human reviews catching in your org?
Another person said:
> 1. Whether the thing should be done in the first place
And you replied:
> All of these are angles an AI reviewer can test for as well, and will (IME) mostly catch mistakes correctly
Which as I noted, is very far from the truth. I neither claimed that you said "agent reviews are a panacea", but when you claim "AI can solve all those things" and two of the first items cannot be addressed by AI (today), then I'm rebuking those specific things, not some other general point you implicitly made.
9dev 6 hours ago [-]
But that’s the thing, when I say "mostly" that sure doesn’t imply it can solve all those things - it can help to a great extent.
embedding-shape 6 hours ago [-]
It is mostly useless at figuring out if "should this thing be built in the first place" and "if it's the correct solution", and mostly cannot help at all with those things.
Where "mostly" means kind of what it says but also not really.
Paria_Stark 10 hours ago [-]
The AI review are still quite far from having the same level of critical thinking and high level knowledge of your application, what you have done in the past and want to do next etc.
If you don't master this for your own project, what's even the point of your job.
rhdunn 8 hours ago [-]
1. Does the implementation fit in the architecture/style of the project?
2. Are there potential security, accessibility, performance, etc. issues?
3. Domain specific knowledge (SQL, ASP.NET, XQuery, etc.) where there are better ways of solving a problem, or possible issues not handled.
4. Sense checking ... is the code easy to read? does it need an explanatory comment? does it need named parameters? etc.
jiggawatts 11 hours ago [-]
Code review tools are designed for less organised dev teams that don’t do PRs and mandatory human reviews already.
It is papering over a lower level of competency without having to invest in actual human oversight or real process improvement.
9dev 10 hours ago [-]
That's a thoroughly uncharitable view. Especially in smaller orgs with a minimum velocity dictated by the company's need to survive, the amount of code required to be written just to keep up with your competitors is massive. Trying to review that all by hand, thoroughly, is draining, thankless, and tedious. You end up with a few fast movers producing most of the code, and some slower movers forced into a reviewer role they never signed up for. It's an unhealthy dynamic.
maxdo 10 hours ago [-]
They do catch important things but it’s really contextual. You can’t grab a model slap it on top and say code review . Hence a dedicated review tool is almost dead . Code review should be part of your pipeline and consume test results from the original task , open spec etc . If you do not have that code review will not help if you do , what is the point of task rabbit just slap <your harness in the sandbox> review against <goal>
bitlad 10 hours ago [-]
It does add lot of noise after a point you start ignoring the suggestions and findings.
Code generated these days with fable and sol are near perfect. What issues they might have is logical errors.
Sharlin 8 hours ago [-]
You have a very interesting definition of "perfect", then.
OtomotO 9 hours ago [-]
> Code generated these days with fable and sol are near perfect.
If you're doing a simple CRUD app, sure.
If you're doing anything more involved they get the job done with dozens of shortcuts that bite you in the ass the moment you have on-call duty.
Way too much code and repetition and hacks.
Especially in GPU code, but also in other fields.
sdeframond 8 hours ago [-]
How do you guys review AI-generated code ?
In our team, frontend work is vibe-coded by the PO and merged as-is without review. Backend is coded by developers, using AI but in a slower, more controlled way.
Recently, our PO has been trying his hand at vibe-coding the backend. I must say he is a smart guy, almost technical but not quite a developer. We've just been handed a burst of stacked PRs amounting for ~15k LOC backend. We do not quite know what do to about it.
I know we are not the only ones in the situation. What's your experience and context ? What do you do ? What works for you what doesn't ?
mobelkh 8 hours ago [-]
throw his garbage out, the time and effort taken to review that is magnitudes more than what it took to prompt it.
have him start with an overall design doc if his change is 15k, it's definitely worth a design doc.
and then have his contributions reviewed in pieces of 200-300 LoC PRs.
any other solution is trading stability and system knowledge, that's 15k LoC no one is truly familiar with, even if you do try to review it
sdeframond 2 hours ago [-]
He's made a bunch of 1-2k LOC PRs and there is a design doc. Everything is AI generated.
The issue is, if he generates all of that without reviewing the code, he will always be far faster than us. And he can't review the code. No matter how he slices it.
Also, he is the CPO/CTO. So we can say no, but there is a natural incentive to go his way. He still doesn't feel confident enough to just bypass the programmers and he's probably right. But it'd nice to find a way to use my expertise to review this amount of code meaningfully, somehow.
embedding-shape 7 hours ago [-]
Yeah, it's the "eager apprentice" problem, common almost everywhere. Solution is to make them stop and double-check before running ahead, in software development, concise design documents outlining what the problem is, what possible solutions are and what the chosen solution is, and why, then review this together with the person, before they can move on to implement it.
sdeframond 1 hours ago [-]
This particular apprentice is also my boss, an overall reasonable guy and has more experience in the software industry than myself, so there's that. He's just not a developer.
fabianlindfors 8 hours ago [-]
We try to avoid reviewing AI-generated code and built our own testing framework and platform to make that possible.
Our principle is that our tests should give us enough confidence to not have to look at the code (which ends up being true for most changes we make). The core thing that makes this possible is that we run our entire code and infra (including fakes for external dependencies) in isolated, forkable environments and write tests against that, so they are as E2E as can possibly be.
The problem then shifts from reviewing code to reviewing tests and that's why we built our own platform. We have a UI that can diff tests, so we know what changed, and a visual way to inspect what the tests actually did. A test could drive a browser like a user would, and in our UI we get a replay of that browser interaction to look at. The browser is talking to a real version of our backend, and the tests can perform assertions against the database and fakes and really anything in our system.
sdeframond 7 hours ago [-]
Well, (AI-generated) test are about half of these PRs' code. So that's still ~8k lines to review...
What techno/service did you base your framework on? How long did it take to set it up? How many are you?
fabianlindfors 7 hours ago [-]
That's the point, we don't review the test code either. Our platform gives us a UI for inspecting not the test code but what actually happened during the test. Like a browser replay, the results of a database query, assertions against those, etc.
This is much more information dense than something like the tests and is a representation of what actually happened during the tests, rather than what the test itself did (which I agree sucks to review, especially AI-generated).
The framework is our own that bundles/adapts some familiar components: Jest-like asssertions, Playwright browser API, typed database client from Bun, Kubernetes client, etc. The tests are written in Typescript but the main code doesn't have to be (just runs containerized in the environment).
We've been building the platform and using it continuously since May but setting it up on a new project takes like 1-2 days of largely autonomous coding agent work. We are just two engineers on our team but have been onboarding other startups to the platform recently so there are a few different teams using it now for their own codebases. It's fully generic so works for any infra or stack.
eterps 8 hours ago [-]
But with full blown e2e browser tests the test suite duration can go through the roof. How do you deal with that?
fabianlindfors 7 hours ago [-]
Forking!
We run the entire stack (browser, frontend, backend, database, etc) in a Linux VM, so latency between each of the pieces is as tiny as can be. This is quite different from "standard" E2E tests I've seen where the test browsers uses something like a persistent staging environment.
The real key is that we can fork that entire Linux VM to take different paths down our testing scenarios, and can run multiple of them in parallel. Tests may look something like:
new user signs up:
|- creates a todo
|- ...
|- ...
|- creates a list
The two nested tests then start from the exact same point, where the previous test left off, but can run in parallel. With enough hardware, the full suite will run as fast as the slowest branch of the test tree. When we switched away from our previous integration test suite to this (not E2E), our tests actually became faster because they share setup through the forking.
eterps 7 hours ago [-]
Which VM technology do you use?
fabianlindfors 6 hours ago [-]
Firecracker, with some tiny modifications to better manage memory for the deep nesting of forks
sdeframond 2 hours ago [-]
Would you mind sharing your infra budget needed to spin these VMs ?
Surely it is reasonable, but also way more than our budget. Id like to compare.
mjmjmjmj 6 hours ago [-]
[dead]
CuriouslyC 6 hours ago [-]
Code review is soon to be an outmoded concept, (un)fortunately. You have to design orthogonal code (e.g. independent modules in a modular monolith, or microservices) and soak test using canaries.
ninkendo 7 hours ago [-]
The answer to this is gonna vary wildly depending on what kind of codebase it is.
A large, mature codebase that predates LLM’s and for which changes need a high level of scrutiny regardless of who made them (think llvm, WebKit, important foundational software), you’re going to want humans in the loop as much as ever… I think reviewing LLM output is the most important thing a human can provide.
But for vibe coded apps where you can just one-shot another one if anything goes wrong? Just vibe the reviews too, who cares. Let the robots review the robots.
Be careful with doing AI review if your codebase is in the former category. Or your codebase will quickly turn into the latter. Complete with “you can just one-shot another one”, because if nobody understands the code any more, there’s not much lost by just you (or your competitors, etc) replacing it wholesale with an AI-written alternative.
I struggle with this a lot. 2 years ago we had a half dozen PR’s a day with a lot of careful review, and now there’s more like 30 of them per day and most people are just rubber stamping them after the AI reviews it. I’m still fighting the good fight trying to review every line of the PR’s I have time to look at, but that constitutes maybe 10% of them. Not only am I barely making a dent, but it’s awkward when I post nitpicks like “this function should go in this module”, etc, the author usually looks at me funny like “why are you even reading this”. Our codebase is gradually becoming more and more vibe coded, and it’s depressing me.
yomismoaqui 8 hours ago [-]
Invest in having a good test suite that validates the functionality introduced by that code. Also AI can review code in an adversarial way and apply those fixes (that ideally will keep the previous tests you did on green)
bushbaba 8 hours ago [-]
I have AI confirm the logic works as expected, but review for system design.
Often in both web/backend I’ve found AI to produce overly duplicative code, or have aspects that could be hard to maintain. Generally less due to the AI, and more because of the prompt itself.
That and even if you’re going to AI slop it up, I’d still demand it be broken up into 1-2k LOC chunks or per meaningful “thing”. This also lets us gradually ramp the change to confirm it actually works earlier on
tosh 7 hours ago [-]
I ran a few toy benches comparing Astra with Sol
and found Astra ~30% faster and at similar cost to Sol for the same outcome
the token efficiency helps Astra even though sticker price is 2.5x that of Sol
vb-8448 5 hours ago [-]
Same experience here, but I have some strange feeling.
Sol I'm used to working a month ago doesn't feel the same I'm using today, slower and less accurate. My gut feeling is that they quantize previous models to prioritize new ones and, who nows, make the new one look better.
Up to July I was using mostly anthropic models and the feeling was the same, so much so that I was able to predict every model release 1 or 2 days before public announcements.
ramon156 12 hours ago [-]
Both OAI and Anthropic seem to have released a model that is slightly better but cost ~2x the previous iteration. Interesting play
torginus 8 hours ago [-]
Generally speaking 'the Fable/Astra built GTA 6' videos are a new phenomenon, so it's clear these models have new capabilities and people will need new ways of interacting with them if they want to leverage these imo.
arthurcolle 11 hours ago [-]
Astra and Sol are the same price when you factor in token efficiency
Squarex 11 hours ago [-]
I don't know, in the Codex app, it burns the limit much faster.
sscaryterry 10 hours ago [-]
Bullshit.
arthurcolle 2 hours ago [-]
Not bullshit
villish 10 hours ago [-]
That likely won’t change if other competitors don’t take the lead at some point. If companies are willing to pay top dollar for the best models AND they get to extract as much money from Chinese labs distilling Astra/Fable it makes no sense to lower prices. Obviously not great for everyday users who don’t have unlimited money.
skrellm 7 hours ago [-]
> If companies are willing to pay top dollar for the best models
The referenced NYT article says Open Source LLM's market share increased to 58% from yesteryear's 10%.
villish 5 hours ago [-]
(58% on OpenRouter)
I believe both things can be true. OAI/Anthropic selling more tokens than ever, and downloadable weight models increasing in market share.
I doubt OpenAI would copy Fables pricing if it wasn’t financially beneficial. They can only burn through VC funding for so long.
dgellow 5 hours ago [-]
I don’t think we have data showing that fable is used significantly
kzrdude 12 hours ago [-]
That should be expected based on the scaling laws that we expect; larger models are more intelligent and cost more. Now it's very unfortunately that they don't publish the size of their models.
jstummbillig 11 hours ago [-]
Roughly how we price (high skilled) human labor.
simianwords 12 hours ago [-]
Interesting comment because it is true that Astra is costlier for the same intelligence tasks as Sol.
But this is not the same for Fable at all.
SneakyZero 11 hours ago [-]
Astra seems to be really slow. Maybe it intends to read more context. But from my experience it is definitely slower than 5.6 sol when handling same tasks.
trvz 10 hours ago [-]
It's a bigger model, of course it's slower.
tosh 7 hours ago [-]
fwiw I found Astra to be faster than Sol w/ both on medium reasoning for simple agentic coding
difficult to compare though because for more open ended, complex tasks Sol might miss something that Astra notices and then Sol might yield a cheaper but worse outcome
dude250711 9 hours ago [-]
Given that Fable is a Sol-class model, should Astra not be compared to Mythos in those tests?
KAdot 6 hours ago [-]
I heavily A/B tested Opus vs Sol for two weeks, giving Claude Code and Codex the same tasks and comparing the results. In my experience Sol is much closer to Opus than to Fable, with Opus often beating Sol. The only area where Sol is better is code reviews, which these benchmarks confirm. I wish they included Fable.
My personal experience as a software engineer, and a former security researcher who did manual code audit, is that this code review tool has such poor results that it isn't worth the "noise" and friction it causes developers during C/I code review
I found the Devin reviewer to be very good, and have heard good things about Cursor's Bugbot. I've also found asking an agent with fresh context or subagent to adversarially review locally is good.
In a manual review, I then expect all “machinery” to already be properly reviewed, and can focus on design / architecture. I would like an AI assisted review tool to make that part easier, not do the actual review for me.
If I were to automate the back and forth, I would get spurious changes that “fix” what wasn’t broken and a removal of the actual interesting bits.
AI code review is startlingly effective. Continually finding things me and my colleagues never would. Well the decent models do. Maybe not so much the cheap ones.
- review for intent fulfillment: is the ticket done?
- review for correctness: race condition bugs, ...
- review for security: check against this list of sources and best practices
- review for api conformity: identify all surfaces of systems outside this codebase touched by the code changes and check against their docs
- etc. pp., same for maintainability, observability & analytics, test coverage, usage of feature flags
The matrix is sparse, so not every model is used with each of the review categories. Effort levels vary, too. The next stage does a consolidation across all findings, then another stage spins up one agent per finding and investigates the whole codebases for identical / similar instances of the finding; finally, it suggests a fix.
This works extremely well for finding deficits, but the amount of noise drives me insane, too. Lots of feedback is technically correct and "by the book", but pretty useless in practical terms – or even detrimental because the amount of code written and thus the size of the change set explodes. I'm not yet sure how to tackle this problem, any suggestions are welcome!
Yes, your prompt need to include to look for certain “quality” aspects you care about. But once that is there it can help find a lot of things.
It can also help in finding edge cases. It is really about the prompt.
I just feel more and more like the effort invested in manual reviews is not worth it
2. If it's the correct solution on a high level
3. Whether it conflicts with or duplicates other parts of the system
4. Whether the comments are actually useful or restating the LLM chat
Also many others but these are the most common IME
The sprawling code comments are becoming the most draining part of code review though, that's really killing me from the inside.
No, none of today's AI would give you enough signal around "should this thing be built in the first place" nor if it's the correct solution on a high level.
They don't understand why you are doing what you are doing, and even if you explain it, they still don't actually understand the motivation and lots of other things.
You'll get them to do guesses and pretend they actually know how to prioritize and will tell you it makes lots of sense, whatever they come up with. But try following it blindly and you'll see where you end up.
This is why "one agent + one good developer" beats "thousands of agents working in a swarm" still today.
> What really important things are human reviews catching in your org?
Another person said:
> 1. Whether the thing should be done in the first place
And you replied:
> All of these are angles an AI reviewer can test for as well, and will (IME) mostly catch mistakes correctly
Which as I noted, is very far from the truth. I neither claimed that you said "agent reviews are a panacea", but when you claim "AI can solve all those things" and two of the first items cannot be addressed by AI (today), then I'm rebuking those specific things, not some other general point you implicitly made.
Where "mostly" means kind of what it says but also not really.
If you don't master this for your own project, what's even the point of your job.
2. Are there potential security, accessibility, performance, etc. issues?
3. Domain specific knowledge (SQL, ASP.NET, XQuery, etc.) where there are better ways of solving a problem, or possible issues not handled.
4. Sense checking ... is the code easy to read? does it need an explanatory comment? does it need named parameters? etc.
It is papering over a lower level of competency without having to invest in actual human oversight or real process improvement.
Code generated these days with fable and sol are near perfect. What issues they might have is logical errors.
If you're doing a simple CRUD app, sure.
If you're doing anything more involved they get the job done with dozens of shortcuts that bite you in the ass the moment you have on-call duty.
Way too much code and repetition and hacks.
Especially in GPU code, but also in other fields.
In our team, frontend work is vibe-coded by the PO and merged as-is without review. Backend is coded by developers, using AI but in a slower, more controlled way.
Recently, our PO has been trying his hand at vibe-coding the backend. I must say he is a smart guy, almost technical but not quite a developer. We've just been handed a burst of stacked PRs amounting for ~15k LOC backend. We do not quite know what do to about it.
I know we are not the only ones in the situation. What's your experience and context ? What do you do ? What works for you what doesn't ?
have him start with an overall design doc if his change is 15k, it's definitely worth a design doc.
and then have his contributions reviewed in pieces of 200-300 LoC PRs.
any other solution is trading stability and system knowledge, that's 15k LoC no one is truly familiar with, even if you do try to review it
The issue is, if he generates all of that without reviewing the code, he will always be far faster than us. And he can't review the code. No matter how he slices it.
Also, he is the CPO/CTO. So we can say no, but there is a natural incentive to go his way. He still doesn't feel confident enough to just bypass the programmers and he's probably right. But it'd nice to find a way to use my expertise to review this amount of code meaningfully, somehow.
Our principle is that our tests should give us enough confidence to not have to look at the code (which ends up being true for most changes we make). The core thing that makes this possible is that we run our entire code and infra (including fakes for external dependencies) in isolated, forkable environments and write tests against that, so they are as E2E as can possibly be.
The problem then shifts from reviewing code to reviewing tests and that's why we built our own platform. We have a UI that can diff tests, so we know what changed, and a visual way to inspect what the tests actually did. A test could drive a browser like a user would, and in our UI we get a replay of that browser interaction to look at. The browser is talking to a real version of our backend, and the tests can perform assertions against the database and fakes and really anything in our system.
What techno/service did you base your framework on? How long did it take to set it up? How many are you?
This is much more information dense than something like the tests and is a representation of what actually happened during the tests, rather than what the test itself did (which I agree sucks to review, especially AI-generated).
The framework is our own that bundles/adapts some familiar components: Jest-like asssertions, Playwright browser API, typed database client from Bun, Kubernetes client, etc. The tests are written in Typescript but the main code doesn't have to be (just runs containerized in the environment).
We've been building the platform and using it continuously since May but setting it up on a new project takes like 1-2 days of largely autonomous coding agent work. We are just two engineers on our team but have been onboarding other startups to the platform recently so there are a few different teams using it now for their own codebases. It's fully generic so works for any infra or stack.
We run the entire stack (browser, frontend, backend, database, etc) in a Linux VM, so latency between each of the pieces is as tiny as can be. This is quite different from "standard" E2E tests I've seen where the test browsers uses something like a persistent staging environment.
The real key is that we can fork that entire Linux VM to take different paths down our testing scenarios, and can run multiple of them in parallel. Tests may look something like:
The two nested tests then start from the exact same point, where the previous test left off, but can run in parallel. With enough hardware, the full suite will run as fast as the slowest branch of the test tree. When we switched away from our previous integration test suite to this (not E2E), our tests actually became faster because they share setup through the forking.Surely it is reasonable, but also way more than our budget. Id like to compare.
A large, mature codebase that predates LLM’s and for which changes need a high level of scrutiny regardless of who made them (think llvm, WebKit, important foundational software), you’re going to want humans in the loop as much as ever… I think reviewing LLM output is the most important thing a human can provide.
But for vibe coded apps where you can just one-shot another one if anything goes wrong? Just vibe the reviews too, who cares. Let the robots review the robots.
Be careful with doing AI review if your codebase is in the former category. Or your codebase will quickly turn into the latter. Complete with “you can just one-shot another one”, because if nobody understands the code any more, there’s not much lost by just you (or your competitors, etc) replacing it wholesale with an AI-written alternative.
I struggle with this a lot. 2 years ago we had a half dozen PR’s a day with a lot of careful review, and now there’s more like 30 of them per day and most people are just rubber stamping them after the AI reviews it. I’m still fighting the good fight trying to review every line of the PR’s I have time to look at, but that constitutes maybe 10% of them. Not only am I barely making a dent, but it’s awkward when I post nitpicks like “this function should go in this module”, etc, the author usually looks at me funny like “why are you even reading this”. Our codebase is gradually becoming more and more vibe coded, and it’s depressing me.
Often in both web/backend I’ve found AI to produce overly duplicative code, or have aspects that could be hard to maintain. Generally less due to the AI, and more because of the prompt itself.
That and even if you’re going to AI slop it up, I’d still demand it be broken up into 1-2k LOC chunks or per meaningful “thing”. This also lets us gradually ramp the change to confirm it actually works earlier on
and found Astra ~30% faster and at similar cost to Sol for the same outcome
https://x.com/__tosh/status/2096201900555170032
the token efficiency helps Astra even though sticker price is 2.5x that of Sol
Sol I'm used to working a month ago doesn't feel the same I'm using today, slower and less accurate. My gut feeling is that they quantize previous models to prioritize new ones and, who nows, make the new one look better.
Up to July I was using mostly anthropic models and the feeling was the same, so much so that I was able to predict every model release 1 or 2 days before public announcements.
But they are not willing to pay. https://news.ycombinator.com/item?id=49566137
The referenced NYT article says Open Source LLM's market share increased to 58% from yesteryear's 10%.
I believe both things can be true. OAI/Anthropic selling more tokens than ever, and downloadable weight models increasing in market share.
I doubt OpenAI would copy Fables pricing if it wasn’t financially beneficial. They can only burn through VC funding for so long.
But this is not the same for Fable at all.
difficult to compare though because for more open ended, complex tasks Sol might miss something that Astra notices and then Sol might yield a cheaper but worse outcome