This makes a ton of sense to me with the huge corpus of web standards that exist. In theory we should be able to generate a browser from those specs, it just took a massive amount of effort before. Maybe if things like this get some traction, some of the human time spent implementing the spec could be spent on creating more/better specs, allowing for even better generation too.
Edit: replies are making good points about AI capabilities and where the effort really goes. Let’s just say that I meant this in an aspirational sense, rather than where the rubber actually meets the road today.
As someone who has spent the last 3 years implementing a browser engine from scratch full time (that currently passes ~half the 200k "css" tests that test style/layout/rendering), I find that (without a lot of close hand holding) the AIs are very far from being able to do this. They'll give you something that passes the tests, but it will do it a ridiculous way be far too slow to be useful (and is wrong architecturally such that it's not going to converge on a better solution).
I guess my hope is that while this may be the case now, perhaps it won’t be eventually. If you can generate from a spec, different implementations may be slower or worse in a variety of ways but still meet the spec.
At that point it’s about optimizing and deciding trade offs, both of which might mean heavily directing the generation in some way. Over time, maybe LLMs or whatever succeeds them won’t go down so many bad architectural paths. We’ve certainly seen that already in the last year or two.
I also don’t think folks like you ever lose your value in these efforts, even if all of it came true tomorrow. At worst it would be you plus the AI, which would always be a more potent combination than say, me with an AI trying to make a browser engine.
What does this have to do with Linux? Do you really think LLMs will only belong to 2 or three companies in the future, when we already have fairly competitive open weights usable from home (or rented servers)? How would this be anything but the opposite of centralization amongst 2 or 3 companies, which is already the case now?
Software has always been displacing things, especially itself.
> I, too, hope that one of the two or three companies will be able to improve and displace everything. Who needs Linux anyway.
In the context of browsers the domination of one or two companies is already complete, and being able to generate new browsers with AI would hopefully allow people to escape the stranglehold of Chrome/Blink.
That said, I doubt a new,competitive browser could be created that way. The (possibly AI) developers of web apps won’t want to test on somebody’s bespoke AI-generated browser.
I recognize the feeling of handholding and have the tingles telling me that given the step change of Opus 5.5 and what passed by in the news lately, agent swarms are about to become mainstream. That's exactly what a project like this would use. I think everybody will be able to create their own engine and out of the ones who try, a few will actually make gains beyond what's available in parts of the engine.
It's a matter of "bring the best non-conflicting wins together" at that point. The existing engines are too big to flip over their codebase without having seen competitors prove it too.
My estimate is that with the release of Opus 5.7, companies need to have their shit together because the work ecosystem will flip over at 6 at the current rate of progress. This generation is the first one that delivers local applications with a better toolset than a small scale SaaS in hours to days.
Not sure why you are being downvoted, but this is definitely a limitation of current models and difficult to optimize for. Anything you are not specifically optimizing for is essentially unconstrained: the model could learn it by pure chance, but the likelihood is exceedingly low.
Techniques, like the newly announced RL-XAR from Meta [1] are being developed that will likely improve reward models and guide RL training to optimize for metrics like software architecture that are hard to verify otherwise.
There’s quite a few edge cases where browsers allow non-compliant code to render as if it were to specs. Because if a site doesn’t render “correctly” from a user perspective, they blame the browser rather than blaming the web developer.
That does happen, but these days they tend to add that kind of thing to the spec (and the test suite). There was a minor drama several years when WHATWG (representing browser vendors) effectively took over the spec process from the W3C because they were fed up with the W3C taking an idealistic viewpoint, and speccing things that they couldn't actually implement.
To be fair - my recollection was "representing browser vendors and pragmatists". It was far from rhe industry stitch-up that that might sound like to modern ears.
That’s good to know. The W3C probably did more to hamper web standards than anyone else (except maybe Microsoft). So I’m not surprised WHATWG took over.
Where in the web standards do you find specifications like how bookmark toolbars and menus should be organized, or that Shift-Ctrl-T should bring back a tab that was closed?
I think browser performance (and security) improvements are really difficult to tackle at the scale of browser complexity, and are both art and engineering. In other words, if this were possible, we would see the results first in existing browsers. This project is very interesting though.
Maybe some day this will be the basis but not the complete solution. Agreed that there’s a lot of semi intangible art that goes into browser engine decisions. I’m guessing a lot of which have to do as much with people and current landscape dynamics as they do with the tech.
> In theory we should be able to generate a browser from those specs, it just took a massive amount of effort before.
No, not really. The absolute vast majority of those specs are human technical and technical-adjacent language, not machine-readable specs.
On top of that many web specs often invent new terminology because a lot of things are specced years or decades after something popular has taken over the term in userland.
Maybe over time the specs will become more machine readable (bounded and concrete) if we take these paths. Although in some sense everything is machine readable these days, just not deterministically so (less helpful for specs, but not nothing either).
Your second point makes a lot of sense to me too. I’m not sure how this could handle that facet of human nature, except possibly to indirectly contribute to speeding up the cycle of spec creation.
Very cool idea, though I think we're a long way away from this. Hopefully they file bugs on the specs when they find ambiguities.
Also the specs are written to define observable behavior, and there's a fair bit of ambiguity that's UA defined, but to be actually compatible with the Web requires doing what Chrome does.
The good news is that you can look at the source of the 3 major engines (and Ladybird!) and your AI agent can do comparisons and figure out optimizations.
I think finding bugs and ambiguities in the spec is actually the highest value thing anyone can do for W3C right now. 17 years after the first draft, we're still finding Flexbox Level 1 ambiguities, and this it still hasn't exited Candidate Recommendation (stage 2 out of 4).
I'm sure this is not the origin, but Bez is a celebrity in England, a member of a 90s band called happy Mondays who didn't sing or play anything but just danced. Famous on the panel show circuit
So far it's only HTML/CSS (and we have our Rust-based framework to write the apps in). But it has been designed for fast incremental rendering, so it could be extended with JavaScript support quite easily.
that looks very interesting. if it's targeting the electron/sciter market I would encourage you to put the binary size and memory footprint in the readme so people can see where it lies on that spectrum - it's one of my top priorities when evaluating a desktop app library.
- Binary sizes start around 8mb if you're using GPU rendering (you can go smaller with CPU rendering, but you probably want the GPU). Our full browser app which pulls in things like sqlite, http cache libraries, etc is 20mb. Those usually compress to about half for distribution (.dmg, .appimage, etc).
- Base memory usage is something like 100mb (mostly from the graphics stack). I'm hoping to be able to bring that down a bit, but I think you can't realistically get much lower than 60mb with modern graphics. And to be perfectly honest we currently have an issue for RAM where it will often jump to more like 300-400mb after a little use. And I haven't fully gotten to the bottom of that yet.
Something like a 4k RGBA texture being 30mb and you need at least 2 for a swapchain. Of course you might not be rendering at 4k, but those are just output buffers before you even start counting application memory.
This makes a ton of sense to me with the huge corpus of web standards that exist. In theory we should be able to generate a browser from those specs, it just took a massive amount of effort before. Maybe if things like this get some traction, some of the human time spent implementing the spec could be spent on creating more/better specs, allowing for even better generation too.
Edit: replies are making good points about AI capabilities and where the effort really goes. Let’s just say that I meant this in an aspirational sense, rather than where the rubber actually meets the road today.
As someone who has spent the last 3 years implementing a browser engine from scratch full time (that currently passes ~half the 200k "css" tests that test style/layout/rendering), I find that (without a lot of close hand holding) the AIs are very far from being able to do this. They'll give you something that passes the tests, but it will do it a ridiculous way be far too slow to be useful (and is wrong architecturally such that it's not going to converge on a better solution).
I guess my hope is that while this may be the case now, perhaps it won’t be eventually. If you can generate from a spec, different implementations may be slower or worse in a variety of ways but still meet the spec.
At that point it’s about optimizing and deciding trade offs, both of which might mean heavily directing the generation in some way. Over time, maybe LLMs or whatever succeeds them won’t go down so many bad architectural paths. We’ve certainly seen that already in the last year or two.
I also don’t think folks like you ever lose your value in these efforts, even if all of it came true tomorrow. At worst it would be you plus the AI, which would always be a more potent combination than say, me with an AI trying to make a browser engine.
> I guess my hope is that while this may be the case now, perhaps it won’t be eventually.
I hear you brother. I, too, hope that one of the two or three companies will be able to improve and displace everything. Who needs Linux anyway.
Hope you are on the right side of the fence.
What does this have to do with Linux? Do you really think LLMs will only belong to 2 or three companies in the future, when we already have fairly competitive open weights usable from home (or rented servers)? How would this be anything but the opposite of centralization amongst 2 or 3 companies, which is already the case now?
Software has always been displacing things, especially itself.
> I, too, hope that one of the two or three companies will be able to improve and displace everything. Who needs Linux anyway.
In the context of browsers the domination of one or two companies is already complete, and being able to generate new browsers with AI would hopefully allow people to escape the stranglehold of Chrome/Blink.
That said, I doubt a new,competitive browser could be created that way. The (possibly AI) developers of web apps won’t want to test on somebody’s bespoke AI-generated browser.
I recognize the feeling of handholding and have the tingles telling me that given the step change of Opus 5.5 and what passed by in the news lately, agent swarms are about to become mainstream. That's exactly what a project like this would use. I think everybody will be able to create their own engine and out of the ones who try, a few will actually make gains beyond what's available in parts of the engine.
It's a matter of "bring the best non-conflicting wins together" at that point. The existing engines are too big to flip over their codebase without having seen competitors prove it too.
My estimate is that with the release of Opus 5.7, companies need to have their shit together because the work ecosystem will flip over at 6 at the current rate of progress. This generation is the first one that delivers local applications with a better toolset than a small scale SaaS in hours to days.
> far too slow to be useful
Speed is probably not part of a web spec.
> wrong architecturally
Nor is software architecture (though web security specs may have some influence here).
Not sure why you are being downvoted, but this is definitely a limitation of current models and difficult to optimize for. Anything you are not specifically optimizing for is essentially unconstrained: the model could learn it by pure chance, but the likelihood is exceedingly low.
Techniques, like the newly announced RL-XAR from Meta [1] are being developed that will likely improve reward models and guide RL training to optimize for metrics like software architecture that are hard to verify otherwise.
[1]: https://facebookresearch.github.io/RAM/blogs/unslop/
> Speed is probably not part of a web spec.
"being a grandmother is not part of the bike frame spec so my grandmother with wheels is a bike"
What do you mean by grandmother with wheels?
Reference to meme:
https://youtube.com/watch?v=2G9ZGIOiPjY
There’s quite a few edge cases where browsers allow non-compliant code to render as if it were to specs. Because if a site doesn’t render “correctly” from a user perspective, they blame the browser rather than blaming the web developer.
That does happen, but these days they tend to add that kind of thing to the spec (and the test suite). There was a minor drama several years when WHATWG (representing browser vendors) effectively took over the spec process from the W3C because they were fed up with the W3C taking an idealistic viewpoint, and speccing things that they couldn't actually implement.
To be fair - my recollection was "representing browser vendors and pragmatists". It was far from rhe industry stitch-up that that might sound like to modern ears.
That’s good to know. The W3C probably did more to hamper web standards than anyone else (except maybe Microsoft). So I’m not surprised WHATWG took over.
Where in the web standards do you find specifications like how bookmark toolbars and menus should be organized, or that Shift-Ctrl-T should bring back a tab that was closed?
I think browser performance (and security) improvements are really difficult to tackle at the scale of browser complexity, and are both art and engineering. In other words, if this were possible, we would see the results first in existing browsers. This project is very interesting though.
Maybe some day this will be the basis but not the complete solution. Agreed that there’s a lot of semi intangible art that goes into browser engine decisions. I’m guessing a lot of which have to do as much with people and current landscape dynamics as they do with the tech.
> In theory we should be able to generate a browser from those specs, it just took a massive amount of effort before.
No, not really. The absolute vast majority of those specs are human technical and technical-adjacent language, not machine-readable specs.
On top of that many web specs often invent new terminology because a lot of things are specced years or decades after something popular has taken over the term in userland.
Maybe over time the specs will become more machine readable (bounded and concrete) if we take these paths. Although in some sense everything is machine readable these days, just not deterministically so (less helpful for specs, but not nothing either).
Your second point makes a lot of sense to me too. I’m not sure how this could handle that facet of human nature, except possibly to indirectly contribute to speeding up the cycle of spec creation.
I'm looking forward to the day when we have fully functioning web browsers that we have full programmatic control over in all aspects.
With luck all chromium/blink-based browsers will go the way of the dodo bird.
I'm afraid internally they would still look like chromium/blink anyway.
Very cool idea, though I think we're a long way away from this. Hopefully they file bugs on the specs when they find ambiguities.
Also the specs are written to define observable behavior, and there's a fair bit of ambiguity that's UA defined, but to be actually compatible with the Web requires doing what Chrome does.
The good news is that you can look at the source of the 3 major engines (and Ladybird!) and your AI agent can do comparisons and figure out optimizations.
I think finding bugs and ambiguities in the spec is actually the highest value thing anyone can do for W3C right now. 17 years after the first draft, we're still finding Flexbox Level 1 ambiguities, and this it still hasn't exited Candidate Recommendation (stage 2 out of 4).
Is this named like Jev intentionally? Or is this just a new naming convention that I haven't been watching?
I'm sure this is not the origin, but Bez is a celebrity in England, a member of a 90s band called happy Mondays who didn't sing or play anything but just danced. Famous on the panel show circuit
Slander! What about the (occasional) maracas?
Those were bowling pins, it was all fake
Maybe it’s one of those weird LLM smells, like the OpenAI goblin problem.
https://openai.com/index/where-the-goblins-came-from/
This is a neat idea. It opens up a whole new way to make a web app into a native app, while also adding native features not available in web views
I have a working (and not vibe coded) implementation of a "browser engine for apps": https://github.com/DioxusLabs/blitz
So far it's only HTML/CSS (and we have our Rust-based framework to write the apps in). But it has been designed for fast incremental rendering, so it could be extended with JavaScript support quite easily.
that looks very interesting. if it's targeting the electron/sciter market I would encourage you to put the binary size and memory footprint in the readme so people can see where it lies on that spectrum - it's one of my top priorities when evaluating a desktop app library.
Good idea - thanks.
FWIW:
- Binary sizes start around 8mb if you're using GPU rendering (you can go smaller with CPU rendering, but you probably want the GPU). Our full browser app which pulls in things like sqlite, http cache libraries, etc is 20mb. Those usually compress to about half for distribution (.dmg, .appimage, etc).
- Base memory usage is something like 100mb (mostly from the graphics stack). I'm hoping to be able to bring that down a bit, but I think you can't realistically get much lower than 60mb with modern graphics. And to be perfectly honest we currently have an issue for RAM where it will often jump to more like 300-400mb after a little use. And I haven't fully gotten to the bottom of that yet.
thanks, those are promising numbers for sure. not being too familiar with the low level details, why is 60MB the realistic floor these days?
Something like a 4k RGBA texture being 30mb and you need at least 2 for a swapchain. Of course you might not be rendering at 4k, but those are just output buffers before you even start counting application memory.
Ok, but will you be able to understand the code?
"Think? Understand? Our robot servants do that for us."
Can we do the same for MacOS, iOS? I want to run them inside Linux.
But can it dance