Rendered at 01:26:27 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
comrade1234 2 hours ago [-]
I spent the last three days (off and on) using Gemini to configure my edge router 4 with my iOS devices on a vpn and it's been awesome. In the past I'd do a google search and read a few sources of documentation, do another google search and read another set of documentation. Now, Gemini aggregates multiple pages together so all of the work of reading source docs from multiple locations is now n a single step.
Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.
paxys 1 hours ago [-]
All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else.
If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
neuralkoi 15 minutes ago [-]
I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse.
Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.
swingandamiss 24 minutes ago [-]
Gemini can just consume the device documents. There's an incentive for device makers to publish this content.
dspillett 5 minutes ago [-]
There had always been some incentive for manufactures to publish device documentation, and yet it has often been quite lacking either in quality or overall existence. I doubt LLM/agents being the readers will change that at all. What I expect AI scraping and using without credit will impact is people publishing their own unofficial help and guidance, and the affect there is likely to be negative. It won't stop all of them, but enough to be noticeable. Another possible negative is the manufactures documentation being AI generated without sufficient review, so possibly more erroneous than before, or intentionally not producing full documentation at all and expecting AI to fill the gap (MS seems to be heading this way: pushing "ask copilot" all over Azure instead of links direct to good reference material). All this would add up to a situation that is somewhere between "a little worse than pre-AI" and "an absolute shit show".
jefftk 45 minutes ago [-]
> If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new?
I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
exmadscientist 15 minutes ago [-]
The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous, are you?), and your idea might get embedded into a set of model weights somewhere. No human will see a byte of it.
Are you actually saying you'd be OK with that?
penneyd 37 minutes ago [-]
Well in the example above the manufacturer still has incentive to provide the manual's and guides that describe how to use their products, and if that is subsequently served by an LLM that's totally fine. The only sites that LLM's would have a negative effect on are those that are only hosting content for the ad views.
Eddy_Viscosity2 28 minutes ago [-]
Manuals don't always well explain how to use their products with everybody else's products because there are too many to do that. But there are lots of people trying things out and might figure out the fine details on how to make various things work. They then publish these how-to pieces (which exist no where else) to the internet, or at least they used to when there were incentives to do so.
satvikpendem 22 minutes ago [-]
People write and create regardless of profit motive, it has been that way for thousands of years.
Havoc 33 minutes ago [-]
That is next earnings quarters problem is the approach being taken
1 hours ago [-]
_dark_matter_ 1 hours ago [-]
I wonder the same thing. I only imagine that what comes next is worse: AI companies using vast resources to develop new training data, in house, locked down. They are already doing this with developers and code at Meta. Information will become locked away behind AI paywalls and chatbots.
matt4711 55 minutes ago [-]
There will be new ways and incentives for content creators to be compensated. Many AI search startups are already talking about this or have created programs that help incentivize content creation.
indigodaddy 34 minutes ago [-]
Believe me they're making money off you one way or the other
snapplebobapple 1 hours ago [-]
its pretty good with logs too. i mostly paste a few pages i suspect hace a problem in them and let it go to town. its almost always correct and is way faster than me
queenkjuul 50 minutes ago [-]
I've been mostly enjoying Gemini, but it also clearly and definitively told me something i was trying to do was not possible with the library im using, so i wrote a different implementation, an hour later to discover that the library does in fact do precisely what i wanted in exactly the way i wanted with less headache. If I'd just gone right to the documentation instead, it actually would have saved me time.
jjulius 1 hours ago [-]
And I've spent the last few days irritated that everything I ask Gemini is answered with something that's blatantly wrong and I'm not even a subject matter expert. A quick Google search for the same questions gives plenty of results that counter what the LLM gave me.
YMMV.
-0_0- 49 minutes ago [-]
Confidently wrong summaries are the bane of Google search, and unfortunately I'm finding the AI seems to bleed into the actual search results too now, often turning up pages that back up what the summary is (wrongly) suggesting instead of surfacing actually relevant results for what I'm asking.
ls612 2 hours ago [-]
I built a simple Claude Code container on my homelab to do the same. I can just SSH in and get tech support when I need it.
stringfood 2 hours ago [-]
no, no, no. you are supposed to romanticize the hunt for correct information /s
nixonpjoshua 1 hours ago [-]
Are you sure the lack of advertising isn't long term feasible? It seems to me that AI models have proven themselves to be something consumers ARE willing to pay a subscription - or even pay per use/token for.
jjulius 1 hours ago [-]
Are they profiting from these subscriptions yet?
queenkjuul 49 minutes ago [-]
I've been using Gemini free tier exclusively for all my AI "needs" and would not be willing to pay for it
theplumber 1 hours ago [-]
It took you 4 days because you used Gemini. Gemini is the worst AI model I ever used. It is way behind even open models. It looks like Google just reached its AOL moment.
noosphr 1 hours ago [-]
Gemini isn't great as a model, googles search and ability to cite textbooks down to the paragraph make it better than every other model for human in the loop tasks.
I end up using the gemini api for with search enabled for the cases that I don't have access to good grounding data even in agentic tasks.
queenkjuul 47 minutes ago [-]
Come now, copilot is worse in every way
satvikpendem 21 minutes ago [-]
Copilot can use any model like Sol or Opus and is just a harness so not sure why people say this.
ilamont 20 minutes ago [-]
After publishers successfully sued the Internet Archive over its digital lending program, calling it unauthorized copying
No. The court specifically determined that the Internet Archive was guilty of unauthorized copying. It was not simply an unfounded or unproven allegation by publishers, book authors, and writers. The Authors Guild, the National Writers Union, the European Writers Council, and the Society of Authors in the UK all came out against the Internet Archive, and supported the suit.
news organizations started blocking the Wayback Machine’s crawlers out of fear that archived pages can provide AI companies with an indirect source of copyrighted material. Each new restriction limits the archive’s ability to act as a comprehensive backstop.
This self-inflicted damage to the wayback machine is the real tragedy of this entire affair. When IA was asked to stop CDL - many times - founder Brewster Kahle continued. The National Writers Union tried to open a dialogue as early as 2010 but was ignored:
The Internet Archive says it would rather talk with writers individually than talk to the NWU or other writers’ organizations. But requests by NWU members to talk to or meet with the Internet Archive have been ignored or rebuffed.
When the requests to abandon CDL turned into demands, Kahle dug in his heels. When the inevitable lawsuits followed, and IA lost, he insisted that he was still in the right and plowed ahead with appeals. And here we are today.
novafunc 2 hours ago [-]
I occasionally use Google Search when DuckDuckGo fails to give me relevant. Almost always, Google has better results.
Though I can find its AI answers annoying aggressive. I'll look up like two search terms and the AI will bullshit multiple paragraphs out of despite having zero context of what I am looking for.
DuckDuckGo seems to have detection of whether it should give an AI answer. And it allows you to have more granular control of when you want to get an AI answer. And is overall less distracting than Google's.
8organicbits 1 hours ago [-]
Interesting, I've seen much better results on DDG. Most recently was the search: `site:feeds.bbci.co.uk inurl:rss.xml` which works on DDG but gives zero results on Google. As far as I can tell, Google just decided not to index these.
jamesfinlayson 1 hours ago [-]
Yeah I mostly still use Google out of habit but there have been a few times where Google has decided something isn't worth indexing (too niche, doesn't use SSL).
I miss when Google was like a grep for the entire visible Internet. Now it tries to second-guess my search and direct me to a bunch of sites which all have identical information that isn't what I'm looking for.
a2ff6eeb0 59 minutes ago [-]
That was true until a few months ago. Now, it almost never has relevant results, and I've given up on it.
i don’t agree with the google has better results thing. sometimes it does. most of the time it’s just that google has the site i want higher in the ordering than DDG. personally i’m fine scrolling down a little bit more. it’s rare i need to go to google for something that DDG doesn’t have at all in their results, but it does happen.
i do have to go to google for maps/directions/planning travel. a lot that’s annoying.
or you can press the gear button -> "Ai features: Manage" -> Search assist
verdverm 25 minutes ago [-]
that control (so I can turn it off completely) is why I picked DDG to replace Google, who force feeds us the hallucinations
I've stopped using DDG now because of result quality. I now use a "meta" search backed by EXA, Tavily, and SearXNG in parallel. It can be agentically de-dupped or summarized as needed. Search as we knew it is done, largely because clicking through to evaluate result relevance before diving deeper sucks. Now we have agents that can do that portion and perform multiple searches, building on information in the last batch, to collect good results
amatecha 1 hours ago [-]
Yeah, there are times where DDG has like, literally three results. Yet, I know for an absolute fact, there are hundreds of pages on the web that contain the terms I specified. Web search is becoming utter garbage.
slig 39 minutes ago [-]
Try brave search.
CommieBobDole 19 minutes ago [-]
The article touches on something that I've been thinking about with regards to Google's AI strategy; the automatically-generated AI search summaries are not great. They very frequently confidently misinterpret what the user is searching for and generate half a page of useless information that pushes actual results down the page, and they are occasionally hilariously incorrect, with hallucinated facts.
This is probably a difficult-to-solve problem; given that they generate billions of these a day, not even Google can afford to devote enough compute to each query to reliably generate quality results. You can see this by selecting the "AI mode" from the search interface after getting the mediocre summary - the results are much better and generally perfectly usable. Though even that is probably a special minimal-compute version of the lowest tier of Gemini, it's still maybe an order of magnitude more capable than whatever generates the search summaries.
The bigger problem is that these search summaries are the default and by far the most common interaction that the general public has with "AI", and because this experience sucks, they just assume that all LLMs are similarly stupid and mostly useless. In non-technical spaces I frequently see the argument that "AI" is not useful for anything, all it generates is garbage hallucinations, and almost invariably they cite some actual terrible experience with the Google AI search summary. I would argue that the strategy of adding LLM summaries to every search is the worst of both worlds - it makes classic search worse while poisoning users against the idea of actual LLM-assisted search.
umvi 2 hours ago [-]
I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI.
There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.
jessetemp 2 hours ago [-]
All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias
hadlock 2 hours ago [-]
I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector
dylan604 1 hours ago [-]
Otherwise they'd be slurping in their own slop
satvikpendem 20 minutes ago [-]
This already exists, there are archives of Reddit or other sites, and Anna's Archive for papers and books.
asawfofor 42 minutes ago [-]
Isn’t this what the paper-bound encyclopedia companies do, albeit shallowly
anigbrowl 59 minutes ago [-]
Overbroad claim. Dramatic corollary
This clickbaity headline format cannot die fast enough
stdatomic 40 minutes ago [-]
It can't. People will simply stop clicking links.
ElProlactin 2 hours ago [-]
There's too much going on in this article and while I think some of the points are valid, others go too far.
A lot of the "cultural record" the author refers to is just digital junk. Random digital content that very few people care about, if we're being honest. Trying to hoard every bit of digital information ever produced is not the same thing as preserving "culture".
Case in point:
> Even the increasing use of ephemeral formats like Instagram Stories and WhatsApp status updates means that large portions of cultural, social, and political communication are never conserved in the first place. As a society, we can probably survive bad search results and come up with another way to schedule a sunset make-out session. But we can’t aspire to sovereignty if we can’t retain and retrieve our collective memory.
For most of human history, nobody was trying to "conserve" every cultural, social or political communication ever produced, and I fail to see how Instagram Stories and WhatsApp status updates, many of which aren't even truly broadcast publicly for all to see, are part of some imaginary "collective memory."
If you find a web page, see an Instagram Story or receive a message that's important to you, save it or take a screenshot. But let's not pretend all these things belong in a global Digital Civilizational Archives.
jbm 1 hours ago [-]
While I dislike Instagram reels and stories, I do feel like there is almost certainly scientific studies on radicalization that would benefit from actual histories of dumb memes. I feel like that about a lot of things, really.
There probably are some important hidden discord groups that would explain the origin of many political positions. Unlike smokey meetings in scummy bars, that exists now and is on a database somewhere.
bitwize 1 hours ago [-]
On the other hand, when modern archaeologists discover "Claudius has a small dick" graffiti on the side of some God-forsaken wall in Pompeii, they're fascinated. The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid. What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
ElProlactin 55 minutes ago [-]
> The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid.
It does, but do you think that people at that time thought anywhere near as much about preserving their scribbles as we do?
I'd venture a guess that we've created more "content" since the advent of the internet than in all of human history prior, and most of it is stored on things that aren't even designed to last a human lifetime without failure.
The idea that we're going to save every piece of digital junk for posterity just isn't realistic or healthy.
> What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
You're right, but you're also assuming that they're going to care that much, and that we're going to survive that long.
queenkjuul 29 minutes ago [-]
I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
ElProlactin 20 minutes ago [-]
> I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
Well as far as digital letters, photos, menus, etc. are concerned, there's nothing stopping someone with the means and motivation from investing in a doomsday storage facility specifically designed to store these things for posterity. If people can do this for crypto they can do it for digital content.
As for physical items, do you know how much junk Americans have in storage units? The US self-storage industry generates over $40 billion in annual revenue. We're probably keeping more "stuff" in storage units where it has a chance of surviving a zombie apocalypse than at any point in human history.
qudat 2 hours ago [-]
Funny as I just cancelled my Kagi sub to get the Gemini ai pro sub. The deal was too good to pass up
amatecha 1 hours ago [-]
The deal will be great for now, while they lure you in and get you to drop the competition -- then they will raise prices later.
satvikpendem 19 minutes ago [-]
Then you cancel and go to another provider, rinse and repeat. This is already what is happening with streaming platforms.
AlexandrB 54 minutes ago [-]
This used to be called "dumping" or predatory pricing[1] and would be fined in the physical retail space. Too bad our regulators are asleep at the wheel.
Kagi does have AI too, for what it’s worth. I found it pretty damn good, but I wish I could give them more money. I have a year subscription so I’m stuck without being able to give them *anything* until it’s done.
dylan604 1 hours ago [-]
Sign up for another account?
t-writescode 1 hours ago [-]
That is an option for sure, but a very disappointing one. And it comes with a meaningful danger of ending up with 2 subscriptions: one yearly and one monthly. They really, really need to finish their “pay as you go” system. I don’t know why it’s not done yet and there’s been no word about it to my knowledge.
dylan604 1 hours ago [-]
Maybe they should just enable a tip option? What do donations do to a company's tax liability and would it be worth their effort to enable something like that?
t-writescode 33 minutes ago [-]
They already have a system for “prepaying”, but it can’t be used to put more credits into your account, they just sit there, waiting to be used by a future subscription.
flinux 1 hours ago [-]
I see only one simple solution (though we should discuss the more complex ones): if Google directly answers a search query, then it must be held accountable for it, for better or for worse, and therefore assume all the benefits and (legal) liabilities that this entails
noosphr 1 hours ago [-]
Gemini's highest tier of plan has the utility of Google search circa 2010. I'm paying $400 for the privilege.
schaefer 2 hours ago [-]
Kagi search today is better than Google search ever was.
And it’s clear that Google’s Ad model ultimately created a priority inversion. The advertisers became the customer.
I am so glad Kagi came along with a business model that is actually working.
onemoresoop 2 hours ago [-]
I was wondering how would Kagi scale/expand if all of a sudden google were to stop serving search altogether (not likely) or alter search such that users look for alternatives.
Skunkleton 1 hours ago [-]
Kagi is an aggregator for other, some paid, search APIs. They have, at least in the past, served some percentage of their results from Bing's API among others for example. Kagi seems to me to be dependent on these APIs being available, if they were to go away, so would Kagi.
I am a happy subscriber of Kagi though, they provide a really excellent service.
jeorb 1 hours ago [-]
I'm a long time Kagi user and I haven't used Google search in about a year.
I tried out Google search for a few technical searches recently and it was surprisingly ad and AI free. Not bad at all and much better than I remember from last year.
Then I put in some non-technical searches and it was all ads and AI and basically unusable.
bigstrat2003 1 hours ago [-]
I wouldn't say it's better, but it's certainly on par with Google in their best years. And it's light years better than what Google is now, or using an LLM.
georgemcbay 2 hours ago [-]
From a user perspective, Google search is the most useful it has been in years, though that doesn't feel entirely like intentional improvement, just a lucky side effect of the move to "AI mode".
And yes, if you take what the AI tells you at face value it could be wrong. But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
And also, yes, the old balance of Google driving clicks to sites that will then generate revenue off more Google Ads being shown after you click through to them creating a virtuous cycle is completely busted, and that sucks. It does not impact me directly but it certainly seems like unless a better system is devised that it is one of a few ways in which AI is likely to stall out its own training funnel.
bloaf 2 hours ago [-]
As someone who regularly reads things online, then wants to read them again like 3 years later, Google has been monotonically declining in quality.
Also I realized the other day how hard it is to find song lyrics for anything other than quite mainstream songs.
Terr_ 1 hours ago [-]
It used to be I could always find the quote I wanted from a book. Nowadays I need to keep my own copies...
> Hister is a private search engine for the pages you visit and the files you keep. It indexes their full contents so you can find information again from the web interface, terminal, or an AI assistant connected through MCP.
georgemcbay 2 hours ago [-]
> As someone who regularly reads things online, then wants to read them again like 3 years later, Google has been monotonically declining in quality.
I generally agree, but I think AI mode actually improved things somewhat compared to how things were just prior to it existing.
And I'm not saying what we have now is better than Golden Age Google, but things were just getting worse and worse for almost a decade. AI didn't fix the decade worth of decline, but it is the first thing I've seen from Google that at least partially reversed it for my own usage.
bloaf 2 hours ago [-]
I think they make things worse, because they very very often present straight inaccurate information.
Just the other day I was trying to find out "What american tree species have the deepest roots". And all the AI responses were giving me back generic lists of big trees and claiming that roots going 20ft deep were the deepest. I know for a fact the mesquite trees behind my house can easily grow roots > 100 ft deep.
If I had clicked on the articles with generic lists of big trees, I would have realized they were all low quality clickbait sources and moved on. But the AI presentation makes you think that the information comes well-researched.
Gualdrapo 2 hours ago [-]
> But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
The point is not about 'quicker' requests but precise requests. It definitely has worsened, though not on a single degree on al levels like the HN hivemind claims, but some aspects are still somewhat precise but others are definitely crap.
i.e. when searching about my neighborhood it still returns better results than bing, yahoo, ddg, yandex and what have you. But they are buried into a load of crap of alleged "relevant" results (those things past the ai stuff) that aren't relevant in any way.
bloaf 2 hours ago [-]
Yandex is the only search engine left which still feels like the "old" web. It feels like you're actually getting a best effort search, and not just the results that someone paid to put in front of you.
peyton 1 hours ago [-]
What’s come next for me has been much better. I use ChatGPT cranked to Pro with “extended” thinking to one-shot whatever I would’ve spent time looking into with Google. It’ll plan the whole sunset bike ride or promposal or whatever from TFA.
amazingamazing 2 hours ago [-]
[2011] The Atlantic - Why Google Won't Survive the Facebook Threat.
I’ll add this article to the list of incorrect predictions lol
charcircuit 2 hours ago [-]
It would be beneficial for the author to look at the revenue for Google Search. Revenue is still growing.
panarky 2 hours ago [-]
Indeed.
Not only is search revenue growing, but it is growing at an accelerating rate.
At the same time, operating margins are expanding.
I don't know the name of the logical fallacy where someone personally uses an LLM instead of Google Search and then infers that the search business is dying, without ever reading a financial statement.
Chinjut 2 hours ago [-]
Indeed. Money over everything. It is absurd to complain about decreasing quality of a product that makes increasing money for increasingly rich people (thus, by definition, better).
carlosjobim 57 minutes ago [-]
If business owners are paying for ads, then it doesn't matter a iota to Google or Facebook if real people use their services. It would be even better for them if real people didn't use their services, since that would save some costs. Business owners are going to keep paying for the ads, as long as they get some number about how many (bot) impressions their ad generated.
echelon 2 hours ago [-]
There is probably a lag time before advertisers give up on AdSense.
A lot of corporate ad spend is already planned, and Google can adjust the costs up as much as they like. They hold the lever.
charcircuit 2 hours ago [-]
If search engine competitors were really eating Google's lunch then ad impressions would go down.
forgetfreeman 2 hours ago [-]
The most frustrating part of all of this is the underlying premise that the internet is, has been, or could ever be a credible cultural record is deeply stupid. Or maybe more charitably it's both historically and technically illiterate. It has always taken continuous unwavering effort on some person's part to keep any given piece of content online. And while managing a simple hosting account and updating domain registration periodically doesn't take a tremendous amount of effort 20 years is a long time to expect anyone to maintain enthusiasm. The internet has always been a frothy, ever changing blend of the odd nugget of truth drifting in a sea of unadulterated bullshit. Treating this, or worse what comes from statistically averaging it, as a source of capital T truth is totally unhinged. From whence did this mythology of online truth spring?
danpalmer 2 hours ago [-]
It's hard to get past the beginning of this article and take it at all seriously. The quote someone who missed a sunset because they asked google and supposedly got the wrong time... but they didn't think to look at the sun or lack thereof to check? Also when I put "when does the sun set today" I get a single exact figure at the top of my results, not from AI, which is honestly the best kind of result – an exact correct answer.
tayo42 2 hours ago [-]
They were planning their day for a specific time to be somewhere. If you're looking at the sun it's to late...
BeetleB 2 hours ago [-]
That's fine, but arguing that Google should give accurate times based on where you are is ... crazy?
In the pre-LLM days it was cool that Google did weather, unit conversions, sports results, etc. But that's not even close to their value proposition. Even 5 years ago if someone told me they had planned a photograph and got it wrong because Google gave them the wrong time for sunset, I would have called them a moron for relying on Google! There are sites and apps dedicated to this. Use one of them!
Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.
If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.
I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
Are you actually saying you'd be OK with that?
YMMV.
I end up using the gemini api for with search enabled for the cases that I don't have access to good grounding data even in agentic tasks.
No. The court specifically determined that the Internet Archive was guilty of unauthorized copying. It was not simply an unfounded or unproven allegation by publishers, book authors, and writers. The Authors Guild, the National Writers Union, the European Writers Council, and the Society of Authors in the UK all came out against the Internet Archive, and supported the suit.
news organizations started blocking the Wayback Machine’s crawlers out of fear that archived pages can provide AI companies with an indirect source of copyrighted material. Each new restriction limits the archive’s ability to act as a comprehensive backstop.
This self-inflicted damage to the wayback machine is the real tragedy of this entire affair. When IA was asked to stop CDL - many times - founder Brewster Kahle continued. The National Writers Union tried to open a dialogue as early as 2010 but was ignored:
The Internet Archive says it would rather talk with writers individually than talk to the NWU or other writers’ organizations. But requests by NWU members to talk to or meet with the Internet Archive have been ignored or rebuffed.
https://nwu.org/nwu-denounces-cdl/
When the requests to abandon CDL turned into demands, Kahle dug in his heels. When the inevitable lawsuits followed, and IA lost, he insisted that he was still in the right and plowed ahead with appeals. And here we are today.
Though I can find its AI answers annoying aggressive. I'll look up like two search terms and the AI will bullshit multiple paragraphs out of despite having zero context of what I am looking for.
DuckDuckGo seems to have detection of whether it should give an AI answer. And it allows you to have more granular control of when you want to get an AI answer. And is overall less distracting than Google's.
I miss when Google was like a grep for the entire visible Internet. Now it tries to second-guess my search and direct me to a bunch of sites which all have identical information that isn't what I'm looking for.
i don’t agree with the google has better results thing. sometimes it does. most of the time it’s just that google has the site i want higher in the ordering than DDG. personally i’m fine scrolling down a little bit more. it’s rare i need to go to google for something that DDG doesn’t have at all in their results, but it does happen.
i do have to go to google for maps/directions/planning travel. a lot that’s annoying.
or you can press the gear button -> "Ai features: Manage" -> Search assist
I've stopped using DDG now because of result quality. I now use a "meta" search backed by EXA, Tavily, and SearXNG in parallel. It can be agentically de-dupped or summarized as needed. Search as we knew it is done, largely because clicking through to evaluate result relevance before diving deeper sucks. Now we have agents that can do that portion and perform multiple searches, building on information in the last batch, to collect good results
This is probably a difficult-to-solve problem; given that they generate billions of these a day, not even Google can afford to devote enough compute to each query to reliably generate quality results. You can see this by selecting the "AI mode" from the search interface after getting the mediocre summary - the results are much better and generally perfectly usable. Though even that is probably a special minimal-compute version of the lowest tier of Gemini, it's still maybe an order of magnitude more capable than whatever generates the search summaries.
The bigger problem is that these search summaries are the default and by far the most common interaction that the general public has with "AI", and because this experience sucks, they just assume that all LLMs are similarly stupid and mostly useless. In non-technical spaces I frequently see the argument that "AI" is not useful for anything, all it generates is garbage hallucinations, and almost invariably they cite some actual terrible experience with the Google AI search summary. I would argue that the strategy of adding LLM summaries to every search is the worst of both worlds - it makes classic search worse while poisoning users against the idea of actual LLM-assisted search.
There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.
This clickbaity headline format cannot die fast enough
A lot of the "cultural record" the author refers to is just digital junk. Random digital content that very few people care about, if we're being honest. Trying to hoard every bit of digital information ever produced is not the same thing as preserving "culture".
Case in point:
> Even the increasing use of ephemeral formats like Instagram Stories and WhatsApp status updates means that large portions of cultural, social, and political communication are never conserved in the first place. As a society, we can probably survive bad search results and come up with another way to schedule a sunset make-out session. But we can’t aspire to sovereignty if we can’t retain and retrieve our collective memory.
For most of human history, nobody was trying to "conserve" every cultural, social or political communication ever produced, and I fail to see how Instagram Stories and WhatsApp status updates, many of which aren't even truly broadcast publicly for all to see, are part of some imaginary "collective memory."
If you find a web page, see an Instagram Story or receive a message that's important to you, save it or take a screenshot. But let's not pretend all these things belong in a global Digital Civilizational Archives.
There probably are some important hidden discord groups that would explain the origin of many political positions. Unlike smokey meetings in scummy bars, that exists now and is on a database somewhere.
It does, but do you think that people at that time thought anywhere near as much about preserving their scribbles as we do?
I'd venture a guess that we've created more "content" since the advent of the internet than in all of human history prior, and most of it is stored on things that aren't even designed to last a human lifetime without failure.
The idea that we're going to save every piece of digital junk for posterity just isn't realistic or healthy.
> What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
You're right, but you're also assuming that they're going to care that much, and that we're going to survive that long.
Well as far as digital letters, photos, menus, etc. are concerned, there's nothing stopping someone with the means and motivation from investing in a doomsday storage facility specifically designed to store these things for posterity. If people can do this for crypto they can do it for digital content.
As for physical items, do you know how much junk Americans have in storage units? The US self-storage industry generates over $40 billion in annual revenue. We're probably keeping more "stuff" in storage units where it has a chance of surviving a zombie apocalypse than at any point in human history.
[1] https://en.wikipedia.org/wiki/Predatory_pricing
And it’s clear that Google’s Ad model ultimately created a priority inversion. The advertisers became the customer.
I am so glad Kagi came along with a business model that is actually working.
I am a happy subscriber of Kagi though, they provide a really excellent service.
I tried out Google search for a few technical searches recently and it was surprisingly ad and AI free. Not bad at all and much better than I remember from last year.
Then I put in some non-technical searches and it was all ads and AI and basically unusable.
And yes, if you take what the AI tells you at face value it could be wrong. But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
And also, yes, the old balance of Google driving clicks to sites that will then generate revenue off more Google Ads being shown after you click through to them creating a virtuous cycle is completely busted, and that sucks. It does not impact me directly but it certainly seems like unless a better system is devised that it is one of a few ways in which AI is likely to stall out its own training funnel.
Also I realized the other day how hard it is to find song lyrics for anything other than quite mainstream songs.
https://github.com/asciimoo/hister
> Hister is a private search engine for the pages you visit and the files you keep. It indexes their full contents so you can find information again from the web interface, terminal, or an AI assistant connected through MCP.
I generally agree, but I think AI mode actually improved things somewhat compared to how things were just prior to it existing.
And I'm not saying what we have now is better than Golden Age Google, but things were just getting worse and worse for almost a decade. AI didn't fix the decade worth of decline, but it is the first thing I've seen from Google that at least partially reversed it for my own usage.
Just the other day I was trying to find out "What american tree species have the deepest roots". And all the AI responses were giving me back generic lists of big trees and claiming that roots going 20ft deep were the deepest. I know for a fact the mesquite trees behind my house can easily grow roots > 100 ft deep.
If I had clicked on the articles with generic lists of big trees, I would have realized they were all low quality clickbait sources and moved on. But the AI presentation makes you think that the information comes well-researched.
The point is not about 'quicker' requests but precise requests. It definitely has worsened, though not on a single degree on al levels like the HN hivemind claims, but some aspects are still somewhat precise but others are definitely crap.
i.e. when searching about my neighborhood it still returns better results than bing, yahoo, ddg, yandex and what have you. But they are buried into a load of crap of alleged "relevant" results (those things past the ai stuff) that aren't relevant in any way.
I’ll add this article to the list of incorrect predictions lol
Not only is search revenue growing, but it is growing at an accelerating rate.
At the same time, operating margins are expanding.
I don't know the name of the logical fallacy where someone personally uses an LLM instead of Google Search and then infers that the search business is dying, without ever reading a financial statement.
A lot of corporate ad spend is already planned, and Google can adjust the costs up as much as they like. They hold the lever.
In the pre-LLM days it was cool that Google did weather, unit conversions, sports results, etc. But that's not even close to their value proposition. Even 5 years ago if someone told me they had planned a photograph and got it wrong because Google gave them the wrong time for sunset, I would have called them a moron for relying on Google! There are sites and apps dedicated to this. Use one of them!