Ai generated content is inherently a regression to the mean and harms both training and human utility. There is no benefit in publishing anything that an AI can generate, just ask the question yourself. Maybe publish all AI content with <AI generated content> tags, but other than that it is a public nuisance much more often than a public good.
Following this logic, why write anything at all? Shakespeare's sonnets are arrangements of existing words that were possible before he wrote them. Every mathematical proof, novel, piece of journalism is simply a configuration of symbols that existed in the space of all possible configurations. The fact that something could be generated doesn't negate its value when it is generated for a specific purpose, context, and audience.
Invented might be a bit strong, but he is certainly the first written record of the word. Dress existed as a verb already, as did the generic reversing “un”, but before Shakespeare there is no evidence that they were used this way. Prior to that other words/phrases, which probably still exist in use today, were used instead. Perhaps “disrobe” though the OED lists the first reference to that as only a decade before Taming Of The Shrew (the first written use of undress) was published, so there are presumably other options that were in common use before both.
It is definitely valid to say he popularised the use of the word, which may have been being used informally in small pockets for some time before.
Following that logic, we should publish all unique random orderings of words. I think there is a book about a library like that, but it is a great read and is not a regression to the mean of ideas.
Writing worth reading as a non-child surprises, challenges, teaches, and inspires. LLM writing tends towards the least surprising, worn out tropes that challenge only the patience and attention of the reader. The eager learner, however will tolerate that , so I suppose that I’ll give them teaching. They are great at children’s stories, where the goal is to rehearse and introduce tropes and moral lessons with archetypes, effectively teaching the listener the language of story.
FWIW I am not particularly a critic of AI and am engaged in AI related projects. I am quite sure that the breakthrough with transformer architecture will lead to the third industrial revolution, for better or for worse.
But there are some things we shouldn’t be using LLMs for.
Just like googling, AIing is a skill. You have to know how to evaluate and judge AI responses. Even how to ask the right questions.
Especially asking the right questions is harder than people realize. You see this difference in human managers where some are able to get good results and others aren’t, even when given the same underlying team.
No, new more-capable and/or efficient models have been forged using bulk outputs of other models as training data.
These inproved models do some valuable things better & cheaper than the models, or ensembles of models, that generated their training data. So you could not "just ask" the upstream models. The benefits emerge from further bulk training on well-selected synthetic data from the upstream models.
Yes, it's counterintuitive! That's why it's worth paying attention to, & describing accurately, rather than remaining stuck repeating obsolete folk misunderstandings.
Less than you might think! Some of the frontier-advancing training-on-model-outputs ('synthetic data') work just uses other models & automated-checkers to select suitable prompts and desirable subsets of generations.
I find it (very) vaguely like how a person can improve at a sport or an instrument without an expert guiding them through every step up, just by drilling certain behaviors in an adequately-proper way. Training on synthetic data somehow seems to extract a similar iterative improvement in certain directions, without requiring any more natural data. It's somehow succeeding in using more compute to refine yet more value from the original non-synthetic-training-data's entropy.
The training sets can already include direct data series about the world, where the "work of human beings" is just setting up the the collection devices. So models can absolutely "experience the world".
But I'm not suggesting they'll advance much, in the near term, without any human-authored training data.
I'm just pointing out the cold hard fact that lots of recent breakthroughs came via training on synthetic data - text prompted by, generated by, & selected by other AI models.
That practice has now generated a bunch of notable wins in model capabilities – contra the upthread post's sweeping & confident wrongness alleging "Ai generated content is inherently a regression to the mean and harms both training and human utility".
How does the banana bread taste at the café around the corner? What's the vibe like there? Is it a good place for people-watching?
What's the typical processing time for a family reunion visa in Berlin? What are the odds your case worker will speak English? Do they still accept English-language documents or do they require a certified translation?
Is the Uzbek-Tajik border crossing still closed? Do foreigners need to go all the way to the northern crossing? Is the Pamir highway doable on a bicycle? How does bribery typically work there? Are people nice?
The world is so much more than the data you have about it.
Of course, training on synthetic data can't do everything! My main point is: it's been doing a bunch of surprisingly-beneficial things, contra the obsolete beliefs about model-output-worthlessness (or deleteriousness!) for further training to which I was initially responding.
But also: with regard to claims about what models "can't experience", such claims are pretty contingent on transient conditions, and expiring fast.
To your examples: despite their variety, most if not all could soon have useful answers answers collected by largely-automated processes.
People will comment publicly about the "vibe" & "people-watching" – or it'll be estimable from their shared photos. (Or even: personally-archived life-stream data.) People will describe the banana bread taste to each other, in ways that may also be shared with AI models.
Official info on policies, processing time, and staffing may already be public records with required availability; recent revisions & practical variances will often be a matter of public discussion.
To the extent all your examples are questions expressed in natural-language text, they will quite often be asked, and answered, in places where third parties – humans and AI models – can learn the answers.
Wearable devices, too, will keep shrinking the gap between things any human is able to see/hear (and maybe even feel/taste/smell) and that which will be logged digitally for wider consultation.
> data series about the world, where the "work of human beings" is just setting up the the collection devices. So models can absolutely "experience the world"
But not experience it the way humans do.
We don’t experience a data series; we experience sensory input in a complicated, nuanced way, modified by prior experiences and emotions, etc. remember that qualia is subjective, with a biological underpinning.
Sure, and there are many such writings that can be useful. No denying. But the LLM cannot experience like humans do and so will forever be outside our circle. Whether it also remains outside our circle of empathy, or us outside of its, remains to be discovered.
One example of useful output does not negate the flood of pollution. I’m not denying or downplaying the usefulness of AI. I am doubting the wisdom of blindly publishing -anything- without making at least a trivial attempt to ensure that it is useful and worth publishing. It is a form of pollution.
The problem is that it lowers the effort required to produce SEO spam and to “publish” to nearly zero, which creates a perverse incentive to shit on the sidewalk.
The amount of AI created, blatantly false blog posts about drug interactions, for example. Not advertising, just banal filler to create site visits, with dangerously false information.
It’s not like shitting on the sidewalk was never a problem before, it’s just that shitting on the sidewalk as a service (SOTSAAS) maybe is something we should try to avoid.
I didn’t mean to imply that -no- ai generated content is useful, only that the vast, vast majority is pollution. The problem is that it is so cheap to produce garbage content with AI that writing actual content is disincentivized, and doing web searches has become an exercise is sifting through AI generated slop.
That at least will add extra work to filter usable training data, and costs users minutes a day wading through the refuse.
So IMHO an right thing is to add "AI rewritten" label to your blog.
hm.. I wonder where this kind of label should live? For a personal blog, putting it on every post seems redundant, as if author uses it, it's likely they use it for all posts. And many blogs don't have dedicated "about this blog" section.
I wonder if things will end up like organic food labeling or "made in .." labels. Some blogs might say "100% by human", some might say "Designed by human, made by AI" and some might just say nothing.
In context of writing text, keyboard and text editor are inanimate tools because they cannot introduce text user did not come up with.
Spellcheck and autocorrect can come up with new words, and so is often anthropomorphized, it's not 100% "inanimate tool" anymore.
AI can form its own sentences and come up with its own facts for a much greater degree, so I would not call it "inanimate tool" at all (again, in context of writing text). It is much closer to editor-for-hire or copywriter-for-hire, and I think it should be treated the same as far as attribution goes.
hm.. looks like I am convincing myself into your point :)
After all, if another human edits/proofreads my posts before publish, I don't need to disclose that on my post... So why should AI's editing be different?
If I ask the question myself then there's no step where a human expert has vetted the content and put their name on it. That curation and vouching is of value.
Now your mind might have immediately went "pffff as if they're doing that" and I agree but only to the extent that it largely wasn't happening prior to AI anyway. The vast majority of internet content was already low quality and rushed out by low paid writers who lacked expertise in what they were writing about. AI doesn't change that.
Completely agree. We are used to thinking of authorship as the critical step. We're going to have to adjust to thinking of publication as the critical step. In an ideal world, publication of a piece would be seen as vouching for that piece. Putting your reputation on the line.
I wonder if we'll see a resurgence in reputation systems (probably not).
Yes, and deep research was junk for the hard topics that I actually needed to sit down and research. Anything shallower I can usually reach by search engine use and scan; deep research saves me about 15-30 minutes for well-covered topics.
For the hard topics, the solution is still the same as pre-AI - search for popular survey papers, then start crawling through the citation network and keeping notes. The LLM output had no idea of what was actually impactful vs what was a junk paper in the niche topic I was interested in so I had no other alternative than quality time with Google Scholar.
We are a long way from deep research even approaching a well-written survey paper written by grad student sweat and tears.
I've found getting a personalized report for the basic stuff is incredibly useful. Maybe you're a world class researcher if it only saves you 15-30 minutes, I'm positive it has saved me many hours.
Grad students aren't an inexhaustible resource. Getting a report that's 80% as good in a few minutes for a few dollars is worth it for me.
Steel-man angle: A desire for data provenance is a good thing with benefits that are independent of utopias/humans vs machines kinds of questions.
But, all provenance systems are gamed. I predict the most reliable methods will be cumbersome and not widespread, thus covering little actual content. The easily-gamed systems will be in widespread use, embedded in social media apps, etc.
Questions:
1. Does there exist a data provenance system that is both easy to use and reliable "enough" (for some sufficient definition of "enough")? Can we do bcrypt-style more-bits=more-security and trade time for security?
2. Is there enough of an incentive for the major tech companies to push adoption of such a system? How could this play out?
Yes, but GP's idea of segregating AI-generated content is worth considering.
If you're training an AI, do you want it to get trained on other AIs' output? That might be interesting actually, but I think you might then want to have both, an AI trained on everything, and another trained on everything except other AIs' output. So perhaps an HTML tag for indicating "this is AI-generated" might be a good idea.
My 2c is that it is worthwhile to train on AI generated content that has obtained some level of human approval or interest, as a form of extended RLHF loop.
Ok, but how do you denote that approval? What if you partially approve of that content? ("Overall this is correct, but this little nugget is hallucinated.")
It apparently doesn't matter unless you somehow consider the entire Internet to be correct. They didn't only feed LLMs correct info. It all just got shoveled in and here we are.
I can see the value of labeling all AI can be trained on purely non-AI generated content.
But I don’t think that’s a reasonable goal. Pragmatic example: There’s almost no optional HTML tags or optional HTTP Headers which are used anywhere close to 100% of the times they apply.
Also, I think field is already muddy, even before the game starts. Spell checker, grammar.ly, and translation all had AI contributions and likely affect most of human-generated text on the internet. The heuristic of “one drop of AI” is not useful. And any heuristic more complicated than “one drop” introduces too much subjective complexity for a Boolean data type.
Yes, it's impossible. We'd have to have started years ago. And then people wouldn't have the discipline to label content correctly or at all. It can't be done.
Just ask any person who works in teaching or any of the numerous faulty AI detectors (they're all faulty).
Any current technology which can used to accurately detect pre-AI content would necessarily imply that that same technology could be used to train an AI to generate content that could skirt by the AI detector. Sure, there is going to be a lag time, but eventually we will run out of non-AI content.
No, that's the problem. Pre-AI era content a) is often not dated, so not identifiable as such, and b) also gets out of date. What was thought to be true 20 years ago might not be thought to be true today. Search for the "half-life of facts".
The observation that humans poop is not sufficient justification for spending millions of dollars building an automated firehose that pumps a torrent of shit onto the public square.
I make no claim to the overall value of LLMs. I'm just pointing out that your analogy is a fallacy. The fact that group A does a small bad thing is not a justification for allowing group B to do a large bad thing. That is true regardless of whether group B does there non-bad things.
It may be the case that the non-bad things B does outweigh the bad things. That would be an argument in favor of B. The another group doing bad things has no bearing on the justification for B itself.
From my experience the people spending "millions" are hoping they get those millions * 10 back because a buddy of theirs told them "this AI thing" is going to replace the most expensive part of companies, the staff costs, not because they think the product is any good. We're getting AI forced down our throat because VC is throwing cash in like there's no tomorrow, not because of whatever value might or might not be there.