Progress is progress, and it's good to see tech like AI used for good but I don't like the idea of these proprietary models being used in profit making endeavors which is what I assume will happen next.
I'm wondering why they have restricted file types. You can't check a PDF for example... surely the main use case for people will be to check if a document was produced or edited by an LLM? That could be an attractive (if not misunderstood) proposition for academics
> I'm wondering why they have restricted file types. You can't check a PDF for example...
TFA/page actually seems incomplete. Text uses a completely different watermark format (an actual watermark as opposed to a provenance/authenticity signature), so it makes sense to me that they're not claiming to be able to scan PDFs when they can't yet incorporate that signal.
On a linked page, they say:
> Watermark detection is currently in private preview [...]
Luckily all pdfs I made in academia have been generated from source (tex or derivatives, asciidoc) and llm are better in generating source than pdf. Even people that didn’t use tex used Word to generate the pdf.
Source can have markers inserted in it as weighted word choice or phrase choices, how commas are inserted, etc., so that the output can still can be identified. In other words just because it's source doesn't mean it can be watermarked.
Do you mind me asking why you have no fear of OpenAI etc publishing a biology paper? With increasing model capability and compatibility with lab hardware could we not be in a scenario soon(ish) where these agents are able to autonomously complete and publish experimental results?
I was debating this with a friend the other day and the consensus we came to was that a highly trained scientist would (or should) always review output like that described above, but that's starting to feel like a weakening argument!
Firstly the underlying worry here is about privacy which hinges on the fact that AI companies are stealing ideas in the first place. Stealing from your customers is an incredibly bad business model and I think if they were to steal IP (intellectual property) from researchers mathematicians or computer scientists would be first.
Now why I think biology is safer:
1) Producing novel biology still has to be done in a lab. It requires laboratories, equipment, experimental protocols, trained personnel, regulatory and safety infrastructure, and often substantial institutional organization all of which there is no indication they're heading for. Also I disagree that lab hardware is near a "soon state" where labs can be full autonomous, (liquid handlers are really good at niche tasks but lack any type of experimental general ability [not AI-bounded], especially for in vivo work where its footprint is non-existent). Even the most automated Labs I know where robots do 80% of experimental work, they still have grad students to carry out that last 20% and to oversee.
2) Even if AI could do the pipeline it's not worth it for AI LLM companies to dedicate capital to it currently. A lot of biology research itself doesn't produce a sellable product, in fact most of it never does. It seems currently and for at least the next couple years at least, AI capital is best spent growing compute to research better models, train better models, and sell inference.
reply