Hacker Newsnew | past | comments | ask | show | jobs | submit | dhx's commentslogin

Great in theory, but what are US enterprises going to do _if_ their private data is later found to be used for training?

1. Not use AI technology and fall behind the rest of the world.

2. Use Chinese AI technology, either hosted by Chinese companies or the models self-hosted.

3. Sue US AI companies for damages, but not enough to have any meaningful impact to such companies that it'd impact US national security goals (per US government contribution to NY Times copyright lawsuit).


There's a huge difference between AI companies exploiting a grey area like training on public corpora and violating a private contract that they explicitly entered into with another party. The latter is very explicitly illegal and would never survive trial in Delaware Chancery court. And all of that is before we get into Federal contracts where training on TS/SCI data could lead to criminal charges.

There's a huge market in the US for providing AI services while respecting client privacy. It makes sense for at least one major provider to offer this.


For (C) -- I think it could additional create jobs funded by government, philanthropic and other private institutions. For example, a government funded museum may already be participating in Wikimedia GLAM projects (e.g. uploading historical images to Wikimedia Commons with complete metadata). Perhaps this type of open source contribution may increase if organisations realise their mission can be better accomplished by contributing this same open data into LLMs, in addition to Wikimedia Commons. If the museum's mission is to educate the public on the history of life in ACMEville, having LLMs be able to provide historical information and images to a prompt of "What is the history of ACMEville?" may be a good pursuit.

I'm sceptical though whether use of LLMs would encourage creation of data that doesn't already exist. For example, if you ask an LLM "What are the top 100 most prevalent flora endemic to ACME National Park", this data may not currently exist _at all_, and to collect, would require paying botanists to do an extensive field survey. If no one has done this work yet--why? Is it relevant to the scientific community, to making government decisions, etc, or just an obscure academic curiosity. There are certainly some journal articles on _other_ national parks describing some of their common endemic flora, but perhaps there was a reason for this. Such as a scientist funded by a one-off government program trying to determine how to preserve or even create habitat for a specific endangered species.

Consider for the prompt of: "What are the top 100 most prevalent flora endemic to ACME National Park"

An LLM may reply: "I couldn't find any journal article or other prior work that may answer this question. Typically such survey field work may cost $X to complete, require expert botanists, and take 6-12 months to complete. Let me know if you want further information on how to find and select a company to conduct such a botanical field survey."

Would this type of LLM response grow the industry of botanical field surveys, or do nothing, perhaps because anyone likely to fund botanical field surveys is already doing so regardless of whatever is happening with AI.


100% agreed

It'd be great to see a description of even just a subset of training datasets. It feels very much under-reported how much expense is worth investing in preparing and selecting training datasets versus just using masses of random quality unprepared training data. This dashboard appears to be good though in showing the limits quickly reached when throwing parameters and compute at the problem.

For example, if they were to train on Wikipedia dumps, do they consider every article to be the same quality across each language, or have they done more work beyond Wikipedia's own article quality ratings to make training decisions such as "Ignore cebwiki it's machine-generated spam" and "Treat dewiki articles with coordinates within Germany as being higher quality (weight it higher) than their equivalent enwiki articles".

And let's say one of the datasets is all the source code of packages in the Gentoo package repository. Not every software package is a good example of how to write code. You perhaps wouldn't want to train your LLM on 1990s era PHP web application source code as an example of how to write code in 2026. Instead, you'd possibly want to use such PHP web application source code as a negative training example of what _not_ to write. But when training an LLM to detect software bugs, maybe outdated PHP source code is good for training.

Similarly for translation, perhaps UN treaty documents translated into 4+ languages are good translation examples because of high accuracy needed, professional translators being used, and bigger budgets. However this training data would perhaps be a negative training example towards translating chat messages, movie subtitles, etc because it doesn't use everyday slang and could result in output of nonsense such as "Pending Your Excellency's response, please accept, Your Excellency, my sincere greetings." for a prompt asking to write a birthday card for a child.

Preparing training data and deciding how to best use it for training I assume would be the largest expense (cost of labour -- mostly expert labour too) and also the greatest opportunity in the future for LLMs to improve. It seems to me somewhat irrelevant if the dashboard indicates a compute expense of $1m or $5m if good training datasets (prepared by experts in their fields) cost $10m/y to maintain. For example, hiring expert software developers to tag 1000's of open source software packages according to their quality, on different metrics, such as human readability, performance optimisation with choice of algorithms, reasonable trade-off between coherence and coupling in the software architecture, currency with state of the art programming trends/preferred dependencies/operating system APIs, etc. And keeping that metadata continually updated rather than a rapidly obsolete once off tagging project completed in 2005.


In the case of Assange, Australian politicians did a lot of work to get him released in the end[1], including:

* Sending a delegation representing all major political parties to the US to argue for the release of Assange. Imagine picking the Republican and Democrat politician LEAST likely to want to cooperate on anything, and those two would have been Australia's equivalent representatives in this delegation. Reports afterwards of the meeting at DOJ HQ indicated it wasn't the type of meeting where the Australians would have brought Tim Tams to share around the room.[2]

* The Australian parliament voted publicly 2:1 on a motion for Assange's release.

* Repeated petitioning through ambassadors in the UK and US, official visits of Australian politicians, etc. Not in private either, as is typically the case for diplomatic affairs.

* Australian politicians attending UK extradition hearings.

* After getting agreement to a plea deal, flying Australian ambassadors for the UK and US to the court of a one-pub-town in the middle of the Pacific Ocean no one has heard of (Northern Mariana Islands) in support of Assange, then all of them flying back to the Australian prime minister's aircraft terminal for a welcome home bevvy.

This was all at a time too where "Free Assange" posters and graffiti was _widely_ distributed across Australian cities.

[1] https://en.wikipedia.org/wiki/Julian_Assange#Plea_bargain_an...

[2] https://www.abc.net.au/news/2024-06-27/inside-the-us-austral...


> authors won’t publish if anyone can republish their work for free

There's plenty (even a majority?) of authors that publish and will continue to publish without any expectation of direct remuneration. Open source software developers and companies hiring such developers. Not-for-profit organisations increasing awareness of a cause. Private companies wanting to reach an audience for marketing reasons.[1] Government organisations. Researchers funded by government grants. Universities publishing books or coursework openly (they're in the business of selling their stamps on degrees, not selling books).

[1] Even includes the likes of Warner Music with CC-BY music videos on YouTube for some artists, seemingly for marketing reasons to try and build the name and following of a particular artist.


Again, incentives.

The people who publish are people who have reason to publish when they can be copied. Typically either they have already been paid, or they expect to gain market share by being free.

People who need renumeration to continue working will not publish.

Intellectual property rights, as much as I dislike the RIAA and MPAA, created a way for more players to enter the market, because it created a way for their needs to be met.


"Sweat of the brow" doctrine has been rejected in most countries.[1] Even Europe's Database Directive, probably the closest thing to an implementation of this doctrine, largely doesn't do much in practice.

An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown (this being one of the books in the series), 3500 beaches were visited and documented along 12000km of coastline.[2]

AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book).

If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including:

1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing?

2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing?

3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of.

"Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole.

[1] https://en.wikipedia.org/wiki/Sweat_of_the_brow

[2] https://sydneyuniversitypress.com/products/9781920898168

[3] https://en.wikipedia.org/wiki/Tragedy_of_the_anticommons


I wasn't aware of the "sweat of the brow" doctrine until you mentioned it. But you seem to be using it incorrectly. The doctrine only states that creativity or originality isn't required to make a work copyright-able.

Even if this were an accepted principle, that wouldn't change the principle of free use. In all of your examples, only re-printing all or substantial portions of the books of Andre D. Short would be copyright violations. Just referencing facts from Short's books, or even including small quotes, in your own new work is not a violation.


The parent comment I replied to is concerned with "life's work got appropriated without consideration, compensation or consent". To alleviate this concern^, "sweat of the brow" doctrine would be required, but it doesn't exist in most jurisdictions. Today in most jurisdictions copyright laws do not care the slightest about an LLM ingesting databases -- phone directories, sport fixtures and results, someone's life work measuring the dimensions of frogs, etc. 100% of the original factual data could be learned by the LLM, and 100% could be output all at once.

^ Of course there are other ways to alleviate the concerns too such as universal basic income, government grants, etc for someone who wants to dedicate their life to measuring the dimensions of frogs, or whatever else their interest may be. There would however be some geopolitical/trade issues involved--a population would have to be comfortable doing the heavy lifting only to have another country simply use the work freely and instead dedicate their lives to something less favourable such as building missiles.


> To alleviate this concern^, "sweat of the brow" doctrine would be required, but it doesn't exist in most jurisdictions

No, it wouldn't. "Sweat of the brow" applies to collections of facts whose compilation required effort. "Life's work" is a superset of that. Originality and creativity, which are required to copyright something, are also work.


The US government's official position on LLMs is (very simply paraphrased) that LLMs are sufficiently transformative and do not hamper the potential market of authors of training material, therefore, copyright claims arising from training material should not be successful.[1] For original and creative training material, for example, a Harry Potter novel, seemingly the US government is asking the courts to set aside some previous questionable findings such as copyright existing very loosely in the likeness of fictional characters (impacting the likes of fan fiction). Can a human -- or LLM -- create a story about children travelling on a train from New York to a school of magic in the "wild west", with many loose similarities to Harry Potter for those familiar with those books? The US government appears to be saying this is OK, especially with the view that the market for Harry Potter is not diminished by a "wild west magic school" book in its likeness.

However, LLMs do sometimes output training data almost 1:1 without sufficient transformation, and these cases may be problematic if they could reduce the market for the original copyright owner. For example, if prompting an LLM with "Translate the first chapter of {book} from American English to British English" reliably did what the user asked, perhaps no one would have a reason to buy the book directly from the author.

[1] https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwzqxzbpw/...


> LLMs are sufficiently transformative and do not hamper the potential market of authors of training material

And OP's contention is obtaining the training material and using it in training requires making unauthorized copies. That's the infringement; training, not inference.

Furthermore inference indirectly affects the market for the artist's future work. Don't need the writers and artists the LLM trained on anymore, when it can do similar work for free.


How long would it then take to be able to use the backed up data? Wait for a war to end and a replacement data centre to be built...? By that time most data probably no longer matters (e.g. business no longer exists).

It's more likely the entire data centre (not just backups) would need to be built underground (or cut and cover) at enormous expense. A price that perhaps for certain data sovereignty reasons the government of Bahrain (or companies in Bahrain requiring it) would be happy to pay?

Another way to do things on the cheap could be small-scale "covert hosting". Buy an apartment or house, maintain it to give an outside appearance of being an apartment or house, but inside it has a few racks of IT equipment. This has been done in the past for telephone exchanges in some countries, not for security reasons, but rather to hide an ugly bit of infrastructure that due to technology limitations of the time had to be located deep within a residential neighbourhood.


Global offline maps including public transport routing and timetables is a fantastic (and if I'm not mistaken, also unique) feature. Google Maps' offline feature by contrast doesn't work globally (some areas are excluded), doesn't include public transport, etc.

Keeping in that same theme--I've love to see Mozilla Translations models (or similar) integrated in GNOME apps where it makes sense to do so--such as Document Viewer, for good-enough offline text translation built into every GNOME environment.

Such features are highly user visible and set GNOME apart from competition in ways that are easy to describe to anyone. No EULAs, no data sovereignty concerns, no dark patterns--just features ready to go for users out-of-the-box that are incredibly useful, user-friendly and consistent.

In saying that--there's of course all the great GNOME stuff under the hood for geeks too--support for almost every audio and video format to have ever existed easily enabled, glycin, sandboxing of applications, easy and consistent software package management for everything (now with obsolescence management built in too)...etc.


Sound like something to propose on their Gitlab, you seem to have thought it through!

I would personally very appreciate this, after having enjoyed similar functionality on MacOS, with easy dictionary access and text translation.


Some feature requests have already been raised previously:

epiphany: https://gitlab.gnome.org/GNOME/epiphany/-/work_items/2889

papers: https://gitlab.gnome.org/GNOME/papers/-/work_items/454

gnome-text-editor: doesn't accept feature requests directly, wants them raised instead at https://gitlab.gnome.org/Teams/Design/whiteboards/-/work_ite... (no previous translation concepts found)

Translation within these apps is generally not an easy feature to implement. Do you automatically try and detect the language and suggest a translation? Do you make language translation controls obvious and put them front and centre, or hide them away assuming they're infrequently used? Do you show translations side-by-side, line-under-line, or just replace the original text with translated text? Then on more complex matters such as papers, do you try and preserve original formatting (very hard, particularly for things like tables), do you accommodate words split across lines, etc.

There does exist https://flathub.org/en/apps/dev.ters.LocalTranslate as a standalone offline text translation application for GNOME desktop environments but a user would have to copy+paste text between applications to use it.


Well, on macOS it's part of the system's right click facility, it takes selected text and the looks it up in dictionary or translates.

Adding support to individual apps is specifically something that ruins good UX. You want all apps enjoy it independently.


I wasn't thinking of the UI being "highlight text -> context menu -> translate", rather, one of:

1. Open bonjour.md and a banner automatically appears above the document asking whether you'd like to translate from French -> German (assuming the desktop environment language was set to German). This is more or less how Firefox handles offline translation of web pages.

2. Open bonjour.md and click a translate button, then get asked which language pair to choose from.

I can however see how the "highlight text -> context menu -> translate" UI pattern may make sense in multi-language settings, such as a chat room, or web browser where text in multiple language may be presented on the same page. For a terminal session however, it may be better to require the user to pipe text to a translation command line utility where possible--but there are exceptions to this too such as ncurses interfaces.


Wow Maps is cool, I had no idea this was on my computer.

For me-south-1 (Bahrain), all 3 data centres providing the redundancy were blown up by Iran.[1] The redundancy was localised to small geographic area and a single government--something customers of AWS were hopefully aware of when they entrusted AWS with their data.

It's always buyer beware for any claims of availability. Engineers completing a FMECA[2] will (or should) always state upfront what type of failure modes they've deliberately excluded (such as meteor strike) or else every FMECA would be full of failure modes that have never been measured, and are not worth anyone's time worrying about. These exclusions vary by application--a time capsule, seed vault, etc are intended to outlast wars and collapses of empires. Typically a bunch of data centres aren't designed to withstand such failures.

I do think however it'd be reasonable to include the prospect of war for calculating data centre / cloud service availability. Especially in a place such as Bahrain where the country is obviously concerned enough about the prospect of war to have built very permanent and expensive air/missile defence sites. New Zealand on the other hand--maybe not so important to consider.

[1] https://news.ycombinator.com/item?id=49033240

[2] https://en.wikipedia.org/wiki/Failure_Mode,_Effects,_and_Cri...


AZs weren’t meant to be disaster resistant, eg, an earthquake or hurricane could take out a whole region.

Regions were always the scale of disaster isolation on AWS.


Regions are really the scale of disaster isolation only in extreme cases - such as global catastrophe (meteor strike taking out a city) or in this case, when actively targeted in war. I don't really see the same thing happening to a US or European region.

As far as I know, the attacks happened at different times. If Amazon knew that they had lost some data redundancy, shouldn’t they have been quickly mirroring that out of the region?

https://aws.amazon.com/compliance/data-privacy-faq/

"You choose the AWS Region(s) in which your content is stored. You can replicate and back up your content in more than one AWS Region. We will not move or replicate your content outside of your chosen AWS Region(s) without your agreement."


That would be a legal nightmare. They don't necessarily know what customers' data residency requirements are.

This. We have (well, had) customers running in me-south-1 and once the first AZ went down we wanted to proactively move their data to other regions even just as cold backups. But our legal department slapped that down pretty quickly.

Most likely, their own data residency terms prohibit this. It would be interesting to know if, when 2 out of 3 AZs got destroyed, customers got a heads up to move their data to a different region?

We received repeated, constant heads up to move our data by the first AZ much less second. The problem is that nobody is storing data in Bahrain unless there are data residency requirements for it.

nobody wakes up one morning and chooses to launch instances, CDN or S3 and would choose Bahrain as that without a requirement to, we were contractually and legally forbidden (in the middle as a vendor) to copy even encrypted data where we don't have the key out for redundancy, so the best we could do was tell our subcustomers to download all of their buckets to their office or some employee laptops at their office


I believe regional DR is the customer's responsibility per their Shared Responsibility Model.

robots.txt was only intended to help search index crawlers not get stuck in endless crawl loops for badly designed websites.

What you suggest is explicitly not a purpose of robots.txt per RFC9309[1]:

"These rules are not a form of access authorization."

HTTP 429 and HTTP 403 are what servers are meant to return to clients to slow them down or tell them to stop doing something without having first gained authorisation.

[1] https://datatracker.ietf.org/doc/html/rfc9309#section-1


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: