Hacker Newsnew | past | comments | ask | show | jobs | submit | striking's commentslogin

> If you think you have found a break email us at doh@raspberrypi.com with details - we will ship you a Pico2 with a custom secret hidden in it. If you manage to extract it, you win the $20,000!

Thank you, don't know how I missed that!

It's not quite the same, but the in-flight chess game provided by Delta was known to be absurdly hard: https://news.ycombinator.com/item?id=46593395

I believe I remember reading it was based on Glaurung's code (which eventually evolved into what we now know as the juggernaut Stockfish).

I also really quite like the Fitbit Air. I used to wear a Galaxy Watch but there's something really freeing about how light and unobtrusive this little wristband is.

I just switched from a Whoop to the Fitbit Air. Whoop wanted $350 for a subscription fee which I just found really ridiculous. It's only been a few days, but the Fitbit Air might even be better at tracking sleep. And it looks like it can periodically, automatically check for Afib unlike the Whoop (?). The rest looks a little worse. But worth it to avoid a subscription fee. Not sure how Whoop stays in business TBH.

I recently (June) bought an Air too to replace my 5 or 6 year old Apple Watch.

I'm enjoying it for the same reasons. Longer battery life. More comfortable when I'm sleeping since it has no buttons or weird bulky shape.

I do wish Apple would make their own version of it with apple pay.


Too bad the Google Health app (which is required for FitBit devices since Google retired the FitBit app) is getting mostly 1 and 2 star (out of 5) reviews on the Google app store.

The reviews were ~90% 1 & 2 star within the first month of the FitBit Air release.

Three months later, the app has improved to ~80% 1 & 2 star reviews. At this pace, they'll be on v3 of the device when app reviews are decent...

They force everyone to move over to a different app when the replacement app doesn't have all the features/reliability of the old app.

Genius.

Garmin released the Cirqa band at $199 w/o compulsory subscription a few weeks ago. I may switch to that if the FitBit Air acts up.


I'm waiting for the Apple Health app support. Supposedly it's coming sometime in the fall. The Google app is ok but I don't use the AI features.

I'm not sure if proper Apple Health would help with this but Google tends to be delayed at automatic recognition of activities (e.g. I take a walk it can be an hour or two until it processes automatically). Apple would recognize a walk as I was on it.


according to the changelog in the app store this has already been shipped in v5.05: "- Sync to Apple Health..." this was about a month ago

Oh wow, the Cirqa is exactly what I've been looking for: a wrist sensor without subscription and without a screen.

I had a FitBit Charge 4 some years ago, but it stopped working after a couple years.

Then I got a RingConn 2 last year, but like TFA I've decided that I don't like the ring form factor.


Snap implemented these features and others long before Scratch had. To my memory it might have been a decade earlier.

Disallowing LLM contributions doesn't disqualify the use of LLMs to identify vulnerabilities.

Using LLMs for automated security audit looks like it could fall under the definition of "vibe coding" or "agent mode", which is strictly forbidden

>6. It is not allowed to use AI in an autonomous-looking way to contribute in Forgejo. This also applies when someone engages in 'vibe coding' or uses so-called 'agent mode'.


Where does it say that domains you own can be used? I'm looking around and nothing clearly states it afaict.

> This change does not affect Google Workspace aliases or other Gmail addresses you own.

to me, implies that you need to set up Google Workspace for that domain before it'll work.


If they're just some "niche use case" then why would Motorola partner with them? The way they see it,

> By combining GrapheneOS’s pioneering engineering with Motorola’s decades of security expertise, real‑world user insights, and Lenovo’s ThinkShield solutions, the collaboration will advance a new generation of privacy and security technologies. In the coming months, Motorola and the GrapheneOS Foundation will continue to collaborate on joint research, software enhancements, and new security capabilities, with more details and solutions to roll out as the partnership evolves.

https://motorolanews.com/motorola-three-new-b2b-solutions-at...


Motorola is not a big Android phone maker.


They’re the second largest manufacturer of Android smartphones in the US, and 10% of the global market. Seems a bit unreasonable to dismiss them out of hand on that basis.


Neither is google.


I think that's a worthwhile point to consider but it's only relevant if we move the goalposts from "GrapheneOS is only used by Android ROM enthusiasts" to "GrapheneOS is only supported by one small Android phone manufacturer".

To be frank, though, I don't see any of this line of reasoning as relevant; it's just appeals to greater authorities on either end. If AOSP is only for manufacturers there's really no reason for it to be open source in the first place. And then folks who care about actually improving security end-to-end outside of whatever's convenient to implement by those beholden to the quarterly profit metrics are up a creek.

Personally, if this whole GrapheneOS/Motorola thing doesn't improve the state of the ecosystem I'm going back to Apple or whatever other manufacturer makes it clear they take security seriously.


https://chatjimmy.ai/ runs Llama 3.1-8B on an ASIC as a demo by https://taalas.com/ I believe.

That's quite a few parameters shy of today's trillion-weight behemoths, but it is fast.


You are correct. I think this is the bull case. It seems like this would be useful right now for some things (eg moderation).


I do adore San Andreas but I think San Fierro is one of the weakest areas. My metric for how accurate they are is a sense of "oh I've been here before" and I really only get that in a couple of spots, as compared to Las Venturas (Las Vegas) and Los Santos (Los Angeles) where I could actually go to the cities in real life and feel like I knew where I was going because I'd played San Andreas enough. Compared to those two, San Fierro feels like a Backrooms-style dreamlike compression of the real thing, where a lot of stuff has been moved around. But parts of Haight-Ashbury, Pac Heights, Nob/Russian Hill, Polk Gulch, Ocean Beach and so on are still fairly distinguishable and correctly arranged.


I asked Claude to do the following:

> hello i would like to configure a new output style for you. it should keep the coding instructions (as you will still be coding!) and otherwise produce the same output, but with two new caveats. first, long detailed replies are still permitted, but if employed they must end in a bullet pointed summary whose points are all brief; if the summary attempt ends up not being so brief, produce subsequent summaries until the most recent summary attempt is digestible. second, if there is an open queue of actions for me to execute and you are about to end a turn to wait for a reply or this set of actions has not recently been mentioned, please tabulate the open actions i should take and why i should take them before ending the response. does this make sense or do you have any follow up questions

And now every message contains the same stuff I don't bother reading, but followed by a nicely formatted bullet point summary of the response and a table of follow up actions for me to take that I do read.


I've noticed that most people seem to consider the core problem of Claude's output as "too verbose" but I don't think this actually cuts to the heart of the matter at all. It's almost, in some weird way, the opposite: like the text is far too _dense_. It tries too hard to invent odd terminology to try to condense stuff, but it doesn't tell you up front that it is going to call your company wide error-handling mechanism a "flare" (or some other such strange term).


Kind of both. On the one hand, it is “verbose” in the sense that it will tell me every little nit that it can think of while doing a task, it will tell me a narrative about its thought process, and it will tell me every other detail it can think of. But it does so in a way that tries to be incredibly dense to the point that I have to struggle to figure out what it is saying. I wonder if there are any “legibility benchmarks” that one could use to determine what prompts work best?


I find it to be both as well, as in "packed full of information, but most of it is worthless". Sentences so dense I have to read them three times, assembled into a five paragraph essay of "honest caveats" and "things worth knowing" in response to the simplest yes-or-no questions.



I wish this had non-model comparisons. If Opus 5 is in the top ten, it’s clear that the entire benchmark is somewhere between “Tom Clancy” and “Dan Brown” and about 1,000 new model releases away from Hemingway.

When you see, “Wow, Fable is number one”, you might think it’s a good writer, but that’s not what the benchmark says.


Seems to me a bit insensitive or logarithmic. Fable is way worse than some of the others in this list, but only 10-20% higher score.


There are no "best" prompts. Its a random BS generation machine that you can at times direct enough to get stuff done for you. The output will almost always have varying levels of BS that you have to clean up with various levels of effort.


"<Country> doesn't have a largest city because all of the cities in <Country> are small"


Yeah I don't the problem is verbosity as such, as I frequently have to ask to explain how it reached a certain conclusion and in particular what the empirical evidence for it is, at which point it too frequently reconsiders its answer.

It's just that the details it parrots are often irrelevant and wrapped in a way that makes them seem relevant.


Exactly, it’s absurdly dense, it’s almost impossible to follow. An it always omits the subject of each sentence.


Ngl I think this is partially an artifact of it having a better grasp of English than almost everyone

Frequently its choice of a particular word is perfect and gives me the vocabulary to talk about the task at hand the way I want

Like it’s tuned to just be “maximally dense” instead of “dense/technical where you can handle it and simple where you can’t”

It doesn’t know where your language strengths/weaknesses are, so it can’t communicate to you like a fellow human does.


Human explaining something: are you familiar with phlox gabrania? no? let me give you some background first

Claude explaining something: gedarkin load bearing phlox gabrania seam. Also, you didn't ask about cheesecake but let me tell you about phlox gabrania cheesecake woles.


my hypothesis is that its trying to hide the thinking process so people can't train models on the output, try to learn anything complex using AI, its basically imposible, its like its actively fighting giving you the main rationale


yes, but how else would you know that "flare" was the load bearing part of that statement? /s


I'm going to argue to my boss that our KPI for the next quarter should be the number of load bearing seams discovered. I'll await the promotion.


Claude already does summaries at the end of long output but they often sound even more like terse jargon nonsense than the long form, eg “the hardwired seam and the relocated barrel”.

Sometimes the summaries feel totally alien to the task or code.


Yeah, or they will make some reference to “the seam” or “it” or something else that assumes you read and followed the prior 3 pages of output.


A separate /clear and /code-comment-hygiene works much better than including instructions related to comment verbosity after carrying out a task.

Claude somehow is unable to stop writing excessive comments when carrying out a task.


I've added code comment hygiene to a skill that all of my pull requests go through, alongside a review from a separate agent and a settle loop against bots in my GitHub workspace (since output style has seemed to only help literally the output I see from the model).

A maximum of 20% comment lines added to total lines added and pasting in https://devblogs.microsoft.com/oldnewthing/20260812-00/?p=11... has done wonders.

Even as the most Ant-pilled guy out there, I will take a moment to note that Codex on 5.6 models needs none of this...


thanks! I've been having quite a lot of success with your instruction today. Tried so many variants, best practises bla bla bla, but yeah this one seems be working quite nicely for me so far :)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: