my biggest code smell is debugging. Trying to step through the code is slow because of how much the debugger has to keep track off and it's a lot of step into / jumps until you get to the piece you actually care about. Then you also see stuff like "if langchain_v2: <branch> elif langchain_v3: <branch> else ...." ... breaking changes are ok! I feel like at this point they're hopelessly spaghettied and need a more opinionated approach to clean up a lot of the mess.
Since others are asking: My current goto is pydantic ai and it's been doing pretty well for my use cases.
Cool to see, but wondering what the upside of this is vs what notion is building with their own agents. I can imagine notion agents can do something very similar to this (and with the same amount of control / security) and at what point would it make sense for a company to build their own vs buying into one..
so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)
Not that anthropic models are very good at this, but due to the changes in tokenizers and thinking tokens: cost per token is not as helpful anymore as cost / task.
Did wonder about the review, it probably sold well in the news papers and part of me thinks that the Odyssey team might even agree with her, but IMHO it's mostly overblown?
The movie is good an wildly successful, and maybe I've gotten less jaded, but felt like her criticism did not land ( saying that as someone who loves to poke holes at things and reading a good polemic).
She's entitled to her views, and others are entitled to think, like you, that her review is overblown.
My view is she has some good points about script and plot, given her expertise is in writing, but she's under-emphasising the importance of the visual arts.
Remember that Slashdot's review of the iPod was "No wireless. Less space than a nomad. Lame.". Not everybody has to like the popular thing.
I quite like watching Yahtzee's Fully Ramblomatic reviews. He's known for giving most games a kicking, but he does highlight what's good in a game as well as its bad points. I'd rather hear what he has to say than get a adverising puff-piece review from IGN. Even for games that I strongly enjoyed, he usually raises good points about them that I'd agree with.
Shower thought: But how often per week do you run the pelican these days?
And do you have it automated at this point or would the automation take out the meaning of the benchmark?
Since others are asking: My current goto is pydantic ai and it's been doing pretty well for my use cases.
reply