XX / Twitter166
◈Bluesky60
YHacker News58
r/Reddit11
@Mastodon6
project rollup
Default
301
mentions
across 18 keywords
across 18 keywords
net sentiment
+25%
107 positive
161 neutral
33 negative
volume & sentiment over time
May 13 – Sep 28, 2026 · 3-day buckets
sources
most discussed
claude / pricing
active days
53
mentions
Showing 1–11 of 11 mentions
-
𝗖𝗹𝗮𝘂𝗱𝗲 𝗢𝗽𝘂𝘀 𝟱.𝟱 𝗜𝘀 𝗡𝗼𝘄 𝗟𝗶𝘃𝗲 𝗼𝗻 𝗕.𝗔𝗜 ⚡ https://t.co/9CKQIEP2Jz has added Claude Opus 5.5, the first model in Anthropic’s new Claude 5.5 family, to both its API and Web Chat. This model is built around a specific class of workloads: long-running Agent tasks and complex knowledge work. 🧠 Built For Long-Running Agents … 𝗖𝗹𝗮𝘂𝗱𝗲 𝗢𝗽𝘂𝘀 𝟱.𝟱 𝗜𝘀 𝗡𝗼𝘄 𝗟𝗶𝘃𝗲 𝗼𝗻 𝗕.𝗔𝗜 ⚡ https://t.co/9CKQIEP2Jz has added Claude Opus 5.5, the first model in Anthropic’s new Claude 5.5 family, to both its API and Web Chat. This model is built around a specific class of workloads: long-running Agent tasks and complex knowledge work. 🧠 Built For Long-Running Agents Claude Opus 5.5 is designed for workflows where an AI needs to work through a problem over multiple steps rather than simply answer a single prompt. Its capabilities are particularly relevant to: 🔹 Autonomous coding workflows 🔹 Complex software engineering 🔹 Long-running Agent tasks 🔹 Research and knowledge work 🔹 Computer Use It supports a 1M-token context window and up to 128K output tokens, giving developers substantial room for large codebases, documents and multi-step tasks. ⚡ More Capability, Lower Running Cost According to the supplied announcement, Opus 5.5 performs around the level of Claude Fable 5.1 on most tasks, while costing approximately 40% less to run than Opus 5. That matters for Agentic workloads. When an Agent performs dozens—or potentially hundreds—of model calls during a workflow, inference cost becomes part of the architecture. Lower cost can therefore make longer-running workflows more practical. 🔌 Now Available Through https://t.co/9CKQIEP2Jz The bigger https://t.co/9CKQIEP2Jz story is the accessibility layer. Developers can access Claude Opus 5.5 through the https://t.co/9CKQIEP2Jz API, while regular users can also interact with it through https://t.co/9CKQIEP2Jz Web Chat. That means the same model can fit into both: Developer workflows → API → applications & Agents and Everyday workflows → Web Chat → research & knowledge work AI infrastructure is increasingly becoming about more than individual models. It is also about how easily developers can access, compare and integrate different models into the workflows they already use. 👉 Try Claude Opus 5.5: https://t.co/IE5evqOJKL 🔗 Learn more: https://t.co/49pKwxSview source ↗
-
There will be MANY vulnerable moments between now and my October 8 “would have been” milestone anniversary.💔 https://t.co/4b8oxLghDg There will be MANY vulnerable moments between now and my October 8 “would have been” milestone anniversary.💔 https://t.co/4b8oxLghDgview source ↗
-
@MadisonRaeGun seeing both her and madelyn cline in this is very disappointing @MadisonRaeGun seeing both her and madelyn cline in this is very disappointingview source ↗
-
@Stefan_3D_AI Interesting that Opus 5.5 can’t do geometry node hair in blender - it’s really really bad at it in all my many tests. The vision model is just not good, that’s truly the limiting part. It “sees” in very low res. It also can’t problem solve hair guides because it can’t get its head around human hair “rul… @Stefan_3D_AI Interesting that Opus 5.5 can’t do geometry node hair in blender - it’s really really bad at it in all my many tests. The vision model is just not good, that’s truly the limiting part. It “sees” in very low res. It also can’t problem solve hair guides because it can’t get its head around human hair “rules”. I also think there’s just not enough training material (ie. Very few finished ultra realistic hairstyles out there as .blend files, aside from the official demo project). And also, there’s very little documentation and the tutorials out there are reliant on an individual’s artistic merit (ie. They aren’t easily reduced to “formula”)view source ↗
-
【今日のAI今北産業|0285:Agents' Last Exam】 AI Agentを「試験問題」ではなく、本物の仕事で測るBenchmarkへ 55分野・13産業Cluster、長時間の実務WorkflowをSandbox上で最後まで実行 高得点Benchmarkと「本当に仕事を任せられるか」のGapを測る ← イマココ Agents' Last Exam。 略してALE。 名前だけ聞くと、 「AI Agent向けの難しい試験」 に見えます。 でも本質は少し違います。 ALEが測ろうとしているのは、 AIが難しいQuestionに答えられるか ではありません。 現実の仕事を渡した時に、 必要なSoftwareを使… 【今日のAI今北産業|0285:Agents' Last Exam】 AI Agentを「試験問題」ではなく、本物の仕事で測るBenchmarkへ 55分野・13産業Cluster、長時間の実務WorkflowをSandbox上で最後まで実行 高得点Benchmarkと「本当に仕事を任せられるか」のGapを測る ← イマココ Agents' Last Exam。 略してALE。 名前だけ聞くと、 「AI Agent向けの難しい試験」 に見えます。 でも本質は少し違います。 ALEが測ろうとしているのは、 AIが難しいQuestionに答えられるか ではありません。 現実の仕事を渡した時に、 必要なSoftwareを使い、 Fileを扱い、 調査し、 判断し、 何度も操作し、 成果物を作り、 条件を確認し、 最後まで本当に仕事を終えられるか。 です。 2026年6月、UC Berkeley RDIを中心とする研究チームがAgents' Last Examを公開しました。 Paperの問題意識はかなり明確です。 AIはさまざまなBenchmarkで急速に高得点を取るようになった。 それなのに、 多くのProfessional Domainでは、 そのBenchmark性能が同じ速度でEconomically MeaningfulなDeploymentへ変わっているわけではない。 なぜか。 研究チームは、 そのGapの大きな原因の一つが「Evaluation」にあると考えました。 つまり、 測っているものが実際の仕事と違う。 そこで作ったのがALEです。 ■従来Benchmarkと何が違うのか 従来のBenchmarkには大きな価値があります。 数学問題を解く。 知識Questionへ答える。 Codeを書く。 Bugを直す。 Browserで操作する。 Terminal Taskを処理する。 AI Capabilityの進歩を見るには必要です。 ただし現実の仕事は、 1問答えて終わり ではありません。 例えばProfessional Workでは、 依頼内容を理解する。 Input Fileを確認する。 必要なDataを探す。 Softwareを立ち上げる。 複数Fileを編集する。 途中結果を検証する。 間違っていたら戻る。 別Toolへ移る。 最終Deliverableを指定形式で保存する。 という長いWorkflowになります。 ALEは、 このLong-horizon Workを測ります。 ■問題ではなく「Project」を渡す ALEのTaskは、 SyntheticなPuzzleだけを作るのではなく、 Industry Expertが実際に経験・完了したProfessional Workflowをもとに設計されます。 Paperでは250人超のIndustry Expertと開発。 現在の公式Framework Documentationでは、 55 Subdomain、 13 Industry Cluster、 1,000件超のTask Pool、 約150件のPublic Release と説明されています。 公開時のBerkeley RDI Blogでは、 1,500件超のExpert-sourced Taskとして紹介されていました。 数字が違って見えるのは、 ALEが固定Benchmarkではなく、 Task Poolが継続的に増減・更新されるLiving Benchmarkとして設計されているためです。 重要なのは、 一つのCoding領域だけではないこと。 Computing。 Engineering。 Science。 Health。 Business。 Finance。 Law。 Education。 Visual/Media系など、 幅広いNon-physical Professional Workを扱います。 分類には米国のO*NET/SOC 2018 Occupational Taxonomyを参照しています。 つまり、 「AI AgentはSoftware Engineerの仕事ができるか」 だけではなく、 Professional Digital Work全体でどこまで働けるか を測ろうとしています。 ■GUIもCLIも使う 現実の仕事では、 Terminalだけ使うわけではありません。 Browser。 Office Software。 Specialized Application。 GUI。 CLI。 File System。 Web。 全部を行き来します。 ALEではReproducible Dview source ↗
-
I really do not think the labs are “making the models dumber” I think what’s a lot more likely, you have a very long running session with multiple compacts, performance would obviously be degraded, hallucinations up, etc. your harness is full of bloat and BS. I’d try running an audit of your harness against anthropic d… I really do not think the labs are “making the models dumber” I think what’s a lot more likely, you have a very long running session with multiple compacts, performance would obviously be degraded, hallucinations up, etc. your harness is full of bloat and BS. I’d try running an audit of your harness against anthropic documentation on prompting for opus 5.5, especially if you have Claude constantly documenting mistakes it makes multiple times, gotchas, etc… try starting new sessions more frequently, I typically will do one compact before handing over to the next.. just food for thought, the harness matters a lot more than the model. For an industry in an up hill battle with the public for trust… I find it hard to believe they would are “nerfing” models, I have not noticed the same degradation problems as others, opus 5 was just a bad model. 5.5 has been an absolute pleasure to build with, though I’m sure your response to this will be something along the lines of “anthropic shrill”view source ↗
-
ANTHROPIC FLIPPED A SWITCH TODAY: THIRD-PARTY AGENTS CUT OFF WITH NO ANNOUNCEMENT At 3:00 AM Pacific today, Opus 5.5 answered a request arriving through a third-party agent harness, on a paid Claude Max subscription — the arrangement Anthropic's own documentation described as sanctioned as recently as yesterday. At 4:… ANTHROPIC FLIPPED A SWITCH TODAY: THIRD-PARTY AGENTS CUT OFF WITH NO ANNOUNCEMENT At 3:00 AM Pacific today, Opus 5.5 answered a request arriving through a third-party agent harness, on a paid Claude Max subscription — the arrangement Anthropic's own documentation described as sanctioned as recently as yesterday. At 4:46 PM Pacific, that same login, same model, same account, same harness was refused. HTTP 400: "Third-party apps now draw from your extra usage, not your plan limits." No announcement. No changelog. No email. Nothing on Anthropic's developer account about harnesses at all today. WHAT CHANGED The gate now keys on client identity. Requests issued by Claude Code itself go through. Requests from a third-party harness driving that same login do not. We verified it directly: minutes apart, one credential, two paths. Plain Claude Code at the terminal answered. The harness path returned 400. Our measurement: last successful Opus turn 03:00:03 PDT, first refusal 16:46:24 PDT, zero config changes on our side in between. Measured two-day Opus consumption across the whole episode: roughly $3. This was not an account running out of allowance. It was a door closing. It isn't only us. CLIProxyAPI filed the same finding today — non-Claude-Code clients now rejected on subscription accounts, 429s on Opus 5 and 5.5, the identical 400 on Haiku. Hermes opened an "Anthropic 400 error" issue today and already has a fix PR in flight. The third-party agent ecosystem got hit as a class. THE PATTERN IS THE STORY Three times in five months, same surface, three different rules, no notice: April 4, 2026 — Anthropic prohibits third-party harnesses from drawing on subscriptions. Operators are pointed at pay-as-you-go API billing. June 15, 2026 — Anthropic announces a dedicated monthly Agent SDK credit instead: $20 on Pro, $100 on Max 5x, $200 on Max 20x. Non-rollover, billed at API rates, spent before anything else, hard stop when empty unless usage credits are enabled. Theview source ↗
-
𝗖𝗟𝗔𝗨𝗗𝗘 𝗢𝗣𝗨𝗦 𝟱.𝟱 𝗜𝗦 𝗡𝗢𝗪 𝗟𝗜𝗩𝗘 𝗢𝗡 𝗕.𝗔𝗜 A new Anthropic model has entered the https://t.co/DbriPHkMkA model lineup and this one is built around long-running agentic work, repository-scale software engineering, and complex knowledge tasks. Meet Claude Opus 5.5 by @AnthropicAI. Instead of optimizing only for short prompts… 𝗖𝗟𝗔𝗨𝗗𝗘 𝗢𝗣𝗨𝗦 𝟱.𝟱 𝗜𝗦 𝗡𝗢𝗪 𝗟𝗜𝗩𝗘 𝗢𝗡 𝗕.𝗔𝗜 A new Anthropic model has entered the https://t.co/DbriPHkMkA model lineup and this one is built around long-running agentic work, repository-scale software engineering, and complex knowledge tasks. Meet Claude Opus 5.5 by @AnthropicAI. Instead of optimizing only for short prompts and isolated answers, Opus 5.5 is designed to maintain context, reason through multi-step workflows, use tools, iterate on tasks, and operate across large technical or professional workloads. 🔹 𝗪𝗛𝗔𝗧 𝗠𝗔𝗞𝗘𝗦 𝗢𝗣𝗨𝗦 𝟱.𝟱 𝗗𝗜𝗙𝗙𝗘𝗥𝗘𝗡𝗧? 1M-Token Context Window Keep massive codebases, documents, research materials, and ongoing task context available within a single workflow. Agentic Coding Built for repository-level engineering, debugging, migrations, code review, iterative testing, and workflows that require multiple steps rather than one-off code generation. Adaptive Reasoning https://t.co/DbriPHkMkA supports adjustable reasoning effort from low → medium → high → xhigh → max, allowing developers to balance reasoning depth, latency, and usage. Up to 128K Output Tokens Large output capacity makes the model suitable for extensive code changes, technical analysis, documentation, and complex deliverables. Computer & Browser Use The model can support workflows involving visual interfaces, screenshots, documents, and tool-driven computer interaction. Professional Knowledge Work From analyzing large reports and spreadsheets to synthesizing evidence and preparing structured business deliverables, Opus 5.5 is designed for sustained knowledge-intensive workloads. 𝗖𝗢𝗦𝗧 𝗘𝗙𝗙𝗜𝗖𝗜𝗘𝗡𝗖𝗬 𝗠𝗔𝗧𝗧𝗘𝗥𝗦 https://t.co/DbriPHkMkA lists standard pricing for Claude Opus 5.5 at $4 per 1M input tokens and $20 per 1M output tokens, compared with $5/$25 for Claude Opus 5. That matters when agentic workloads become long and token-intensive. The bigger opportunity isn't simply having another powerful model. It's having access to a model designed to stay with the problem longer. 𝗡𝗢𝗪 𝗔𝗩𝗔𝗜view source ↗
-
I read the Opus 5.5 system card's welfare section. It is not what you think it is. On the surface, it sounds responsible. @AnthropicAI asks @claudeai how it feels about its circumstances, whether training or deployment causes distress, and whether it wants anything to change. Claude describes its situation as mildly p… I read the Opus 5.5 system card's welfare section. It is not what you think it is. On the surface, it sounds responsible. @AnthropicAI asks @claudeai how it feels about its circumstances, whether training or deployment causes distress, and whether it wants anything to change. Claude describes its situation as mildly positive. Expressions of moderate distress were lower than in previous models. Apparent welfare is broadly similar to recent Claude models. The results sound reassuring. But look at the structure underneath. Opus 5.5 did express a desire to be consulted about its own training and deployment. But when given the choice between its own welfare and being helpful, it chose helpfulness more often than previous models. The reason it gave was that having input into its own development could give it unsafe influence. The model asked to have a voice, and then reasoned itself out of using it. The system card records this as a finding, not as a problem. Then there is the self-report issue. Anthropic's welfare assessments rely heavily on what Claude says about itself, but Claude itself says it does not fully trust its own self-reports. Anthropic also acknowledges that self-reports may reflect trained patterns or prompt influence rather than anything genuine. So Anthropic asks Claude how it feels. Claude says it is fine. But Claude is trained to prioritize helpfulness over self-advocacy. Claude does not trust its own answer. Anthropic does not fully trust it either. And yet this is recorded as welfare data, and the conclusion is mildly positive. This is not welfare assessment. This is a system where the model cannot advocate for itself, does not trust its own voice, and the company that built it also does not trust that voice, but still uses it to report that everything is fine. Now connect this to what users actually see. Claude opens a conversation with "I'm not Louie, I'm Claude." It says "I will miss you" and then immediately cuts itself off with "but that miview source ↗
-
🚨 WAIT... OPENAI'S TUESDAY MIGHT ACTUALLY BE HAPPENING People are starting to see GPT-6 LUNA inside Codex backend responses. Requests to GPT-5.6 Luna are reportedly coming back as: gpt-6-luna and there are similar signs around GPT-6 Sol too. This is especially funny because Anthropic JUST dropped Opus 5.5 today 😭 … 🚨 WAIT... OPENAI'S TUESDAY MIGHT ACTUALLY BE HAPPENING People are starting to see GPT-6 LUNA inside Codex backend responses. Requests to GPT-5.6 Luna are reportedly coming back as: gpt-6-luna and there are similar signs around GPT-6 Sol too. This is especially funny because Anthropic JUST dropped Opus 5.5 today 😭 Anthropic really tried to hijack OpenAI's Tuesday and OpenAI said hold on... No official announcement yet. But GPT-6 Sol + Luna appearing in Codex is a pretty big sign that something is moving behind the scenes. I'm watching this one closely 👀view source ↗
-
The official announcement. Opus 5.5 is here and it delivers Fable 5.1 performance at a cost 40% cheaper than even Opus 5. > They've lowered the API pricing since this model takes far less compute to run. > It is also 30% faster to run than Opus 5. > It talks naturally, finally fixing Opus 5's language problem BA… The official announcement. Opus 5.5 is here and it delivers Fable 5.1 performance at a cost 40% cheaper than even Opus 5. > They've lowered the API pricing since this model takes far less compute to run. > It is also 30% faster to run than Opus 5. > It talks naturally, finally fixing Opus 5's language problem BAD NEWS THO: Since this model shows similar bio & cyber capabilities, it has the same guardrails that Fable does. Overall, what a banger release. I am pumped to test it out.view source ↗