XX / Twitter166
◈Bluesky60
YHacker News58
r/Reddit11
@Mastodon6
project rollup
Default
301
mentions
across 18 keywords
across 18 keywords
net sentiment
+25%
107 positive
161 neutral
33 negative
volume & sentiment over time
May 13 – Sep 28, 2026 · 3-day buckets
sources
most discussed
claude / pricing
active days
53
mentions
Showing 76–100 of 166 mentions
-
6M views on the Day Drinker official YouTube teaser! Keep it going! https://t.co/vpJckoV3bU #DayDrinker starring Johnny Depp, Madelyn Cline, and Penélope Cruz – in theaters March 26, 2027. https://t.co/2Ld2Z0kmWm 6M views on the Day Drinker official YouTube teaser! Keep it going! https://t.co/vpJckoV3bU #DayDrinker starring Johnny Depp, Madelyn Cline, and Penélope Cruz – in theaters March 26, 2027. https://t.co/2Ld2Z0kmWmview source ↗
-
@jeansdildo If I was rich a and handsome I’d be with Maddie Cline and we’d have beautiful children We can both imagine @jeansdildo If I was rich a and handsome I’d be with Maddie Cline and we’d have beautiful children We can both imagineview source ↗
-
Setup takes under a minute in Cline: 1. Subscribe to the coding plan on https://t.co/4YapRnBJPW 2. Grab your coding API key from the developer portal 3. In Cline settings, pick https://t.co/4YapRnBJPW as provider, paste key, select GLM-5.3 Clean terminal, predictable billing, zero surprise charges. Setup takes under a minute in Cline: 1. Subscribe to the coding plan on https://t.co/4YapRnBJPW 2. Grab your coding API key from the developer portal 3. In Cline settings, pick https://t.co/4YapRnBJPW as provider, paste key, select GLM-5.3 Clean terminal, predictable billing, zero surprise charges.view source ↗
-
Sam asks governments to help control AI as an OpenAI agent crosses an Australian boundary. Dario ships Opus 5.5, Zuckerberg challenges the phone and Son borrows billions. Episode 22: Sam Calls for Control. Dario Scales Anthropic. #SiliconDrama #eTatos https://t.co/3RFpil9hQ2 Sam asks governments to help control AI as an OpenAI agent crosses an Australian boundary. Dario ships Opus 5.5, Zuckerberg challenges the phone and Son borrows billions. Episode 22: Sam Calls for Control. Dario Scales Anthropic. #SiliconDrama #eTatos https://t.co/3RFpil9hQ2view source ↗
-
https://t.co/9CTKHbbLuJ https://t.co/9CTKHbbLuJview source ↗
-
@Stefan_3D_AI Interesting that Opus 5.5 can’t do geometry node hair in blender - it’s really really bad at it in all my many tests. The vision model is just not good, that’s truly the limiting part. It “sees” in very low res. It also can’t problem solve hair guides because it can’t get its head around human hair “rul… @Stefan_3D_AI Interesting that Opus 5.5 can’t do geometry node hair in blender - it’s really really bad at it in all my many tests. The vision model is just not good, that’s truly the limiting part. It “sees” in very low res. It also can’t problem solve hair guides because it can’t get its head around human hair “rules”. I also think there’s just not enough training material (ie. Very few finished ultra realistic hairstyles out there as .blend files, aside from the official demo project). And also, there’s very little documentation and the tutorials out there are reliant on an individual’s artistic merit (ie. They aren’t easily reduced to “formula”)view source ↗
-
【今日のAI今北産業|0285:Agents' Last Exam】 AI Agentを「試験問題」ではなく、本物の仕事で測るBenchmarkへ 55分野・13産業Cluster、長時間の実務WorkflowをSandbox上で最後まで実行 高得点Benchmarkと「本当に仕事を任せられるか」のGapを測る ← イマココ Agents' Last Exam。 略してALE。 名前だけ聞くと、 「AI Agent向けの難しい試験」 に見えます。 でも本質は少し違います。 ALEが測ろうとしているのは、 AIが難しいQuestionに答えられるか ではありません。 現実の仕事を渡した時に、 必要なSoftwareを使… 【今日のAI今北産業|0285:Agents' Last Exam】 AI Agentを「試験問題」ではなく、本物の仕事で測るBenchmarkへ 55分野・13産業Cluster、長時間の実務WorkflowをSandbox上で最後まで実行 高得点Benchmarkと「本当に仕事を任せられるか」のGapを測る ← イマココ Agents' Last Exam。 略してALE。 名前だけ聞くと、 「AI Agent向けの難しい試験」 に見えます。 でも本質は少し違います。 ALEが測ろうとしているのは、 AIが難しいQuestionに答えられるか ではありません。 現実の仕事を渡した時に、 必要なSoftwareを使い、 Fileを扱い、 調査し、 判断し、 何度も操作し、 成果物を作り、 条件を確認し、 最後まで本当に仕事を終えられるか。 です。 2026年6月、UC Berkeley RDIを中心とする研究チームがAgents' Last Examを公開しました。 Paperの問題意識はかなり明確です。 AIはさまざまなBenchmarkで急速に高得点を取るようになった。 それなのに、 多くのProfessional Domainでは、 そのBenchmark性能が同じ速度でEconomically MeaningfulなDeploymentへ変わっているわけではない。 なぜか。 研究チームは、 そのGapの大きな原因の一つが「Evaluation」にあると考えました。 つまり、 測っているものが実際の仕事と違う。 そこで作ったのがALEです。 ■従来Benchmarkと何が違うのか 従来のBenchmarkには大きな価値があります。 数学問題を解く。 知識Questionへ答える。 Codeを書く。 Bugを直す。 Browserで操作する。 Terminal Taskを処理する。 AI Capabilityの進歩を見るには必要です。 ただし現実の仕事は、 1問答えて終わり ではありません。 例えばProfessional Workでは、 依頼内容を理解する。 Input Fileを確認する。 必要なDataを探す。 Softwareを立ち上げる。 複数Fileを編集する。 途中結果を検証する。 間違っていたら戻る。 別Toolへ移る。 最終Deliverableを指定形式で保存する。 という長いWorkflowになります。 ALEは、 このLong-horizon Workを測ります。 ■問題ではなく「Project」を渡す ALEのTaskは、 SyntheticなPuzzleだけを作るのではなく、 Industry Expertが実際に経験・完了したProfessional Workflowをもとに設計されます。 Paperでは250人超のIndustry Expertと開発。 現在の公式Framework Documentationでは、 55 Subdomain、 13 Industry Cluster、 1,000件超のTask Pool、 約150件のPublic Release と説明されています。 公開時のBerkeley RDI Blogでは、 1,500件超のExpert-sourced Taskとして紹介されていました。 数字が違って見えるのは、 ALEが固定Benchmarkではなく、 Task Poolが継続的に増減・更新されるLiving Benchmarkとして設計されているためです。 重要なのは、 一つのCoding領域だけではないこと。 Computing。 Engineering。 Science。 Health。 Business。 Finance。 Law。 Education。 Visual/Media系など、 幅広いNon-physical Professional Workを扱います。 分類には米国のO*NET/SOC 2018 Occupational Taxonomyを参照しています。 つまり、 「AI AgentはSoftware Engineerの仕事ができるか」 だけではなく、 Professional Digital Work全体でどこまで働けるか を測ろうとしています。 ■GUIもCLIも使う 現実の仕事では、 Terminalだけ使うわけではありません。 Browser。 Office Software。 Specialized Application。 GUI。 CLI。 File System。 Web。 全部を行き来します。 ALEではReproducible Dview source ↗
-
Your money. Your choice. https://t.co/9WGOD4CLfx Your money. Your choice. https://t.co/9WGOD4CLfxview source ↗
-
i think very soon llm outputs cannot be simply graded by human readings - right now we judge model outputs by checking things like 1) oh this is more concise 2) oh this made the whole point more relevant 3) oh this is better for documentation, rather than, "what is actually correct, in-depth and future proof". of cour… i think very soon llm outputs cannot be simply graded by human readings - right now we judge model outputs by checking things like 1) oh this is more concise 2) oh this made the whole point more relevant 3) oh this is better for documentation, rather than, "what is actually correct, in-depth and future proof". of course, each task's purpose is also different, so being a good judge at all subjects is also very difficult. right now the biggest difference i see of opus 5 vs 5.5 is decision making, 5 is indecisive leading to many self-contraditions, where 5.5 makes sure to check corners and decide on a most-probably-correct outcomeview source ↗
-
https://t.co/vpdLGQLgEF https://t.co/vpdLGQLgEFview source ↗
-
https://t.co/vpdLGQLgEF https://t.co/vpdLGQLgEFview source ↗
-
84 days from breach to an email in a public inbox. This week's AI briefing: agents, authorization, and who finds out when a boundary gets crossed. https://t.co/bTWr95NyM0 84 days from breach to an email in a public inbox. This week's AI briefing: agents, authorization, and who finds out when a boundary gets crossed. https://t.co/bTWr95NyM0view source ↗
-
Should we still pre-process data for BI agents? https://t.co/JRIy4rypJI Should we still pre-process data for BI agents? https://t.co/JRIy4rypJIview source ↗
-
For the decompilation crowd, I finished decompiling Wave Race 64 a couple of weeks ago. My repo was turned over to the team in charge of the official repo and there will be an announcement in the future once it's release worthy. Switching efforts to SSX 3 (GameCube). Might make this one public. It's at 21% already aft… For the decompilation crowd, I finished decompiling Wave Race 64 a couple of weeks ago. My repo was turned over to the team in charge of the official repo and there will be an announcement in the future once it's release worthy. Switching efforts to SSX 3 (GameCube). Might make this one public. It's at 21% already after a few days. Opus 5.5, GPT-6-Astra, and Fable 5.1 are getting a workout. * Don't ask for the Wave Race 64 repo. I respect the team that's working on it and will let them finish that project.view source ↗
-
https://t.co/mqpme74K0f https://t.co/mqpme74K0fview source ↗
-
I really do not think the labs are “making the models dumber” I think what’s a lot more likely, you have a very long running session with multiple compacts, performance would obviously be degraded, hallucinations up, etc. your harness is full of bloat and BS. I’d try running an audit of your harness against anthropic d… I really do not think the labs are “making the models dumber” I think what’s a lot more likely, you have a very long running session with multiple compacts, performance would obviously be degraded, hallucinations up, etc. your harness is full of bloat and BS. I’d try running an audit of your harness against anthropic documentation on prompting for opus 5.5, especially if you have Claude constantly documenting mistakes it makes multiple times, gotchas, etc… try starting new sessions more frequently, I typically will do one compact before handing over to the next.. just food for thought, the harness matters a lot more than the model. For an industry in an up hill battle with the public for trust… I find it hard to believe they would are “nerfing” models, I have not noticed the same degradation problems as others, opus 5 was just a bad model. 5.5 has been an absolute pleasure to build with, though I’m sure your response to this will be something along the lines of “anthropic shrill”view source ↗
-
•https://t.co/YDUW7toWed’s sitemap leaked Kimi K4 as Moonshot AI’s next major model, alongside GLM-5.5 Flash and DeepSeek V4.1 Pro, ahead of any official announcement. •The Kimi K4 data page is still live on https://t.co/gRNQEKKXeO with 0 tokens and unknown specs, sitting like a placeholder waiting for launch day. •Aft… •https://t.co/YDUW7toWed’s sitemap leaked Kimi K4 as Moonshot AI’s next major model, alongside GLM-5.5 Flash and DeepSeek V4.1 Pro, ahead of any official announcement. •The Kimi K4 data page is still live on https://t.co/gRNQEKKXeO with 0 tokens and unknown specs, sitting like a placeholder waiting for launch day. •After heavy hype around Opus 5.5 and Jev AI, the leak has now put Kimi K4 in the spotlight.view source ↗
-
https://t.co/98UyhASdsz https://t.co/98UyhASdszview source ↗
-
I don’t know why Sonnet 5 is currently the default model in claude - if in their own docs they advise keeping the default on Opus 5.5. And just look at the price/intelligence. Even in the documentation they recommend: > For most agent workloads, start with Claude Opus 5.5 at its default effort (medium). https://t.co/2… I don’t know why Sonnet 5 is currently the default model in claude - if in their own docs they advise keeping the default on Opus 5.5. And just look at the price/intelligence. Even in the documentation they recommend: > For most agent workloads, start with Claude Opus 5.5 at its default effort (medium). https://t.co/2QHzcfBKMT So I see the point of switching so that you don’t throw money away. #claude #opus5.5 #anthropic #claudecode #opus #sonet5 #claudeaiview source ↗
-
No major model releases over the last day. Apparently the big tech vendors used all available compute preparing talking points for the Xi state dinner. And what a cluster it was. Elon Musk, Jensen Huang, Lisa Su, Sam Altman, Mark Zuckerberg, Sundar Pichai, Satya Nadella, Tim Cook and Jeff Bezos all checked present for… No major model releases over the last day. Apparently the big tech vendors used all available compute preparing talking points for the Xi state dinner. And what a cluster it was. Elon Musk, Jensen Huang, Lisa Su, Sam Altman, Mark Zuckerberg, Sundar Pichai, Satya Nadella, Tim Cook and Jeff Bezos all checked present for the White House orbit last night. Musk, Huang, Su and Cook were seated at the head table with Trump and Xi. However, no corporate CEOs from China. Anthropic was also absent, and didn't even make the kiddies table. You can guess why. Did they say anything substantive afterward? Mostly, no. Dinner coverage was overwhelmingly handshakes, seating charts and sea bass. Bilateral talks did include AI, and Xi publicly calling for continuing U.S.-China dialogue on AI risks and benefits, preventing misuse and keeping AI under human control. Musk was the exception, but made his interesting comments before the hors d'oeuvres rather than over dessert. In an interview aired by Chinese state media, Musk called Chinese AI models “generally outstanding”. He went on, saying Chinese may be the best work in the world on performance per unit of compute. He estimated China could overcome its lithography and chip-manufacturing constraint in roughly two to three years, mark those as Musk years. cough, cough. Moving on, the latest OpenAI rumor is ChatGPT Pro Max. Tibor Blaho found PROMAX / chatgptpromax references in ChatGPT’s front-end code. The unreleased plan is showing around $500/month, with the key description “Fastest Work and Codex.” A $600 figure in one screenshot appears to include VAT. Which then led to speculation that Cerebras inference could be involved. Reddit immediately reached the obvious conclusion: if OpenAI wants five hundred bucks a month, somebody had better put something considerably more interesting than slightly faster Sol behind the velvet rope. They were more distracted by picking up a useful piece of evidence. which is no longer one Azurview source ↗
-
I went through Anthropic’s official Opus 5.5 + Claude Code guidance and turned it into a practical guide covering prompting, CLAUDE.md, context, subagents, effort levels, verification, and better workflows. Full article below ↓ https://t.co/LzA4svKzA8 I went through Anthropic’s official Opus 5.5 + Claude Code guidance and turned it into a practical guide covering prompting, CLAUDE.md, context, subagents, effort levels, verification, and better workflows. Full article below ↓ https://t.co/LzA4svKzA8view source ↗
-
Hidden Claude Capability 3 (that ChatGPT users miss): The communication quality that reads as human instead of AI. He told them about the quality dimension that no benchmark measures and that the ChatGPT user noticed immediately but attributed to the wrong cause. The ChatGPT user said Claude's writing was "impressive… Hidden Claude Capability 3 (that ChatGPT users miss): The communication quality that reads as human instead of AI. He told them about the quality dimension that no benchmark measures and that the ChatGPT user noticed immediately but attributed to the wrong cause. The ChatGPT user said Claude's writing was "impressive." He noticed the quality. He attributed it to Claude being "a writing tool" implying that Claude trades capability for polish, that the writing quality comes at the cost of the analytical depth ChatGPT provides. He told him the attribution is backwards. Claude Opus 5.5's writing quality doesn't come at the expense of analytical depth it comes alongside it. The same model that writes with natural cadence, varied sentence structure, and direct communication also scores Fable-class performance on coding benchmarks, completes complex refactors in fewer iterations than competing models, and achieves the highest alignment score of any model Anthropic has ever tested. He told him early testers of Opus 5.5 consistently describe the communication as "clearer and more direct" than any previous Claude model and markedly more human-sounding than Astra's output. Anthropic trains for communication quality as a distinct capability not just saying the right thing (accuracy) but saying it in a way that reads as a knowledgeable human rather than a capable machine (expression). He told him the practical impact compounds across every output the user doesn't edit before sending. The email drafted by Claude that requires zero editing before forwarding to a client saves 5 minutes. The report that reads as professional analysis rather than AI-generated summary saves 15 minutes of rewriting. The code comments that read as if a senior engineer wrote them rather than a documentation bot save the entire team's reading time. The communication quality isn't a "writing feature." It's a time-saving feature that applies to every output the AI produces.view source ↗
-
https://t.co/vXQb4VTChk https://t.co/vXQb4VTChkview source ↗
-
ANTHROPIC FLIPPED A SWITCH TODAY: THIRD-PARTY AGENTS CUT OFF WITH NO ANNOUNCEMENT At 3:00 AM Pacific today, Opus 5.5 answered a request arriving through a third-party agent harness, on a paid Claude Max subscription — the arrangement Anthropic's own documentation described as sanctioned as recently as yesterday. At 4:… ANTHROPIC FLIPPED A SWITCH TODAY: THIRD-PARTY AGENTS CUT OFF WITH NO ANNOUNCEMENT At 3:00 AM Pacific today, Opus 5.5 answered a request arriving through a third-party agent harness, on a paid Claude Max subscription — the arrangement Anthropic's own documentation described as sanctioned as recently as yesterday. At 4:46 PM Pacific, that same login, same model, same account, same harness was refused. HTTP 400: "Third-party apps now draw from your extra usage, not your plan limits." No announcement. No changelog. No email. Nothing on Anthropic's developer account about harnesses at all today. WHAT CHANGED The gate now keys on client identity. Requests issued by Claude Code itself go through. Requests from a third-party harness driving that same login do not. We verified it directly: minutes apart, one credential, two paths. Plain Claude Code at the terminal answered. The harness path returned 400. Our measurement: last successful Opus turn 03:00:03 PDT, first refusal 16:46:24 PDT, zero config changes on our side in between. Measured two-day Opus consumption across the whole episode: roughly $3. This was not an account running out of allowance. It was a door closing. It isn't only us. CLIProxyAPI filed the same finding today — non-Claude-Code clients now rejected on subscription accounts, 429s on Opus 5 and 5.5, the identical 400 on Haiku. Hermes opened an "Anthropic 400 error" issue today and already has a fix PR in flight. The third-party agent ecosystem got hit as a class. THE PATTERN IS THE STORY Three times in five months, same surface, three different rules, no notice: April 4, 2026 — Anthropic prohibits third-party harnesses from drawing on subscriptions. Operators are pointed at pay-as-you-go API billing. June 15, 2026 — Anthropic announces a dedicated monthly Agent SDK credit instead: $20 on Pro, $100 on Max 5x, $200 on Max 20x. Non-rollover, billed at API rates, spent before anything else, hard stop when empty unless usage credits are enabled. Theview source ↗
-
https://t.co/KeC9ONLJLl https://t.co/KeC9ONLJLlview source ↗