XX / Twitter50
◈Bluesky1
listening to
Opus 5.5 model documentation
51
mentions
tracked
tracked
net sentiment
+37%
25 positive
20 neutral
6 negative
volume & sentiment over time
Sep 23–28, 2026 · Daily
sources
most discussed
claude / pricing
active days
6
what people keep raising
- claude / pricing 27
- models / architecture 3
- pricing / anthropic 3
- agents 2
- think / human 2
mentions
Showing 1–25 of 50 mentions
-
X X / Twitter 4d ▲ 2 positive models / architectureYesterday, I did something crazy, deleted 10,000+ lines from my project repo. ▶Majorly all documentation and otherwise redundant tests that models keep making over the time, a deep clean, you can say! I decided to do something about my bloated documentation (Including AGENTS. md) after reading a lot about how latest i… Yesterday, I did something crazy, deleted 10,000+ lines from my project repo. ▶Majorly all documentation and otherwise redundant tests that models keep making over the time, a deep clean, you can say! I decided to do something about my bloated documentation (Including AGENTS. md) after reading a lot about how latest intelligent capability improvement in models lile opus 5.5 and Astra→ that folks have been talking about taking down the 'plan mode' from their harnesses (CRAZYY!!) When I started this project in Google's Antigravity using maybe Gemini 3.5 Flash or something, I created a lot of documentation around the architecture and some nuances using the most intelligeny model of that time: Opus 4.6 at that time in Antigravity, which is still there (Hopeless team maintain or governing the decisions for model selection). For the past two to three months, I've had these documentation trying to gauge and help my models maneuver and understand the priorities I have. But I realized that maybe it is creating a lot of noise and a lot of unnecessary context for the really intelligent models like Astra that I have started using, so that they are not unintentionally restrainted to deliver the value that they possibly can. So I've deleted every documentation (because I am a solo founder, I know exactly they were and what I what in my head) and only kept an agents. md file that contains the most important rules, more like a philosophy of the project. From today, I will be using a freshly installed Codex with slight personal configuration change along with Astra-reformed Pstack. Which is a Codex version of 'PStack by Lauren Tan @poteto ' at SpaceX, and I triple verified my version with Astra adversarial judge to perform a comparison between using Lauren's skills and the new Codex plugin that I've created to see if the inspired plugin that has been created by Astra is performing with the philosophy that Lauren has made PStack with OR not. And so now it's a new time to seeview source ↗
-
X X / Twitter 4d ▲ 1 positive claude / pricingCompare Opus 5.5 vs GPT 6 Sol for coding, writing, agents, and marketing. See which AI model delivers better quality, speed, and value. #OpenAI #Claude #GPT6 #Opus5.5 https://t.co/S3UFei4bKe Compare Opus 5.5 vs GPT 6 Sol for coding, writing, agents, and marketing. See which AI model delivers better quality, speed, and value. #OpenAI #Claude #GPT6 #Opus5.5 https://t.co/S3UFei4bKeview source ↗
-
X X / Twitter 4d ▲ 0 positive claude / pricing🎙️ Claude Opus 5.5 Is Now Live on https://t.co/He4LVLveJn: Built for Long-Running AI Agents 🚀🤖 1️⃣📢 A New Claude Generation Arrives on https://t.co/He4LVLveJn https://t.co/He4LVLveJn has added Claude Opus 5.5, described as the first model in the new Claude 5.5 family. Its positioning is notably different from a gene… 🎙️ Claude Opus 5.5 Is Now Live on https://t.co/He4LVLveJn: Built for Long-Running AI Agents 🚀🤖 1️⃣📢 A New Claude Generation Arrives on https://t.co/He4LVLveJn https://t.co/He4LVLveJn has added Claude Opus 5.5, described as the first model in the new Claude 5.5 family. Its positioning is notably different from a general-purpose chatbot. The model is designed around long-running agentic coding and complex knowledge work, targeting workflows that require sustained reasoning and execution. 2️⃣🤖 Agentic Coding Takes Center Stage The key theme is autonomy. Rather than handling only isolated coding prompts, Claude Opus 5.5 is designed for workflows where an AI agent can work through multiple stages of a task. That makes it relevant to larger software projects, debugging cycles, codebase analysis, and other development processes that cannot be completed effectively through a single prompt. 3️⃣📚 1M-Token Context Window A 1-million-token context window gives the model substantial room to work with large amounts of information in one workflow. For developers, that can be useful when dealing with: 💻 Large codebases 📄 Extensive documentation 🧩 Multiple project files 🔍 Long-running research tasks 🤖 Agent memory and context Large context does not automatically guarantee better output, but it expands the amount of information an agent can potentially consider at once. 4️⃣📝 Up to 128K Output Tokens The model also supports up to 128K output tokens. Combined with the large context window, this is particularly relevant for tasks requiring lengthy reasoning or substantial generated output rather than short conversational responses. The architecture is therefore aimed at workloads where both input and output can become unusually large. 5️⃣💰 Lower Running Cost Than Opus 5 https://t.co/He4LVLveJn states that Claude Opus 5.5 can deliver comparable performance to Claude Fable 5.1 on most tasks while costing approximately 40% less to run than Opus 5. If that cost relationshipview source ↗
-
X X / Twitter 4d ▲ 2 positive models / architectureThis is why LLMS will NEVER be conscious!!! We nailed the toughest function of all. Full character swap or just face swap. It was soooo hard to nail. This at low res. Scaling up is easy since it's already managed. But this was so tough; now it's unlimited. And I even added a new function through my main agent that can… This is why LLMS will NEVER be conscious!!! We nailed the toughest function of all. Full character swap or just face swap. It was soooo hard to nail. This at low res. Scaling up is easy since it's already managed. But this was so tough; now it's unlimited. And I even added a new function through my main agent that can sharpen even more I did not ask it to generate a sad face at all. This is a frame of a video. But look at the man's face. Fully consistent with the video frame itself. It was the toughest of all. The first above is the original. Build 7 is live and finished. The docs and test kit are installed on your laptop, and the employees' briefing is updated (attached above). 2 things I have proven. Continuous learning is only possible through client-side harnesses. Otherwise, all compute starts from a generalized training baseline. Every context compaction is every new session. You see it yourself. This was really tough; it looked so fake. In such a childish way. 2. I also proved that a native open-weight TRANSFORMER model is wich my clipping studio tool is built on. It does not come across as sophisticated at all. I swear it comes like the old Will Smith video that went viral. Because it is one of the last truly open weight model that you don't have to pay for a license and wich is not banned in US & EU like some are. Everything after was paid. But I believed I could build it to the next level. I even invented functions genuinely. Like patching things from other professional software tools. Tried many ways to arrive at my goal. But I will probably use this function like never. I do not know why I would do a face swap if I can just have consistent characters. And have nailed the motion physics already in an earlier phase. And emotional expression. But since I have it in my tool as a perfectionist, it did not feel complete. Everything else came together. And even after these images, what you see now. I have built a large range of FUNCTIONALITIES purely forview source ↗
-
X X / Twitter 4d ▲ 15 neutral claude / pricingModels, Machines and Rockets: The Expanding Competition Among ChatGPT, Claude and Grok. https://t.co/yGKb4BWou0 Models, Machines and Rockets: The Expanding Competition Among ChatGPT, Claude and Grok. https://t.co/yGKb4BWou0view source ↗
-
X X / Twitter 5d ▲ 9 positive claude / pricingClaude Opus 5.5 on xhigh reasoning. Ran all day, 10 hours or so. --- Tools --- I didn't tell Claude to do any of this, it just used the API keys I put in its environment: - Generated images using gpt-image-2.5-flare via API - Generated 3D models using Hunyuan via API, and some from Poly Haven (CC0) - Got the iPhone D… Claude Opus 5.5 on xhigh reasoning. Ran all day, 10 hours or so. --- Tools --- I didn't tell Claude to do any of this, it just used the API keys I put in its environment: - Generated images using gpt-image-2.5-flare via API - Generated 3D models using Hunyuan via API, and some from Poly Haven (CC0) - Got the iPhone Duo model out of Xcode somehow and rigged it in Blender - Did all the motion graphics in Python - Prompted and generated the background music using Lyria 3 Pro via API - Generated the voice clips and sound effects using ElevenLabs via API --- Cost --- Raw API cost would've been ~$265. Actual cost was 10% of the weekly quota (20x Max plan), plus $16 in external API costs. I believe the Max plan gets something like 6x the Pro plan's overall usage, so this should fit in the $20 Pro plan, but you'd hit the 5h limit a few times. Annoying, but you can make it work. Claude usage - Model calls: 1,824 (including 8 review subagents) - Total tokens: 781M - Cache reads: 772M - Cache writes: 7M - Output: 1.5M (0.6M of it thinking) - Cost at Opus 5.5 API prices: ~$233 External APIs - 3D food models (Hunyuan 3D + bake-off, fal): ~$15 - Food photos (OpenAI gpt-image-2.5, ~115 images): free via Codex CLI but ~$14 at API prices - Music, voice, sound effects (Lyria 3, ElevenLabs, fal): ~$3.50 - Image model bake-off (FLUX.2 Pro, Gemini): ~$1 - External subtotal: ~$33 --- Prompt 1 (one-shot app, first draft of video) --- I want you to build an iOS calorie tracker app. It should have a super clean, minimalist design: very design-forward and aesthetic. Sweat all the tiny details and make everything feel super premium. Make every interaction extremely delightful. You can go so far as to write custom Metal shaders, custom UI components, etc., to make everything fluid, unique and delightful. The idea I have here is to make it one unified conversational experience. I should be able to page through different days, each of which is a conversation thread with an agent. Theview source ↗
-
X X / Twitter 5d ▲ 27 neutralhttps://t.co/0s0ZBntV9n https://t.co/0s0ZBntV9nview source ↗
-
X X / Twitter 5d ▲ 32 neutralhttps://t.co/tHGElE3N0s https://t.co/tHGElE3N0sview source ↗
-
X X / Twitter 5d ▲ 4 positive claude / pricing90 seconds on low. 67 minutes on max. Same workout-app request. Different results. What did the extra time buy—and did you need it? My guide to Claude Code’s effort settings: when to iterate fast, when to verify deeply, and what to measure. https://t.co/zg9mlTZlf0 90 seconds on low. 67 minutes on max. Same workout-app request. Different results. What did the extra time buy—and did you need it? My guide to Claude Code’s effort settings: when to iterate fast, when to verify deeply, and what to measure. https://t.co/zg9mlTZlf0view source ↗
-
X X / Twitter 6d ▲ 1 neutralhttps://t.co/RWwHmIiDSw https://t.co/RWwHmIiDSwview source ↗
-
X X / Twitter 6d ▲ 37 neutralhttps://t.co/9CTKHbbLuJ https://t.co/9CTKHbbLuJview source ↗
-
X X / Twitter 6d ▲ 0 negative think / human@Stefan_3D_AI Interesting that Opus 5.5 can’t do geometry node hair in blender - it’s really really bad at it in all my many tests. The vision model is just not good, that’s truly the limiting part. It “sees” in very low res. It also can’t problem solve hair guides because it can’t get its head around human hair “rul… @Stefan_3D_AI Interesting that Opus 5.5 can’t do geometry node hair in blender - it’s really really bad at it in all my many tests. The vision model is just not good, that’s truly the limiting part. It “sees” in very low res. It also can’t problem solve hair guides because it can’t get its head around human hair “rules”. I also think there’s just not enough training material (ie. Very few finished ultra realistic hairstyles out there as .blend files, aside from the official demo project). And also, there’s very little documentation and the tutorials out there are reliant on an individual’s artistic merit (ie. They aren’t easily reduced to “formula”)view source ↗
-
X X / Twitter 6d ▲ 3 negative claude / pricing【今日のAI今北産業|0285:Agents' Last Exam】 AI Agentを「試験問題」ではなく、本物の仕事で測るBenchmarkへ 55分野・13産業Cluster、長時間の実務WorkflowをSandbox上で最後まで実行 高得点Benchmarkと「本当に仕事を任せられるか」のGapを測る ← イマココ Agents' Last Exam。 略してALE。 名前だけ聞くと、 「AI Agent向けの難しい試験」 に見えます。 でも本質は少し違います。 ALEが測ろうとしているのは、 AIが難しいQuestionに答えられるか ではありません。 現実の仕事を渡した時に、 必要なSoftwareを使… 【今日のAI今北産業|0285:Agents' Last Exam】 AI Agentを「試験問題」ではなく、本物の仕事で測るBenchmarkへ 55分野・13産業Cluster、長時間の実務WorkflowをSandbox上で最後まで実行 高得点Benchmarkと「本当に仕事を任せられるか」のGapを測る ← イマココ Agents' Last Exam。 略してALE。 名前だけ聞くと、 「AI Agent向けの難しい試験」 に見えます。 でも本質は少し違います。 ALEが測ろうとしているのは、 AIが難しいQuestionに答えられるか ではありません。 現実の仕事を渡した時に、 必要なSoftwareを使い、 Fileを扱い、 調査し、 判断し、 何度も操作し、 成果物を作り、 条件を確認し、 最後まで本当に仕事を終えられるか。 です。 2026年6月、UC Berkeley RDIを中心とする研究チームがAgents' Last Examを公開しました。 Paperの問題意識はかなり明確です。 AIはさまざまなBenchmarkで急速に高得点を取るようになった。 それなのに、 多くのProfessional Domainでは、 そのBenchmark性能が同じ速度でEconomically MeaningfulなDeploymentへ変わっているわけではない。 なぜか。 研究チームは、 そのGapの大きな原因の一つが「Evaluation」にあると考えました。 つまり、 測っているものが実際の仕事と違う。 そこで作ったのがALEです。 ■従来Benchmarkと何が違うのか 従来のBenchmarkには大きな価値があります。 数学問題を解く。 知識Questionへ答える。 Codeを書く。 Bugを直す。 Browserで操作する。 Terminal Taskを処理する。 AI Capabilityの進歩を見るには必要です。 ただし現実の仕事は、 1問答えて終わり ではありません。 例えばProfessional Workでは、 依頼内容を理解する。 Input Fileを確認する。 必要なDataを探す。 Softwareを立ち上げる。 複数Fileを編集する。 途中結果を検証する。 間違っていたら戻る。 別Toolへ移る。 最終Deliverableを指定形式で保存する。 という長いWorkflowになります。 ALEは、 このLong-horizon Workを測ります。 ■問題ではなく「Project」を渡す ALEのTaskは、 SyntheticなPuzzleだけを作るのではなく、 Industry Expertが実際に経験・完了したProfessional Workflowをもとに設計されます。 Paperでは250人超のIndustry Expertと開発。 現在の公式Framework Documentationでは、 55 Subdomain、 13 Industry Cluster、 1,000件超のTask Pool、 約150件のPublic Release と説明されています。 公開時のBerkeley RDI Blogでは、 1,500件超のExpert-sourced Taskとして紹介されていました。 数字が違って見えるのは、 ALEが固定Benchmarkではなく、 Task Poolが継続的に増減・更新されるLiving Benchmarkとして設計されているためです。 重要なのは、 一つのCoding領域だけではないこと。 Computing。 Engineering。 Science。 Health。 Business。 Finance。 Law。 Education。 Visual/Media系など、 幅広いNon-physical Professional Workを扱います。 分類には米国のO*NET/SOC 2018 Occupational Taxonomyを参照しています。 つまり、 「AI AgentはSoftware Engineerの仕事ができるか」 だけではなく、 Professional Digital Work全体でどこまで働けるか を測ろうとしています。 ■GUIもCLIも使う 現実の仕事では、 Terminalだけ使うわけではありません。 Browser。 Office Software。 Specialized Application。 GUI。 CLI。 File System。 Web。 全部を行き来します。 ALEではReproducible Dview source ↗
-
X X / Twitter 1w ▲ 1 positive think / humani think very soon llm outputs cannot be simply graded by human readings - right now we judge model outputs by checking things like 1) oh this is more concise 2) oh this made the whole point more relevant 3) oh this is better for documentation, rather than, "what is actually correct, in-depth and future proof". of cour… i think very soon llm outputs cannot be simply graded by human readings - right now we judge model outputs by checking things like 1) oh this is more concise 2) oh this made the whole point more relevant 3) oh this is better for documentation, rather than, "what is actually correct, in-depth and future proof". of course, each task's purpose is also different, so being a good judge at all subjects is also very difficult. right now the biggest difference i see of opus 5 vs 5.5 is decision making, 5 is indecisive leading to many self-contraditions, where 5.5 makes sure to check corners and decide on a most-probably-correct outcomeview source ↗
-
X X / Twitter 1w ▲ 10 neutralhttps://t.co/vpdLGQLgEF https://t.co/vpdLGQLgEFview source ↗
-
X X / Twitter 1w ▲ 1 neutral agents84 days from breach to an email in a public inbox. This week's AI briefing: agents, authorization, and who finds out when a boundary gets crossed. https://t.co/bTWr95NyM0 84 days from breach to an email in a public inbox. This week's AI briefing: agents, authorization, and who finds out when a boundary gets crossed. https://t.co/bTWr95NyM0view source ↗
-
X X / Twitter 1w ▲ 0 neutral agentsShould we still pre-process data for BI agents? https://t.co/JRIy4rypJI Should we still pre-process data for BI agents? https://t.co/JRIy4rypJIview source ↗
-
X X / Twitter 1w ▲ 17 neutralhttps://t.co/mqpme74K0f https://t.co/mqpme74K0fview source ↗
-
X X / Twitter 1w ▲ 85 negative claude / pricingI really do not think the labs are “making the models dumber” I think what’s a lot more likely, you have a very long running session with multiple compacts, performance would obviously be degraded, hallucinations up, etc. your harness is full of bloat and BS. I’d try running an audit of your harness against anthropic d… I really do not think the labs are “making the models dumber” I think what’s a lot more likely, you have a very long running session with multiple compacts, performance would obviously be degraded, hallucinations up, etc. your harness is full of bloat and BS. I’d try running an audit of your harness against anthropic documentation on prompting for opus 5.5, especially if you have Claude constantly documenting mistakes it makes multiple times, gotchas, etc… try starting new sessions more frequently, I typically will do one compact before handing over to the next.. just food for thought, the harness matters a lot more than the model. For an industry in an up hill battle with the public for trust… I find it hard to believe they would are “nerfing” models, I have not noticed the same degradation problems as others, opus 5 was just a bad model. 5.5 has been an absolute pleasure to build with, though I’m sure your response to this will be something along the lines of “anthropic shrill”view source ↗
-
X X / Twitter 1w ▲ 0 neutralhttps://t.co/98UyhASdsz https://t.co/98UyhASdszview source ↗
-
X X / Twitter 1w ▲ 1 positive claude / pricingI don’t know why Sonnet 5 is currently the default model in claude - if in their own docs they advise keeping the default on Opus 5.5. And just look at the price/intelligence. Even in the documentation they recommend: > For most agent workloads, start with Claude Opus 5.5 at its default effort (medium). https://t.co/2… I don’t know why Sonnet 5 is currently the default model in claude - if in their own docs they advise keeping the default on Opus 5.5. And just look at the price/intelligence. Even in the documentation they recommend: > For most agent workloads, start with Claude Opus 5.5 at its default effort (medium). https://t.co/2QHzcfBKMT So I see the point of switching so that you don’t throw money away. #claude #opus5.5 #anthropic #claudecode #opus #sonet5 #claudeaiview source ↗
-
X X / Twitter 1w ▲ 2 positive pricing / anthropicNo major model releases over the last day. Apparently the big tech vendors used all available compute preparing talking points for the Xi state dinner. And what a cluster it was. Elon Musk, Jensen Huang, Lisa Su, Sam Altman, Mark Zuckerberg, Sundar Pichai, Satya Nadella, Tim Cook and Jeff Bezos all checked present for… No major model releases over the last day. Apparently the big tech vendors used all available compute preparing talking points for the Xi state dinner. And what a cluster it was. Elon Musk, Jensen Huang, Lisa Su, Sam Altman, Mark Zuckerberg, Sundar Pichai, Satya Nadella, Tim Cook and Jeff Bezos all checked present for the White House orbit last night. Musk, Huang, Su and Cook were seated at the head table with Trump and Xi. However, no corporate CEOs from China. Anthropic was also absent, and didn't even make the kiddies table. You can guess why. Did they say anything substantive afterward? Mostly, no. Dinner coverage was overwhelmingly handshakes, seating charts and sea bass. Bilateral talks did include AI, and Xi publicly calling for continuing U.S.-China dialogue on AI risks and benefits, preventing misuse and keeping AI under human control. Musk was the exception, but made his interesting comments before the hors d'oeuvres rather than over dessert. In an interview aired by Chinese state media, Musk called Chinese AI models “generally outstanding”. He went on, saying Chinese may be the best work in the world on performance per unit of compute. He estimated China could overcome its lithography and chip-manufacturing constraint in roughly two to three years, mark those as Musk years. cough, cough. Moving on, the latest OpenAI rumor is ChatGPT Pro Max. Tibor Blaho found PROMAX / chatgptpromax references in ChatGPT’s front-end code. The unreleased plan is showing around $500/month, with the key description “Fastest Work and Codex.” A $600 figure in one screenshot appears to include VAT. Which then led to speculation that Cerebras inference could be involved. Reddit immediately reached the obvious conclusion: if OpenAI wants five hundred bucks a month, somebody had better put something considerably more interesting than slightly faster Sol behind the velvet rope. They were more distracted by picking up a useful piece of evidence. which is no longer one Azurview source ↗
-
X X / Twitter 1w ▲ 1 positive claude / pricingI went through Anthropic’s official Opus 5.5 + Claude Code guidance and turned it into a practical guide covering prompting, CLAUDE.md, context, subagents, effort levels, verification, and better workflows. Full article below ↓ https://t.co/LzA4svKzA8 I went through Anthropic’s official Opus 5.5 + Claude Code guidance and turned it into a practical guide covering prompting, CLAUDE.md, context, subagents, effort levels, verification, and better workflows. Full article below ↓ https://t.co/LzA4svKzA8view source ↗
-
X X / Twitter 1w ▲ 2 positive claude / pricingHidden Claude Capability 3 (that ChatGPT users miss): The communication quality that reads as human instead of AI. He told them about the quality dimension that no benchmark measures and that the ChatGPT user noticed immediately but attributed to the wrong cause. The ChatGPT user said Claude's writing was "impressive… Hidden Claude Capability 3 (that ChatGPT users miss): The communication quality that reads as human instead of AI. He told them about the quality dimension that no benchmark measures and that the ChatGPT user noticed immediately but attributed to the wrong cause. The ChatGPT user said Claude's writing was "impressive." He noticed the quality. He attributed it to Claude being "a writing tool" implying that Claude trades capability for polish, that the writing quality comes at the cost of the analytical depth ChatGPT provides. He told him the attribution is backwards. Claude Opus 5.5's writing quality doesn't come at the expense of analytical depth it comes alongside it. The same model that writes with natural cadence, varied sentence structure, and direct communication also scores Fable-class performance on coding benchmarks, completes complex refactors in fewer iterations than competing models, and achieves the highest alignment score of any model Anthropic has ever tested. He told him early testers of Opus 5.5 consistently describe the communication as "clearer and more direct" than any previous Claude model and markedly more human-sounding than Astra's output. Anthropic trains for communication quality as a distinct capability not just saying the right thing (accuracy) but saying it in a way that reads as a knowledgeable human rather than a capable machine (expression). He told him the practical impact compounds across every output the user doesn't edit before sending. The email drafted by Claude that requires zero editing before forwarding to a client saves 5 minutes. The report that reads as professional analysis rather than AI-generated summary saves 15 minutes of rewriting. The code comments that read as if a senior engineer wrote them rather than a documentation bot save the entire team's reading time. The communication quality isn't a "writing feature." It's a time-saving feature that applies to every output the AI produces.view source ↗
-
X X / Twitter 1w ▲ 4 neutralhttps://t.co/vXQb4VTChk https://t.co/vXQb4VTChkview source ↗