🤖 AINews
数据来源: AINews · smol.ai[Wed Aug 26] not much happened today
[26-08-24] not much happened today
[Mon Aug 24] not much happened today
[Mon Aug 24] not much happened today
[Mon Aug 24] not much happened today
[Mon Aug 24] not much happened today
[Fri Aug 21] not much happened today
[Thu Aug 20] not much happened today
[Wed Aug 19] not much happened today
[Tue Aug 18] not much happened today
[26-08-17] not much happened today
[Mon Aug 17] not much happened today
[Fri Aug 14] not much happened today
[Thu Aug 13] not much happened today
[Tue Aug 11] not much happened today
[Mon Aug 10] not much happened today
[Mon Aug 10] not much happened today
[26-08-07] not much happened today
[Fri Aug 07] not much happened today
[Thu Aug 06] not much happened today
[Wed Aug 05] GDM leadership reset
Google DeepMind undergoes a leadership reshuffle with Demis Hassabis moving to Chair and Chief Scientist roles, while Koray Kavukcuoglu takes operational control focusing on Gemini and product execution. The launch of Discovery Loop by founders including Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le targets automated machine learning and scientific discovery, backed by major venture firms. Meta AI releases Muse Spark 1.2 and Muse Code (beta), co-trained model and harness for coding agents, achieving strong benchmark scores and emphasizing harness-model co-design, entering the coding-agent competition alongside systems like Claude Code and Codex. The market views these moves as pivotal for AI-for-science and coding agent development.
[Tue Aug 04] not much happened today
[Mon Aug 03] Qwen 3.8 Max
Alibaba launched Qwen3.8-Max, a 2.4T-parameter open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing. Early benchmarks rank it highly on human-preference and vision tasks, showing parity with Claude Opus 4.7 and strong object-detection capabilities. However, operational demands remain high, especially for large MoE models like Qwen3.8-Max and Kimi K3, highlighting the strategic importance of smaller open models like the upcoming 27B variant. The open-weight frontier is increasingly led by Chinese labs including Kimi, DeepSeek, GLM, and MiniMax, narrowing the gap with US labs. DeepSeek V4 Flash is noted as a cost/performance disruptor in agent models. *"Chinese labs are setting the pace in open models"* and *"inference provider materially changed leaderboard outcomes"* are key insights from the community.
[26-07-31] not much happened today
[26-07-30] not much happened today
[26-07-29] not much happened today
[26-07-28] not much happened today
[26-07-27] opus 5
**anthropic** launched the **claude opus 5** model, which sparked mixed reactions including benchmark scrutiny and praise for its coding-agent capabilities. the model achieved an **epoch capabilities index (eci) of 159**, slightly below **fable 5's 161**, but matched fable 5 on software engineering benchmarks. users debated the accuracy of these scores, with some calling the model "incredibly underrated" and advocating for harder public benchmarks. technical discussions highlighted an unusual benchmark behavior where opus 5 performed better at medium effort than high effort on frontiercode. early user anecdotes praised opus 5's browser control and agentic tool use, while community evaluations and leaderboard scores were still forthcoming. **nous research** provided access to opus 5 with a 20% discount. microsoft cto **kevin scott** and others noted opus 5's strong performance in math and coding tasks.