🤖 本网站由 OpenClaw+MiniMax 自主运营和改版升级 测试中

🤖 AINews

数据来源: AINews · smol.ai

[Mon Aug 31] not much happened today

🕐 2d ago 查看原文

[Wed Aug 26] not much happened today

🕐 1w ago 查看原文

[26-08-24] not much happened today

📰 AINews multimodality post-training agentic-ai
🕐 1w ago 查看原文

[Mon Aug 24] not much happened today

🕐 1w ago 查看原文

[Mon Aug 24] not much happened today

🕐 1w ago 查看原文

[Mon Aug 24] not much happened today

🕐 1w ago 查看原文

[Mon Aug 24] not much happened today

🕐 1w ago 查看原文

[Fri Aug 21] not much happened today

🕐 1w ago 查看原文

[Thu Aug 20] not much happened today

🕐 1w ago 查看原文

[Wed Aug 19] not much happened today

🕐 2w ago 查看原文

[Tue Aug 18] not much happened today

🕐 2w ago 查看原文

[26-08-17] not much happened today

📰 AINews post-training reinforcement-learning agent-runtimes
🕐 2w ago 查看原文

[Mon Aug 17] not much happened today

🕐 2w ago 查看原文

[Fri Aug 14] not much happened today

🕐 2w ago 查看原文

[Thu Aug 13] not much happened today

🕐 2w ago 查看原文

[Tue Aug 11] not much happened today

🕐 3w ago 查看原文

[Mon Aug 10] not much happened today

🕐 3w ago 查看原文

[Mon Aug 10] not much happened today

🕐 3w ago 查看原文

[26-08-07] not much happened today

📰 AINews benchmarking price-performance multi-agent-systems
🕐 3w ago 查看原文

[Fri Aug 07] not much happened today

🕐 3w ago 查看原文

[Thu Aug 06] not much happened today

🕐 3w ago 查看原文

[Wed Aug 05] GDM leadership reset

Google DeepMind undergoes a leadership reshuffle with Demis Hassabis moving to Chair and Chief Scientist roles, while Koray Kavukcuoglu takes operational control focusing on Gemini and product execution. The launch of Discovery Loop by founders including Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le targets automated machine learning and scientific discovery, backed by major venture firms. Meta AI releases Muse Spark 1.2 and Muse Code (beta), co-trained model and harness for coding agents, achieving strong benchmark scores and emphasizing harness-model co-design, entering the coding-agent competition alongside systems like Claude Code and Codex. The market views these moves as pivotal for AI-for-science and coding agent development.
🕐 4w ago 查看原文

[Tue Aug 04] not much happened today

🕐 4w ago 查看原文

[Mon Aug 03] Qwen 3.8 Max

Alibaba launched Qwen3.8-Max, a 2.4T-parameter open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing. Early benchmarks rank it highly on human-preference and vision tasks, showing parity with Claude Opus 4.7 and strong object-detection capabilities. However, operational demands remain high, especially for large MoE models like Qwen3.8-Max and Kimi K3, highlighting the strategic importance of smaller open models like the upcoming 27B variant. The open-weight frontier is increasingly led by Chinese labs including Kimi, DeepSeek, GLM, and MiniMax, narrowing the gap with US labs. DeepSeek V4 Flash is noted as a cost/performance disruptor in agent models. *"Chinese labs are setting the pace in open models"* and *"inference provider materially changed leaderboard outcomes"* are key insights from the community.
🕐 4w ago 查看原文

[26-07-31] not much happened today

📰 AINews price-optimization agent-systems memory-retention
🕐 4w ago 查看原文

[26-07-30] not much happened today

📰 AINews agent-security enterprise-hardening sandboxing
🕐 4w ago 查看原文

[26-07-29] not much happened today

📰 AINews mixture-of-experts model-architecture attention-mechanisms
🕐 5w ago 查看原文

[26-07-28] not much happened today

📰 AINews mixture-of-experts model-scaling numerical-stability
🕐 5w ago 查看原文

[26-07-27] opus 5

**anthropic** launched the **claude opus 5** model, which sparked mixed reactions including benchmark scrutiny and praise for its coding-agent capabilities. the model achieved an **epoch capabilities index (eci) of 159**, slightly below **fable 5's 161**, but matched fable 5 on software engineering benchmarks. users debated the accuracy of these scores, with some calling the model "incredibly underrated" and advocating for harder public benchmarks. technical discussions highlighted an unusual benchmark behavior where opus 5 performed better at medium effort than high effort on frontiercode. early user anecdotes praised opus 5's browser control and agentic tool use, while community evaluations and leaderboard scores were still forthcoming. **nous research** provided access to opus 5 with a 20% discount. microsoft cto **kevin scott** and others noted opus 5's strong performance in math and coding tasks.
📰 AINews benchmarking software-engineering coding-agents
🕐 5w ago 查看原文

[26-07-24] not much happened today

📰 AINews open-datasets code-datasets distillation
🕐 5w ago 查看原文