SHOW / EPISODE

7月31日 | 代理凌晨自己发 Slack,工程师在睡觉

Episode 183
13m | Jul 30, 2026

本期内容


这期节目把五件看似独立的事情串在一起:OpenAI 把效率收益让利给开发者、同一模型靠两个参数设置把基准分数翻三倍、Every.to 记录的那个凌晨代理自主发消息的真实场景、Ethan Mollick 提醒我们还在用 2024 年的方式使用 2026 年的工具,以及企业里多个 AI 代理之间无法互信的基础设施困局。听完你会发现,能力已经不是今天最难的问题,如何管理、审计、信任这些代理,才是接下来真正要面对的挑战。


本期要点


- GPT-5.6 推出 Luna、Terra、Sol 三个子模型,模型自我优化带来的效率收益直接折进 API 定价,批量调用成本可观下降

- ARC-AGI-3 测试中,仅开启"保留推理上下文"和"压缩"两个参数,GPT-5.6 Sol 的成绩从只解第一关跃升至六关全通

- OpenAI 内部已用 AI 代理维护自身基础设施,工程师的核心工作正在从写代码转向定义任务边界和审查代理行为

- Ethan Mollick 按任务类型推荐工具而非按模型排名,核心判断是大多数人还在用 2024 年的方式使用 2026 年的工具

- 企业多代理场景下,身份验证、权限控制和操作可追溯是最大瓶颈,五家创业公司正分别从不同层面修这个问题


参考资料


Advancing the price-performance frontier with GPT-5.6 — https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark — https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/

Inside OpenAI's Race to Reinvent Software Development for the Agent Era — https://every.to/

An opinionated guide to which AI to use to do stuff: The Summer 2026 Edition — https://www.oneusefulthing.org/

Enterprise AI agents can't talk to each other, can't be trusted with permissions, and can't be audited — 5 startups are already fixing that — https://venturebeat.com/


---


BearTalk 狗熊有话说播客,始于 2012 年。

订阅地址:https://beartalking.com/page/podcast

Paused
Audio Player Image
BearTalk AI 简讯
Loading...