← 返回案例库
Case / 2100928490367230201自动化
案例媒体 / 1MP4 ↗
中文参考
很早就访问了@typesafeai Jev,并花了几个小时对其进行了改进。 我不断碰到的问题是:代理跟踪可以告诉您运行了哪个工具,花费了多长时间,返回了什么,但永远不会告诉您代理实际上是否正在进行任何操作。一个代理编辑、测试和恢复同一个文件六次,在日志中看起来非常健康。 所以我制作了JevScope。它观察人工智能代理的工作,并询问Jev每一步实际上意味着什么,这是否与任务一致,是否进展,是否重复,是否陷入困境。每一步都是随着时间的推移而绘制的。 这是一个编码代理陷入比赛条件的情况。注意黄线。 没有日志行写着“我被困住了。“曲线确实如此。 详细的演示、代码和撰写即将推出。
原帖全文 / EN
Got early access to @typesafeai Jev and spent a few hours building on it. The problem I keep hitting: agent traces tell you which tool ran, how long it took, what it returned but never whether the agent is actually getting anywhere. An agent editing, testing and reverting the same file six times looks perfectly healthy in the logs. So I made JevScope. It watches an AI agent work and asks Jev what each step actually means is this aligned with the task, is it progress, is it repeating itself, is it stuck. Every step, plotted over time. Here's a coding agent getting stuck on a race condition. Watch the yellow line. No log line says "I'm stuck." The curve does. Detailed demo, code and writeup coming soon.
- 浏览
- 1217
- 点赞
- 8
- 收藏
- 1
- 转发
- 1
引用帖
Diogo Almeida @CompleteSkeptic
在共同发明ChatGPT后,我不断问自己:为什么超人聊天模型没有导致AGI? 过去两年,我一直在秘密开发一种新的模型训练方法(RLCD),以及我们今天发布的一种新型前沿人工智能模型:Jev ·快20- 200倍 · 40-400x https://t.co/JSybNG2BKJ
查看 X 原帖 ↗来源
本页为独立社区整理,仅展示公开原帖与来源链接。媒体、文本、账号信息及商标权利归原作者和原平台所有。
