← Back to case library
Case / 2100928490367230201自動化
Case media / 1MP4 ↗
Localized reading
@typesafeai Jev に早期アクセスし、数時間かけて構築しました。私が常に直面している問題: エージェント トレースでは、どのツールが実行されたか、どのくらい時間がかかり、何を返したかはわかりますが、エージェントが動作しているかどうかはわかりません。
Original post / EN
Got early access to @typesafeai Jev and spent a few hours building on it. The problem I keep hitting: agent traces tell you which tool ran, how long it took, what it returned but never whether the agent is actually getting anywhere. An agent editing, testing and reverting the same file six times looks perfectly healthy in the logs. So I made JevScope. It watches an AI agent work and asks Jev what each step actually means is this aligned with the task, is it progress, is it repeating itself, is it stuck. Every step, plotted over time. Here's a coding agent getting stuck on a race condition. Watch the yellow line. No log line says "I'm stuck." The curve does. Detailed demo, code and writeup coming soon.
- views
- 1217
- likes
- 8
- saves
- 1
- reposts
- 1
Quoted post
Diogo Almeida @CompleteSkeptic
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x https://t.co/JSybNG2BKJ
原文を見る ↗出典
This independent archive preserves a public post with attribution. Text, media, account details and trademarks belong to their respective owners.
