DeepSeek Harness Real-World Test | One Extra Item Calculated in To-Do List
This time, I’m not testing installation or discussing concepts.
I gave DeepSeek Harness a real work task: Read 3 meeting notes from the same project and organize them into a follow-up to-do list. 🧪
To see if it could complete the task independently, I didn’t put reference answers in the workspace, nor did I allow it to search online for additional info.
The first round took about 51 seconds. Owners, rescheduled dates, and final budgets were correctly identified, but a subtle issue appeared in the statistics:
⚠️ Harness counted 7 items as "Decided Matters," but after manual verification, only 6 were actionable to-dos.
The extra item wasn’t hallucinated content; it was a misjudgment of criteria:
1. Effective parameters like time and scale were double-counted as to-dos;
2. A suggestion to "interview candidates" was split into a new task.
In the second round, instead of redoing it, I added 3 rules: 📌
1. Only actions requiring delivery, review, confirmation, or assignment count as executable to-dos;
2. Effective times, scales, and budgets serve only as current context;
3. Candidate suggestions should be merged into remarks for corresponding items, not counted separately.
Retest results: 6 executable to-dos, 1 completed, 5 incomplete, 1 undecided suggestion—matching the manual checklist. ✅
Total execution time for two rounds was approx. 70.99 seconds, with estimated model cost around ¥0.15.
This test confirmed one thing: DeepSeek Harness can handle real work, but saying "help me organize to-dos" isn't enough. To make results directly usable, you must define what counts as a to-do, what is background info, and what decisions cannot be made on behalf of humans.
I will continue running real work tasks with DeepSeek Harness to see where it saves time and where human verification is still needed.
#DeepSeek harness #DeepSeek #AI agent
Physix Frontier