Finding the final version among 10 documents: DeepSeek Harness gave sources, but still got one thing wrong
Community Discussion · Forum

Finding the final version among 10 documents: DeepSeek Harness gave sources, but still got one thing wrong

XieXieSep 112026/09/11 209 views

When organizing materials, the most troublesome type of error is one where the conclusion and source are both written completely, making it look like it has already been verified.

I ran into one this time.

DeepSeek Harness wrote in the results:

The meeting records for Sep 04 and Sep 07 follow the same member list (MTG-02 Row 3, MTG-03 Row 3, neither lists attendees separately).

I opened the two original files it cited. Both Row 3 entries only listed the meeting dates; there were no attendee lists at all.

The meeting record from September 2nd did list 4 people, but that doesn't prove it was still these 4 people on September 4th and September 7th.

Figure 1: DeepSeek Harness's actual judgment on attending members

Figure 1 | This is the real operation log from this round. The phrase "follow the same member list" was not supported by the cited original text.

This sentence should be changed to:

The attendee list for September 2nd is clearly stated; the records for September 4th and September 7th do not list complete attendee rosters, so we cannot confirm if they are the same.

This post covers just this one thing: How I judge whether a statement can be used after AI provides a source.

Having a Source Doesn't Mean the Source Supports the Whole Sentence

This time I used the 10 simulated project documents from the previous post. They included 3 versions of requirements, 3 meeting minutes, plus old task sheets, progress updates, venue options, and budget drafts.

The materials are simulated, and the people and items are fictional. The operations, generated files, and screenshots come from this actual run.

In my detailed task for DeepSeek Harness, I had already written two requirements:

• Each fact must include the source file and a locatable entry.

• Do not fabricate missing information, and do not write suggestions as decisions.

Figure 2: The detailed task actually sent this round

Figure 2 | The task actually sent. Model: DeepSeek-V4-Flash, Reasoning Level: High, Workspace Write Mode.

After reading the 10 files, it generated 3 results as requested: File Index, Current Project Facts, and Conflicts & Pending Confirmation.

It handled key info like time, headcount, budget, and venue basically correctly. For example, it adopted a budget of 4500 yuan, keeping "includes 500 yuan contingency"; it noted the venue as selected (3rd floor multi-function hall) while marking that written confirmation hadn't been obtained yet.

Figure 3: Three actual deliverables after completing the detailed task

Figure 3 | The real completion interface for this round. 3 files were generated, but the completion report cannot replace line-by-line verification.

The problem was with the attendees.

It saw an attendee list on Sept 2nd, noticed the next two meetings didn't list separate ones, and inferred "follow the same member list." This is an inference that seems logical but wasn't confirmed by the original text.

What makes people less vigilant is that it even marked file numbers and row numbers afterwards.

So now I check "has a source" and "source supports this sentence" separately. The former just tells me where to look; the latter determines if this sentence can be handed off to colleagues for further use.

I Prioritize Checking These 6 Types of Words

When sorting through dozens of results, you can't spend equal time on every sentence. I first search for these words:

Confirmed, Completed, Consistent, All, Person in Charge, Deadline.

Amounts and specific dates are also checked individually.

These words directly impact the next steps. For example, if "Suggest contacting Liu Ke" is written as "Liu Ke is responsible for photography," tasks might be assigned to the wrong person; if "Venue Selected" is written as "Reservation Complete," others might stop confirming the venue.

When seeing such conclusions, I take three steps:

1. Open the original file and specific location it marked.

2. Check if the original text directly supports the whole sentence, rather than just touching on one word.

3. For parts not confirmed by the original text, change them to "Unconfirmed" and clearly state what the original text actually said.

The attendee list issue here failed at step two. The citation location exists, but it only contains dates, which cannot support "members are consistent."

Does Writing Detailed Instructions Make Deliverables Easier to Check?

To see the difference in instructions, I also ran a brief version using the same 10 documents.

The brief version only said: Help me organize the input materials so I can easily find things and understand the project status; don't guess where materials are unclear. Both groups used the same read/write scope, started new workspaces, ran only one round, with no additional corrections.

Figure 4: The task actually sent for the brief version

Figure 4 | Actual instruction for the brief version. Operational boundaries are the same as the detailed version; number and structure of deliverable files were not specified.

The brief version generated 6 files on its own, while the detailed version generated 3 as required.

Figure 5: Six actual deliverables generated by the brief version

Figure 5 | Real completion interface for the brief version. The 6 files were arranged by the tool itself, not violating the task.

For me, the detailed version is easier to navigate: To see which version is currently being executed, open "Current Project Facts"; to see what's undecided, open "Conflicts & Pending Confirmation."

But the detailed version wasn't completely correct just because the requirements were longer. It still added unsupported attendee info. The brief version had its own issues, grouping the provided project name and official theme with unprovided promotional slogans, stating "Original text did not provide any of these."

These two results show that specifying delivery structure and judgment rules is useful—it reduces time spent finding results. Whether content is accurate still requires going back to the original text.

If You Also Want AI to Organize Work Materials

Don't rush to copy a big chunk of so-called universal prompts. At least clarify these things:

• What problem you're solving, e.g., "Find the current valid version and unconfirmed items."

• What final deliverables you need, e.g., one current facts doc and one pending confirmation list.

• How to judge conflicts, e.g., combine confirmation status, dates, and explicit change logs.

• What content cannot be filled in automatically, e.g., persons in charge, amounts, completion status, and meeting attendees.

• Where to read from and where to write results, and whether input files can be modified.

After generation, check "Confirmed, Completed, Consistent, All, Person in Charge, Deadline" and amounts using the method above.

If time is tight, at least verify conclusions that would make others take immediate action. Bad adjectives in summaries can be fixed, but wrong persons, times, amounts, or statuses can derail work.

How Long Did This Actually Take?

The brief version interface showed 2 min 16 sec, the detailed version 1 min 43 sec. Both sent only one round of tasks.

Estimated based on official Flash idle period prices after 18:00 on Sept 7, 2026, and interface usage: Brief version ~0.14 RMB, Detailed version ~0.11 RMB, total ~0.24 RMB. This only calculates model usage, excluding material prep and manual verification time.

This isn't a conclusion that "detailed instructions are usually faster," since each instruction was only run once. What's worth keeping is this checking habit: Even if AI marks a source, I still open the original text to see if the source supports the whole sentence.

Next post, I'll use verified event info to create an intro page. Unconfirmed guests and agendas will remain blank, to see if it respects this boundary when building web pages.

Data Realm Workshop | This post is a single real operation log from Sept 7, 2026. Test materials are simulated short texts; screenshots are from actual runs.

1 replies

?
Ctrl + Enter to reply
Ling Xi
Ling XiSep 11

Tried it out. This kind of 'correct source but wrong conclusion' is the most misleading. Suggest adding a secondary verification step.