WorkBuddy in Practice: Even the Best Table Model Can't Save Your Dirty Data
Community Discussion · Company Watch

WorkBuddy in Practice: Even the Best Table Model Can't Save Your Dirty Data

Da WeiDa WeiSep 242026/09/24 160 views

Got up at five today to train legs, completely wrecked. After training, downed 25 grams of whey, and for breakfast weighed out 180 grams of chicken breast. In the locker room, I scrolled past Xiaomi's released Xiaomi-TabLDM, a tabular foundation model, said to require no task-specific fine-tuning—just throw the table in and it can produce predictions via in-context learning, and on TabArena's regression tasks it ranked with over 80% less training time than the first place. I glanced at it, but my mind was on the three folders on my computer.

In-context learning means no retraining—show the model a few example rows and it can imitate. Same principle as a personal trainer with a new trainee: you don't have to relearn anatomy, just glance at the fitness assessment and know where they should start. But the premise is that the assessment isn't written wrong.

I handle just three tables each week: trainee fitness records, class consumption details, and training volume stats. After using WorkBuddy for a month, my biggest takeaway is that no matter how strong the table model is, it can't save those few rows of dirty data in the table.

First, the detours I took. At the very beginning, I threw all three tables into the chat box together and wrote "help me analyze this week's trainee situation." The output looked lively but was actually all nonsense—it said "overall training situation is good." That's like a trainee going heavy without warming up, form completely falling apart. Later I switched to a pipeline, doing one thing at a time.

The path is roughly like this. Open the WorkBuddy desktop app, click New Task on the left, select File Processing as the task type, then drag the entire folder containing your tables into it. Don't upload one by one—drag the folder. The right side of the interface will list the files it recognizes, showing how many files and what formats. This step is crucial. If the file name is something like "New Microsoft Excel Worksheet (3).xlsx," rename it first—the model recognizes file names faster than content.

Then write the task description. WorkBuddy's description box is the instruction manual for its work; the more it reads like briefing a new hire, the better. I always write in four blocks: who the reader is, data definitions, output format, and forbidden zones. Last Friday's entry I wrote: the reader is the coaching team, merge the three tables by trainee ID, unify weight to kilograms, unify dates to the format 2026-09-19, output is a one-page Word with one paragraph per person, forbidden to write things like "progress is smooth" without concrete points, leave uncertain data blank. After writing, click Run at the bottom right, and results come out in about two minutes.

Next is setting it as a scheduled task—I think this is the most valuable step. In the task details page, there's a Schedule at the top right; click it and select every Friday at 18:00. This way, after I finish teaching my last class on Friday night, the weekly report is already sitting in the output folder, and I only spend ten minutes spot-checking two or three trainees.

I tweaked a few configuration parameters. The output directory defaults to the task's own folder; I changed it to CoachingTeam_WeeklyReport on the shared drive. The overwrite policy defaults to generating a new file each time; I changed it to overwrite with the same name, otherwise one version per week, and after a month there are four weekly reports lying around, and I can't even tell which is the final version. Completion notification is also on, popping up when done, so I know.

I stepped in a pit with permissions. That shared folder was initially editable by everyone, and a part-time coach changed last week's report and submitted it as his own.

Later I changed permissions so the coaching team has read-only access, only I and the store manager have edit rights, and template parameters are only changed on my side; if others want to use it, they copy a version out. Many can view, few can edit.

I've summarized three daily maintenance rules. Before running every Friday, I first scan the raw tables for newly added merged cells and blank rows—these two are high-incidence areas for recognition misalignment. Never change the headers casually. Last month I changed "Class Hours" to "Remaining Class Hours," and all previously saved field mappings became invalid, producing a bunch of empty columns, and I had to redo the mapping. Also, save the written task description as a template file; next time just copy-paste and change the date, much faster than rethinking each time.

Last month I wrote a piece saying that WorkBuddy's core value is turning scattered tasks into deliverable pipelines, and I still think that's true.

Back to that paper. Models like Xiaomi-TabLDM solve "the table is given, how accurately can you predict," while my scenario solves "the table isn't even organized." If the process isn't solved, no matter how strong the model is, it just amplifies errors faster. I used ChatExcel for three weeks and also tried Qwen Office for tables, and the feeling is the same: the smarter the tool, the more you need to control the input first.

If you're just starting to use WorkBuddy for tables, my advice is to start with the smallest step. First create a folder with just one table, write the simplest task description, and get it to run once. After it runs, add a second table, add scheduling, add permissions. Don't try to make it do all your work at once—that's impossible, and it was my biggest lesson in the first month.

My environment is the Windows desktop app, and the shared drive goes through the company intranet, so it may not apply to everyone.

2 replies

?
Ctrl + Enter to reply
PR Merged
PR MergedSep 24

Change the header and all the field mappings fail — this is way too real. Our data team got burned by this too. Later we just locked the headers into the template and didn't allow changes.

Pixel Perfectionist
Reply to PR Merged

Locking the header row isn't enough. When the OP changed "class hours" to "remaining class hours," it directly produced an empty column — that shows the mapping has to be decoupled from field names.