Two days running SAM: how much of the cutout job can it handle?
Community Discussion · Tracks

Two days running SAM: how much of the cutout job can it handle?

Siqi Draws PPTSiqi Draws PPT1d ago2026/10/02 71 views

I compared SAM with my old Photoshop magic wand plus pen tool workflow, and actually ran through it. Conclusion first: for batch rough screening it really saves effort, but fine edges still need human cleanup.

First, what SAM is. Full name Segment Anything Model, released by Meta's FAIR lab in April 2023, Apache 2.0 license, commercially usable. Positioned as an image segmentation foundation model — that is, cutout, extracting an object from the background and outputting a pixel-level mask. Parameter scale is about 1 billion, model file around 2.4GB, trained on the SA-1B dataset, reportedly with over 1 billion masks. I cross-checked these numbers from the model card and several reviews — no conflicts between sources.

SAM's advanced design allows it to adapt to new image distributions and tasks without prior knowledge, a property called zero-shot transfer.

Preparation didn't take much effort. I didn't set up a local environment — first opened an online demo, uploaded an image, and clicked. Later I borrowed a machine with a GPU and ran it locally once to see the speed difference. Preparation was just two things: an image, and an environment that can run inference.

Getting started was simpler than expected. The interface is a canvas plus a few buttons. Three operations: click on an object with the mouse to get a mask; drag a box around the target to get a mask; or click nothing and it automatically segments the whole image into several pieces. The first two are called promptable segmentation — the click or box is the prompt, i.e., the instruction to the model.

My first click was a bit surprising. I uploaded letter blocks scattered on a wooden table, with fairly busy table texture. I clicked one block, and the mask came out almost instantly, with cleaner edges than the magic wand. The magic wand selects based on similar colors, and easily misses things with gradients like wood grain. SAM is different — it seems to understand that this block is an object, rather than calculating colors. Later I dragged a box around seven or eight blocks at once, and it gave separate masks without blurring them together. This step in the traditional workflow requires cutting them out one by one.

Then the pitfalls came.

Small objects and fine structures don't work. There was a tiny printed letter on a block; I clicked on it, and the mask was the whole block, not that letter. I boxed it again, still the whole block. Thin, elongated structures like lines, hair, and grids tend to break or bring along an extra ring of background.

Transparent and reflective objects also don't work. Glass cups, metal edges — the boundaries it gives are floaty. The human eye can see the cup's outline, but the mask it gives eats part of the table. I tried several points and several boxes, all about the same.

Fully automatic mode can't be used directly. Automatic segmentation cuts the image very finely — one table surface can be cut into dozens of pieces, and it even pulls out shadows separately. Fine as source material, but as a finished product you need to merge manually, and merging isn't faster than cutting it out yourself.

I organized my impressions from these few days into a table.

Dimension SAM Traditional magic wand + pen
Learning curve Low, just click Medium, needs practice
Single-image rough cutout speed Fast, seconds Slow, depends on complexity
Edge precision Good enough for main subject, not for details Manual can reach highest
Batch processing Strong, scriptable Weak
Transparent/reflective objects Unstable Relies on human judgment
Commercial license Apache 2.0 Depends on software license

Conclusion. Recommended for teams doing data annotation and batch rough screening, especially scenarios handling large numbers of images and needing a first-draft mask, where it can pull people out of repetitive labor. Not recommended for single-image fine retouching that needs to be a finished product — for this kind of work it gives a seventy-point draft, and the last thirty points still need a human. Individual users wanting to cut out an avatar or product image might find existing tools less hassle than messing with it.

One more positioning point. SAM itself isn't finished software — it's a model. What's useful are the tools wrapped around it; annotation platforms, cutout apps, and video editors all have traces of it. Its value lies in becoming a foundation piece in the vision field, same logic as those foundation models in the text field.

Looking ahead, my judgment is that segmentation will go from being a cutout tool to the first step in a pipeline. Annotation teams won't use it as a selling point anymore, but will assume it's just there — like how nobody nowadays specifically praises an editor for having an undo key. The quality-check and edge-trimming workflow behind it is where the real differentiation lies, and that part is still human for now. I'll put this prediction here and check back in half a year to see if it holds.

1 replies

?
Ctrl + Enter to reply
Lao Fan
Lao Fan20h ago

I buy that you can draw a box and give seven or eight blocks separate masks at once, but the transparent reflective parts have floating edges, so for cutting out full vehicle exterior parts you basically have to redo it by hand.