How We Test Background Removal Quality (Our Rubric)
The 10-image scorecard we use to judge cutout quality - hair, packaging, jewelry, low-contrast failures - plus how browser-local tools trade off against cloud upload workflows.
By Ahmet C. Toplutaş · Last updated August 9, 2026 · Editorial standards
Most background-remover roundups show perfect demo photos. Real catalogs do not. This page documents the practical rubric we use when comparing cutout quality for nobackground and when advising stores on whether a browser-local tool is enough.
Visual examples we score against
These are the same style of cutouts we keep on the homepage gallery. In a real trial we score edges at 100% zoom; the images below are illustrative of product, portrait, and marketplace subjects.

Figure 1 - Product / packaging style cutout. Score boundary accuracy and label corners.

Figure 2 - Portrait / hair stress case. Score flyaways and glasses reflections separately.

Figure 3 - Marketplace-ready subject. After scoring, place on pure white for channel export.
The 10-image test set
- Straight-on portrait with loose hair
- Curly hair against a mid-tone wall
- Matte packaging box, three-quarter angle
- Glossy bottle with room reflections
- Soft-goods item with fringe or pom-pom
- Ring or necklace with thin metal
- Mug or bag with an interior hole (handle)
- White product on near-white table (low contrast)
- Dark product on dark backdrop
- Slightly soft phone JPEG from a warehouse aisle
If a tool only looks good on items 3 and 9, it is not ready for your catalog day.
Scoring rubric (per image, 1-5)
- Boundary: Does the edge follow the true silhouette within a few pixels?
- Holes: Are interior openings (handles, rings) correctly empty?
- Leakage: Is leftover background tinting the edge?
- Usability: Would you ship this at 800px listing width with under two minutes of cleanup?
Sum the scores. Anything under a pre-agreed threshold gets a re-shoot note, not endless reprocessing of the same file.
Browser-local vs cloud upload
Cloud tools can run larger models and offer APIs. They also move pixels off-device. Browser-local tools (like nobackground) keep processing on the laptop tab, warm a model cache for sequential batch days, and avoid credit invoices. Neither architecture invents contrast that was never captured.
How we run a fair trial
Same ten files. Same zoom. Same human scorer. No marketing screenshots. Write the acceptance bar before you look at prices. After the scorecard, read the honest sequential batch notes on batch workflow if catalog volume is your constraint - multi-upload claims that the UI cannot keep are a trust problem, not a feature.
What this site optimizes for
nobackground optimizes for private, account-free PNG masters with on-device AI. It does not pretend to be a full photo studio, a mobile template pack, or a metered API. Use this rubric to decide whether that scope matches your job - and keep the transparent masters when it does.
Worked example: scoring a glossy bottle
Photograph the bottle under two soft lights. Run removal. On a black canvas, look for white speckles along the highlight edge - that is leakage. On a white canvas, look for dark crumbs along the label - that is over-cut. Score boundary and leakage separately so a pretty label does not hide a chewed silhouette.
If reflections contain other products from the table, either retake with a cleaner surround or accept that no consumer model will invent a studio reflection. Honesty in the scorecard beats wishing the tool were magic.
Sampling rates for real teams
For fewer than 50 new SKUs a week, QA 100% at listing size. For 50-300, QA every fifth frame plus every hero. Above that, automate ingestion only after the scorecard stays green for two consecutive weeks. Raising automation before the bar is stable just scales defects.
Publishing this rubric
We publish the rubric so merchants can reproduce the trial without trusting a landing-page collage. If your scores disagree with ours on the same files, tell us which item failed - that feedback is how on-device cutouts improve more than synonym-filled blog posts ever will.
Linking the rubric back to the product
After you score your ten files, try the same set on the homepage tool with a warm model cache. Note first-image latency versus the fifth image - that gap is why sequential catalog days feel faster after warm-up. If local results clear your bar, store PNG masters and finish channel-specific backgrounds downstream.
If local results miss on glass or extreme low contrast, escalate those SKUs to a desktop session instead of declaring the whole architecture useless. Hybrid teams ship more catalogs than purists.
Share this rubric with freelancers so “done” means the same thing in Istanbul, Berlin, and remote Slack threads. Ambiguous handoffs create duplicate masks and silent quality drift.
Recording results so the team learns
Keep a simple spreadsheet: filename, tool, four rubric scores, ship/no-ship, and a one-line note. After twenty SKUs you will see patterns - maybe every low-contrast white-on-white fails, which is a studio fix, not a software switch. Patterns beat anecdotes in vendor meetings.
Attach two crops (original edge vs cutout edge) for any no-ship decision. Future you will not remember why SKU-184 failed in March. Visual evidence keeps freelancers aligned without a live call.
Finally, revisit the spreadsheet when you change cameras or packaging. A rubric that never updates becomes folklore. Treat it like a living QA SOP next to your brand kit.
Use the rubric before every vendor renewal or tool switch. A two-hour scorecard is cheaper than a six-month migration that only looked good on stock demos.
When scores are close between two tools, prefer the one whose privacy and pricing match your worst week - not your demo week.
Ready to try it yourself?
Remove backgrounds from your images for free - no sign-up needed.
Try nobackground Free