Models you download and run, which answer "what is this and which pixels is it on". This is the other half of the methods you write yourself: reach for these when the object is not one colour, not one shape, or not something you can describe with a rule.
The document ends with training a model of your own, because that is what you do when no downloadable model knows your objects — and with the datasets and labelling tools that step needs.
Every licence below was read from the project's own LICENSE file or model card
in September 2026, not from a blog post. Where the code and the weights carry
different licences, both are given, because that catches people out regularly.
The full picture is in licences and
platforms.
Contents#
- Models you can download
- Models that choose where to grip
- Models you would train yourself
- Datasets
- Labelling tools
1. Models you can download#
Everything in this section has weights you can fetch today and run without training anything. Section 2 is the other case: models you train on your own objects.
1.1 Box detectors#
What they are. Models that return a rectangle and a class for each object they recognise. They are the cheapest useful answer, they run fastest, and for a robot working with well-separated objects on a table they are often all you need.
The table lists the ones worth knowing. Read it as: the first column is the model, the second is what it is good at, and the third is the licence you would be agreeing to.
| Model | What it is good at | Licence (code / weights) | Where |
|---|---|---|---|
| Ultralytics YOLO | the easiest to use, by a distance; fast; huge community | AGPL-3.0 — see licences | ultralytics/ultralytics |
| RT-DETR | transformer detector with no need for non-maximum suppression; accurate at similar speed | Apache-2.0 / Apache-2.0 | lyuwenyu/RT-DETR, weights |
| D-FINE | a refinement of RT-DETR, currently among the strongest real-time detectors | Apache-2.0 | Peterande/D-FINE |
| DEIM | a training scheme that improves DETR-style detectors | Apache | Intellindust-AI-Lab/DEIM |
| RF-DETR | Ultralytics-like ergonomics without the AGPL; actively developed | Apache-2.0 | roboflow/rf-detr |
| YOLOX | anchor-free YOLO with a genuinely permissive licence | Apache-2.0 | Megvii-BaseDetection/YOLOX |
| Faster R-CNN, RetinaNet | the classics, in torchvision, trivially available | BSD-3 | pytorch/vision |
Five jobs a box detector suits:
- picking well-separated objects off a table or a conveyor
- counting things, where the box is only needed to say "one here"
- cueing a promptable segmenter, which needs a box to start from
- tracking objects between frames, where a box is enough to follow
- any job where the object is roughly as wide as it is long, so the box fits it
Five jobs it cannot do:
- gripping an odd shape, where the box's corners are not the object
- measuring, since the box of a tilted object belongs to no real dimension
- separating objects that overlap, where boxes overlap too
- anything needing the outline: avoiding a handle, finding a rim, fitting a profile
- finding an object whose class is not in the model's list
1.2 Mask models#
What they are. Models that return the pixels of each object rather than a box. This is the answer a robot usually wants, because it supports measuring and gripping round a shape.
| Model | What it is good at | Licence | Where |
|---|---|---|---|
| Mask R-CNN | the workhorse; well understood; easy to fine-tune | torchvision BSD-3 throughout; Detectron2 code Apache-2.0 but its weights are CC BY-SA 3.0 | pytorch/vision, detectron2 |
| Mask2Former | stronger masks; one architecture for all three kinds of segmentation | MIT, but the repository is archived | facebookresearch/Mask2Former |
| OneFormer | one model trained once, doing semantic, instance and panoptic | MIT | SHI-Labs/OneFormer |
| SegFormer | efficient semantic segmentation | NVIDIA Source Code License — non-commercial | NVlabs/SegFormer |
| mmdetection / mmsegmentation | a large library of implementations to train from | Apache-2.0 | mmdetection, mmsegmentation |
Two notes that matter more than the accuracy numbers. Mask2Former's repository is archived, which means no fixes and no new dependencies supported; its weights still work and the code will rot. And mmdetection has not had a push since August 2024, so while it is still the widest collection of implementations, it is drifting away from current PyTorch.
Five jobs a mask model suits:
- gripping round a shape, where the outline decides where the fingers go
- measuring an object's silhouette, which is the input to the dimension document
- separating touching objects of the same class, which instance segmentation does by design
- excluding a handle, a spout or a label from a grasp
- any job where you need the area of something rather than a box around it
Five jobs it cannot do:
- run at high frame rate on a small computer, which is where a box detector wins
- recognise objects outside its training classes
- produce reliable boundaries on transparent or reflective objects
- give orientation, which a mask does not contain
- tell you anything in millimetres
1.3 Promptable segmenters: the Segment Anything family#
What it is. You give the model a point, a box, or a rough region; it returns the exact mask containing it. It does not name anything. Its strength is the boundary, which is markedly better than anything trained per-class, and its weakness is that something else must decide where to point.
Why this rather than a mask model. Because it works on objects it has never seen, which a closed-set mask model cannot. In a robot cell this is the difference between handling your five known parts and handling whatever a customer puts on the table.
| Model | What it is | Licence (code / weights) | Where |
|---|---|---|---|
| SAM | the original; excellent boundaries; slow | Apache-2.0 / Apache-2.0 | segment-anything |
| SAM 2 | adds video and is faster; the safe default | Apache-2.0 / Apache-2.0 | sam2 |
| SAM 3 | the newest; segments every instance matching a phrase | bespoke "SAM License", weights gated | sam3 |
| SAM 3.1 | a drop-in update to SAM 3, March 2026 | same as SAM 3 | weights |
| EdgeSAM, EdgeTAM | the two that run properly on Apple hardware | EdgeSAM non-commercial; EdgeTAM Apache-2.0 | EdgeSAM, EdgeTAM |
| MobileSAM | a much smaller SAM for embedded use | Apache-2.0 | MobileSAM |
| FastSAM | a fast approximation, built on Ultralytics | AGPL-3.0 | FastSAM |
If the licence matters to you, SAM 2 is the one to reach for: it is the most
recent of the family that is plainly Apache-2.0 in both code and weights. SAM 3's
licence is a Meta community licence that does permit commercial use, but it is a
bespoke agreement with its own acceptable-use terms rather than a standard open
licence, so it needs reading rather than assuming. FastSAM is AGPL because it is
built on Ultralytics, which is the most commonly missed licence inheritance in
this whole field — and note that FastSAM's own README claims Apache-2.0 while its
LICENSE file is AGPL-3.0. When a repository contradicts itself, the licence file
is the one that counts.
SAM 3 is also more than a faster SAM. It does promptable concept segmentation: given a short phrase it returns every instance matching it, which is the job that previously needed Grounding DINO and SAM chained together. That makes it a replacement for the pairing described in section 1.4, at the cost of a licence that is not a standard open one.
Five jobs the SAM family suits:
- objects the robot has never seen and you cannot enumerate
- turning a detector's rough box into an accurate outline, which is the standard pairing
- labelling data: a human clicks, SAM produces the mask, which is how most labelling tools now work
- cluttered scenes where per-class models fall apart
- anything where boundary quality is what limits you
Five jobs it cannot do:
- start a pipeline, since nothing in it decides what to point at
- name the object, which is the entire "what is it" question
- run fast on a small computer, unless you use MobileSAM and accept the drop
- give consistent object identity across frames, without the video variant
- handle transparent objects, whose boundary is genuinely ambiguous in the image
1.4 Open-vocabulary models#
What they are. Models you prompt with words. They were trained on pictures paired with text, so they can find things that were never a class in any list.
| Model | What it does | Licence (code / weights) | Where |
|---|---|---|---|
| Grounding DINO | text in, boxes out; the standard choice | Apache-2.0 / Apache-2.0 | GroundingDINO, weights |
| Grounded-SAM | Grounding DINO for the box, SAM for the mask: text in, masks out | Apache-2.0 | Grounded-Segment-Anything |
| OWLv2 | open-vocabulary detection from Google; strong and simple to run | Apache-2.0 | weights |
| YOLO-World | real-time open-vocabulary detection | GPL-3.0 | YOLO-World |
| Florence-2 | one small model doing captioning, detection and grounding | MIT | weights |
The pairing worth knowing is Grounding DINO plus SAM, usually packaged as Grounded-SAM. Between them they take a phrase and return a mask, with no training and no class list, and both halves are Apache-2.0. For a robot that has to handle objects you cannot enumerate in advance, this is the current default.
Five jobs open-vocabulary models suit:
- objects you can describe but not collect pictures of
- a long tail of rare items, as in a warehouse or a laboratory
- prototypes, where the class list is still changing every week
- taking an instruction in words — "pick up the blue mug" — and acting on it
- generating training labels for a smaller, faster model you then train yourself
Five jobs they cannot do:
- run in a few milliseconds on a small computer, which they are far from
- give repeatable answers to two phrasings of the same request
- distinguish things whose difference has no ordinary name — two similar valve bodies
- work where the object has no common-language description at all, which is most of manufacturing
- offer any guarantee, which is why safety-relevant decisions are not made this way
1.5 Backbones and features#
What they are. Not object finders, but the feature extractors other things are built on. They matter here because a strong backbone with a small trained head is often the cheapest route to a good custom model.
DINOv2 is Apache-2.0 and produces features good enough that a simple classifier on top of them matches models trained end to end. DINOv3 is newer and stronger, but is published under a bespoke DINOv3 licence rather than Apache, so it needs reading before commercial use.
1.6 Open-vocabulary models that are not downloadable#
One correction that catches people regularly. Grounding DINO 1.5, 1.6, 1.6 Pro
and DINO-X have no open weights. The repositories with those names contain
client code for a paid hosted service, and the Apache-2.0 licence on them covers
the client, not the model. Only the original Grounding DINO has downloadable
weights. A great many blog posts present the later versions as though you could
pip install them.
Similarly, Ultralytics YOLO27 is announced and not released, and Depth Anything V2's Giant checkpoint has said "coming soon" for a long time.
1.7 Transparent and shiny objects#
This deserves its own entry because it is the case that defeats everything above, and because the honest state of it is not what people expect.
The classical approach is the depth hole from section 5.7. The reference work is Lysenkov, Eruhimov and Bradski, RSS 2012, which deliberately used the depth sensor's failure as the segmentation cue. The code from that lineage (wg-perception/transparent_objects) is long unmaintained.
Since then the field has moved to learned depth completion, and the licensing is awkward:
| Project | Licence | State |
|---|---|---|
| ClearGrasp | Apache-2.0 | abandoned in 2021; still the standard citation for the problem |
| TransCG | CC BY-NC-SA 4.0 | abandoned in 2022; the largest real dataset, and non-commercial |
| ReMake | MIT | 2026, and the most usable recent option: a monocular depth model plus an instance mask, completing the depth |
| FoundationStereo | NVIDIA, non-commercial | excellent, and not shippable |
Polarisation imaging is frequently suggested for this and, as section 4.5 says, there is essentially no open-source work behind the suggestion.
2. Models that choose where to grip#
A fifth kind of answer, and the one most likely to be suggested when someone hears "objects we have never seen". A grasp model takes a point cloud and returns ranked gripper poses: not what the object is, not how big, but how to hold it.
| Model | Licence | State |
|---|---|---|
| GPD | BSD-2-Clause | the one permissive option, and last touched in January 2022 |
| graspnet-baseline | academic and non-profit, non-commercial only | trained on GraspNet-1Billion, which is also non-commercial |
| Contact-GraspNet | NVIDIA, not a standard licence — read it | widely used, needs CUDA |
| AnyGrasp | no licence file; a key you apply for, tied to a machine | the strongest of them, and the least free |
| Dex-Net / GQ-CNN | UC Regents custom | the original of the family; old |
The licence column is doing a lot of work there. Of five, one is permissive and stale, three are non-commercial or bespoke, and the best is licence-keyed to a machine you register.
Why they are less of an answer than they look. A grasp network is trained to predict will this grip hold — that is, will the object not slip out. That is one of your constraints and usually not the only one. It has no way to know that a wine glass must be held on the stem rather than the bowl, or that a grip above half the glass's height cannot be inverted afterwards, or that the handle is the one place the fingers must not land. Those are task constraints, and there is nowhere to type them in.
So the rule of thumb is: if your object has a sentence — hold the narrowest part below the widest — the sentence beats the network, because it encodes knowledge the network cannot be given. Grasp models earn their place when no such sentence exists, which is the genuinely open-ended case: a tote of mixed unknown goods.
Five jobs a grasp model suits:
- bin picking of mixed, unknown, opaque objects with no shared structure
- warehouse totes, where the next item is genuinely unpredictable
- a first pass at an object family before you have worked out a rule for it
- suction grasping, where the question really is just "is this patch flat enough"
- generating candidates for something else to filter against task constraints
Five jobs it cannot do:
- respect a constraint that is not about slipping — orientation, reachability afterwards, a part of the object that must not be touched
- work on transparent or reflective objects, having no point cloud to read
- explain why it chose a grip, which matters when one fails
- run without CUDA, for four of the five above
- be shipped commercially, for four of the five above
3. Models you would train yourself#
Everything so far assumes somebody else's classes. The moment your objects are specific — your parts, your products — you train.
You almost never train from scratch. You fine-tune: take a model that already knows what edges, textures and objects look like in general, and teach it your classes with a few hundred labelled pictures. The camera area works through this with eighty pictures of one object, which is enough to see it work.
The choices, in the order most people should consider them:
| Route | When it is right | Licence |
|---|---|---|
| RF-DETR | a permissive detector with good ergonomics; the sensible default in 2026 | Apache-2.0 |
| torchvision references | you want no framework at all, just PyTorch | BSD-3 |
| segmentation_models_pytorch | semantic segmentation with a wide choice of backbones | MIT |
| Hugging Face transformers | fine-tuning DETR, Mask2Former, OneFormer and friends | Apache-2.0 |
| Detectron2 | you specifically need its Mask R-CNN recipes | Apache-2.0 code, CC BY-SA 3.0 weights |
| Ultralytics | the fastest path to a working model, if AGPL is acceptable | AGPL-3.0 |
| mmdetection | you need an implementation that exists nowhere else | Apache-2.0, last updated August 2024 |
A newer route is worth knowing because it changes the economics. Use an open-vocabulary model to label your data — Grounding DINO or SAM 3 generating boxes and masks from a text prompt — then train a small, fast, permissively licensed model on those labels. You get a model that runs in milliseconds on a cheap computer, trained on data nobody had to draw by hand.
Five jobs training your own model suits:
- a fixed set of objects that a general model does not know
- anything that has to run fast on a small computer
- distinguishing objects that differ in ways with no common name
- a task where you need to control and version the model's behaviour
- meeting a licence constraint, by training your own weights on permissive code
Five jobs it does not:
- objects that change every week, where you would retrain every week
- a long tail of thousands of rare items
- projects with no way to collect and label a few hundred pictures
- proving anything before the mechanical and lighting side is settled
- one-off jobs, where an open-vocabulary model costs nothing and works today
3.1 Making the training data in a simulator#
If you already have a simulator — and in this repo you do — it can generate labelled data, and the labels are free and perfect because the simulator knows where everything is. For a project whose objects are procedurally generated anyway, this is by some way the cheapest route to a trained model.
| Tool | Licence | What it is |
|---|---|---|
| Kubric | Apache-2.0 | a Blender and PyBullet pipeline built for generating annotated video and image datasets |
| BlenderProc | GPL-3.0 | photorealistic rendering with annotations, from DLR; widely used for BOP-style pose data |
| Gazebo | Apache-2.0 | the simulator you are already running, scripted to vary the scene and dump labels |
| NVIDIA Replicator | closed, and NVIDIA hardware | the industrial version of the same idea |
Note BlenderProc's GPL-3.0: the data you generate is yours, but the tool is copyleft, which matters if you plan to ship a pipeline that embeds it.
Domain randomisation is what makes synthetic data transfer. Rather than trying to make the render look real, you vary everything you are not trying to teach — lighting, colours, textures, backgrounds, camera pose, object placement — so widely that the real world looks like one more variation. The model then cannot latch onto any of the details that differ between simulation and reality, because none of them were ever constant.
The honest caveat is that a model trained purely on synthetic data usually needs either a lot of randomisation or a small amount of real data to fine-tune on, and which one is cheaper depends entirely on how hard your real data is to collect. What simulation will not tell you is the wider version of the same warning.
4. Datasets#
If you are training, you need data, and the licences here are stricter than the code licences. Read this table as: what the annotations allow, then separately what the images allow, because they are usually different and the images are where the restriction bites.
| Dataset | Annotations | Images | Commercial use |
|---|---|---|---|
| Open Images V7 | CC BY 4.0 | CC BY 2.0 | yes — the only large one that is cleanly clear |
| COCO | CC BY 4.0 | Flickr terms, copyright per image | annotations yes, images are your own risk |
| LVIS | BSD | inherits COCO's | as COCO |
| ADE20K | BSD-3 | non-commercial research and education only | no |
| Cityscapes | — | non-commercial | no |
| Objects365 | CC BY 4.0 | academic purposes only, registration required | no |
| SA-1B | research licence | same | no |
| GraspNet-1Billion | CC BY-NC-SA 4.0 | same | no |
The short version for anyone building a product: Open Images V7 is the one you can use without thinking about it. COCO's annotations are fine and its images are a judgement call. The rest forbid commercial use, and a model trained on them inherits the problem.
5. Labelling tools#
| Tool | Licence | Self-hosted | Notes |
|---|---|---|---|
| CVAT | MIT | yes | the standard; has SAM-assisted labelling built in |
| Label Studio | Apache-2.0 | yes | broader than vision; model-assisted via a backend |
| FiftyOne | Apache-2.0 | yes | for looking at and curating a dataset rather than drawing on it |
| X-AnyLabeling | GPL-3.0 | yes | a wide model zoo for assisted labelling |
| labelme | GPL-3.0 | yes | widely assumed to be MIT; it is not |
| Roboflow | hosted service | no | see the note below |
One thing about Roboflow's free tier that is easy to miss and matters commercially: on the free plan your data and models are public on Roboflow Universe. Keeping a dataset private requires a paid plan.