This document takes one household job and says how it would actually be built: which part of it is a planner, which part is a model, which part is a force loop, and which named tool does each one.
The job is this. Glasses are standing on a table. The robot has to pick up the empty ones, turn each through 180 degrees, and stand it mouth-down on a drying rack — the ordinary kitchen kind with a flat base and vertical pegs standing up from it. The glasses that still have water in them must be left alone.
It sounds trivial and the moving part of it is trivial. Everything hard about it comes from the object. The glass is transparent, so a depth camera cannot see it. It has to be turned completely over, which most arms cannot do from an arbitrary starting pose. It breaks if squeezed. And a glass with water in it must be spotted before the turn, because after the turn the water is on the floor.
This is for somebody who has read the overview and wants the method choices in it applied to one concrete problem. It is a design document, not code. The glossary has the longer explanation of any term.
Contents#
- The task
- What makes it hard
- The four versions
- Which part uses what
- When an unfamiliar glass turns up
- What will bite you
- What to measure
- Where to go next
1. The task#
The sequence, in the order it happens:
- Find the glasses on the table, and find the rack.
- Decide which ones are empty.
- Choose a free peg.
- Grasp an empty glass and lift it clear.
- Confirm it is empty.
- Turn it 180 degrees, so the mouth points down.
- Lower it over the peg until the rim reaches the rack base.
- Let go.
- Check that it is standing.
Every number below comes from one example setup, which is also what the pictures are drawn to. Measure your own before using any of them.
| Thing | Number |
|---|---|
| Glass rim, outside and inside | 70 mm and 64 mm |
| Glass height, wall thickness | 120 mm, 3 mm |
| Glass mass, empty and full | 218 g and 563 g |
| Peg diameter, peg height | 12 mm, 100 mm |
| Pegs on the rack, spacing | 6 in two rows of three, 90 mm apart |
The arm is an ordinary six-axis arm bolted to the table, with a repeatability of around a tenth of a millimetre. Saying that early saves wasted effort: nothing here fails because the arm cannot hit a position accurately enough.
An attempt succeeds when the glass is mouth-down on a peg, still standing ten seconds later, with no neighbour moved and nothing broken. A full glass left untouched is also a success. Water anywhere is a failure, however the run ended.
2. What makes it hard#
Five things, each solved by a different part of the system.
The depth camera cannot see it#
A depth camera, meaning one that reports a distance for every pixel, sends light out and measures what comes back. Most of that light goes straight through a glass, and what does not is bent sideways by the curved wall. The depth picture therefore has a hole exactly where the glass is.
The ordinary camera picture still shows the glass, in its edges and its highlights. So the outline comes from a segmentation model run on that picture, and depth is demoted to two jobs it can still do: measuring the table plane, and confirming a result the model has already produced.
A glass with water in it must be caught before the turn#
There are two signals and they are worth using in order. Looking for a water line in the picture is cheap and can be done before the arm touches anything. Weighing is certain but only available after the lift, because the arm is the scales.
Weighing means reading the payload from the joint torques, or from a wrist force sensor, with the glass lifted clear. The margin is large — 218 g against 563 g — so the threshold is not delicate, and it is set low enough that a glass with a mouthful left in it is still refused.
The turn has to be planned backwards#
Turning the glass over is 180 degrees about a horizontal line. The last joint of most arms has a limited range, commonly plus or minus 175 degrees, and 180 degrees of turn does not fit into that if you start in the middle of it.
The fix is to turn the wrist back before closing the fingers. That means the grasp has to already know the release orientation, which is the thing most people get wrong the first time: they write a grasp, then a turn, and find the turn impossible.
The tolerance is not where you would guess#
Getting a 64 mm glass over a 12 mm peg allows 26 mm of sideways error. The real constraint is the rack filling up.
Pegs 90 mm apart and glasses 70 mm across leave 10 mm each side. Because the glass is tall, orientation spends that budget faster than position does: a 5-degree tilt swings the far end 10.5 mm sideways. So the descent must be vertical, and holding the glass upright matters more than placing it exactly.
Squeezing it wrong breaks it, once#
There is a band of grip force that works. Below it the glass slips; above it the rim cracks. The upper edge belongs to the glass and cannot be moved. The lower edge can, by using fingers with more friction, and that is the only lever there is.
Wet glass, which is what comes out of a sink, pushes the lower edge up and narrows the band further.
3. The four versions#
Build it four times, the way the learning path does. Each version removes one assumption from the one before, and each is a system you can run and measure on its own. Do not start the next version until the current one fails for a reason you can state out loud.
Version 1: everything known in advance#
Assumes: the glass always starts on a marked spot, every glass is the same, a person puts out only empty ones, and the rack was measured once and has not moved.
What it is: a fixed sequence, with two of its steps controlled by force rather than by position.
for peg in pegs: # measured once, in fill order
move_to(pickup, wrist = -90°) # pre-turned, so the flip will fit
close_gripper(force = grip_force)
if gripper_width > 68 mm: stop("no glass")
lift(); turn_wrist(to = +90°) # the glass is now mouth down
move_above(peg, clearance = 130 mm)
descend_until(vertical_force > 2 N) # the rim has met the base
open_slowly(); retreat(); check_standing()
Stack: ROS 2, the Robot Operating System, with URDF, the Unified Robot Description Format, to describe the arm; MoveIt 2 for the moves; ros2_control with an admittance controller for the descent.
Buys you: a complete working loop in days, a cycle time you can measure, and a rig to test everything that follows. Cannot do: anything if the glass moves, the rack moves, or somebody puts out a full one.
Version 2: the glass can be anywhere, and it might be full#
Adds perception. An overhead camera finds every glass and classifies each one as empty or full; a wrist camera checks the last stretch of the approach; the lift confirms the weight before the turn.
There is more than one way to find the glass, and which one is right depends on how fixed your set of objects is. Read the table as a shortlist, with the recommendation underneath it.
| Option | What it gives you | When to pick it |
|---|---|---|
| YOLO segmentation — You Only Look Once, a fast detector that also outlines what it finds | outlines of a fixed set of classes, plus the empty-or-full label in the same pass | the default, once you have labelled pictures of your own glasses |
| SAM 2 — the Segment Anything Model | outlines of objects it was never trained on | prototyping, and labelling the pictures YOLO will be trained on |
| Grounding DINO with SAM 2 | finds things from a written phrase, such as "drinking glass" | when the set of objects is open, or before you have any labels |
| ClearGrasp, TransCG | fills in the hole a transparent object leaves in the depth picture | when you need the 3D shape and not just the outline |
Start with SAM 2, use it to label a few hundred pictures, then ship YOLO. SAM 2 is class-agnostic and heavy; YOLO is fast and gives you the empty-or-full class in the same pass, which is the class you actually need. What that costs you is a labelling job on your own table, and a model that knows nothing outside the classes you labelled.
Stack, on top of version 1: YOLO segmentation with two classes, empty glass and full glass; Open3D to fit a cylinder to each outline against the measured table plane; an AprilTag fiducial marker — a flat printed pattern a camera can locate exactly — on the rack base; BehaviorTree.CPP for the sequencing, because this is the version where recovery branches start to multiply.
Then it has to touch the glass. Perception only says where the glass is. Four of the nine steps in section 1 — grasping, confirming it is empty, coming down onto the peg, and letting go — are settled by force rather than by geometry, and the last step is settled by looking again. That is the difference between a sequence that runs and a sequence that works.
Read the table as one row per moment: what it decides, and what gives you the number.
| Moment | What it decides | What provides it |
|---|---|---|
| Closing on the glass | how hard to squeeze, so it neither slips nor cracks | gripper_controllers in ros2_controllers, commanded as a force rather than as a width |
| While carrying it | whether it is slipping | the same controller's reported finger width, which keeps closing if the glass is sliding through |
| Just after the lift | whether it is empty | force_torque_sensor_broadcaster, publishing the wrist wrench, with the gripper's own weight taken off |
| Coming down onto the peg | when the rim has met the base | admittance_controller, or cartesian_controllers, moving on the force instead of to a height |
| Before opening the fingers | whether the rack is now carrying the glass | the same wrench: if the load has not transferred, the glass is hung up, so lift away rather than let go |
| After retreating | whether it is actually standing | the overhead camera and the same YOLO pass, asking whether there is a glass on that peg |
Two things are worth taking from that table. The first is that none of it is a new framework: it is version 1's ros2_control with two more controllers loaded and one topic read, which is the practical reason the force work belongs here rather than in something written from scratch. The second is that only one of these costs money. You have to buy the force reading — a wrist force-torque sensor, or an arm that estimates it from joint currents well enough to trust — while the slip check is free, because every gripper already reports where its fingers are.
Buys you: the assumption that hurt most in version 1 is gone. The table can be loaded any way round, and full glasses are handled correctly. Costs you: training data and a labelling job, plus a component whose reasoning you cannot read.
Version 3: a recipe for each glass we know#
You know what your glasses look like, and version 2 tells you where they are. That is enough to write down what to do with each one, and writing it down beats learning it while the set is small and fixed. A glass library is a small file with one record per type, and a recipe is the version 1 sequence reading its numbers from that record instead of from constants.
Three types cover a normal kitchen. Read the table as one row per record, where every column is a number the recipe uses directly.
| Type | Rim, height | Empty, full | Water gate | Most tilt allowed | Where it goes |
|---|---|---|---|---|---|
| Tea glass | 55 mm, 90 mm | 128 g, 273 g | 146 g | 6.4° | front row |
| Tumbler | 70 mm, 120 mm | 218 g, 563 g | 261 g | 4.8° | either row |
| Tall glass | 75 mm, 160 mm | 305 g, 854 g | 373 g | 3.6° | back row |
Two of those columns are worked out rather than measured, and both are worth following. The tilt limit is the 10 mm neighbour budget from section 2 divided by the height, which is why the tall glass is the fussy one and the tea glass is forgiving. The water gate sits an eighth of the way up from empty to full, so that a mouthful of water is still caught.
Those gates are the clearest argument for keeping all of this per type rather than global.
Where to hold it
A record needs one more thing: where on the glass to take hold. It is not obliged to be the middle, and for these glasses the middle is the wrong answer. Four criteria argue about it, and three of them agree.
So the record gains four more fields — grasp height, grip width, approach and force — and the question is settled once, by a person with a ruler, rather than worked out at run time.
For a glass with no record, the grasp comes from the cylinder that perception already
fitted: take pairs of opposing points with opposite surface normals, score them by
friction and by how far they sit from the centre of mass, and keep the lowest one that
clears the table. On a cylinder that is arithmetic rather than a search. The MoveIt 1
package for this, moveit_grasps, has no ROS 2 branch, so it is a short piece of your
own code instead.
There are also grasp-proposal networks — Contact-GraspNet and GraspGen are the open ones — which take a point cloud and return ranked six-degree-of-freedom grasps. They are the right tool for a bin of objects you cannot list in advance, and they are worth knowing exist. They are not right here, for two reasons: a transparent glass gives no point cloud to feed them, and they rank grasps by whether the object will slip, having been trained on objects that cannot break. A grasp across the rim scores well.
However the grasp was chosen, check it rather than trust it. All three checks use sensors the task already has.
| When | What to look at | What it tells you |
|---|---|---|
| As the fingers close | force against finger width | force rising at the width you predicted means the glass is where you thought it was |
| Just after the lift | the torque in the wrist wrench | an unexpected moment means you gripped higher, lower or further off the axis than intended |
| Before the turn | tilt 20 to 30 degrees slowly | whether the glass shifts in the fingers, while a shift is still recoverable |
A failed check costs a regrasp. Not checking costs a glass.
Stack: no new frameworks at all. The library is a file, and the recipe is the version 1 sequence with its constants replaced by lookups. The perception from version 2 gains one job: say which of the three types it is looking at, which for these three is a question of size and is already answered by the cylinder fit.
Buys you: three glass types handled properly, with numbers you can read, change and argue about, and no training run anywhere. Costs you: it does not extend by itself. Every type is a record somebody wrote, and section 5 is about the fourth one.
Version 4: an instruction decides the goal#
If the job is only ever "put the glasses on the rack", a language model adds nothing and you should not fit one. There is a single goal, it never changes, and a sentence describing it is strictly worse than a constant.
It earns its place when the goal changes, and the natural way for that to happen here is sorting. Real instructions look like this:
- "Put the tea glasses on the front row and the tall ones at the back."
- "The wine glasses don't go on the rack — put them on the tray."
- "Leave the tea glasses out, we're using them."
- "The rack is full. Stack whatever is left on the tray."
Each of those changes which glasses are picked, in what order, and where they end up. None of them changes how a glass is held. That split is the whole design.
Stack: an open vision-language model such as Qwen3-VL reads the instruction and the overhead picture and writes a short programme that calls the version 3 recipes. That pattern — the model writes the plan, ordinary code runs it — is the one Code as Policies established, and the learned methods document covers the family properly.
Buys you: goals that change daily without a code change, and an operator who does not have to be a programmer. Costs you: a new kind of failure. The model will cheerfully produce a plan the cell cannot carry out — a peg that is taken, a glass type it has never seen, a tray that is not there — and nothing downstream will notice that the goal was wrong, only that some step failed. So a checker sits between the model and the recipes, in ordinary code, and refuses anything that does not match the library and the current state of the rack.
4. Which part uses what#
This is the whole shortlist in one place. Read it as one row per job: what we use, and the obvious alternative we are not using, with the reason.
| Job | What we use | Rather than |
|---|---|---|
| Simulate it first | Gazebo | MuJoCo — better contact, but simulated cameras and ROS in the loop matter more here |
| Describe the arm | URDF | a bespoke model nothing else can read |
| Find the glasses | YOLO segmentation | colour thresholding, which has nothing to work with on a transparent object |
| Label the training pictures | SAM 2, or Grounding DINO with a phrase | outlining a few hundred glasses by hand |
| Tell empty from full | the same model, two classes, then confirmed by weight | vision alone, because the turn cannot be undone |
| Fill the depth hole, if needed | ClearGrasp, TransCG | waiting for the depth camera to get better |
| Fit the shape and the table | Open3D | PCL, the Point Cloud Library — capable, but heavier than this needs |
| Locate the rack | an AprilTag marker, read by apriltag_ros | FoundationPose, unless you cannot glue a marker on |
| Plan the motion | MoveIt 2 | hand-written waypoints, which stop working the day the rack moves |
| Plan grasp, turn and place together | MoveIt Task Constructor | planning each stage separately, then finding the grasp forbids the release |
| Drive the joints | ros2_control | your own control loop, where the hard part is the timing |
| Descend onto the rack | the admittance controller in ros2_controllers, or cartesian_controllers | commanding a height into a rigid base |
| Sequence and recover | BehaviorTree.CPP | a state machine, which turns illegible once recovery branches multiply |
| Hold the per-glass numbers | a file, one record per type | constants spread through the code |
| Choose where to hold an unfamiliar glass | antipodal sampling on the fitted cylinder, written yourself | Contact-GraspNet, GraspGen — they need a point cloud a transparent glass does not give, and they rank by slip rather than by breakage |
| Turn an instruction into a goal | Qwen3-VL in the Code as Policies pattern | a menu of buttons, which cannot cover what people actually ask for |
| Learn a skill, if recipes run out | LeRobot | the original ACT and diffusion policy repositories, which are quiet now |
| Record every attempt | rosbag2, which ships with ROS 2 | log lines, which cannot show you the frame before the drop |
The hardware choices are shorter, and only four of them matter.
| Part | What we use | Why |
|---|---|---|
| Gripper | two-finger parallel, with soft silicone pads | the pads raise friction, which widens the low side of the force band; suction fails on a curved, wet, inverted object |
| Where it holds | the end that becomes the top after the turn | the fingers then stay clear of the pegs and the neighbouring rims during the descent |
| Cameras | one overhead, one on the wrist | the wrist camera removes the camera-to-arm calibration error, which is the error that actually sinks this task |
| Force | a wrist force sensor, or joint-torque estimates | it is needed twice: to weigh the glass, and to stop the descent on contact |
Three of those carry a cost worth knowing before you commit. MoveIt Task Constructor is a noticeably steeper learning curve than plain MoveIt, and version 1 does not need it. BehaviorTree.CPP adds a second representation, written in XML, that a newcomer must learn before they can read your logic at all. And the admittance controller needs a force reading you trust plus an afternoon of stiffness tuning with nothing to show for it.
5. When an unfamiliar glass turns up#
Version 3 knows three glasses. This section is about the fourth, and it is the question that decides whether the system is a demonstration or a product.
Noticing it is the safety-relevant half#
The dangerous case is not a new glass. It is a new glass treated as an old one — a cylinder fitted to a wine glass returns an answer, and the answer is confident and wrong. Three cheap checks catch nearly all of it, and any one of them failing should stop the attempt:
- the cylinder fit reports a poor fit against the outline;
- the measured rim and height match no record in the library;
- the weight after the lift matches neither the empty nor the full figure for the type the system thinks it is holding.
That third one is the quiet hero, because it runs after the grasp and before the turn, which is exactly where a wrong guess is still recoverable.
Then there are three routes, and you will use all of them#
Read the table as three responses to the same event, in the order you would reach for them.
| Route | What happens | What it costs | When it is right |
|---|---|---|---|
| Refuse it | the glass is left on the table and flagged for a person | nothing | always the default, and the only correct answer for a wine glass, which cannot go on a peg at all |
| Measure and add | somebody writes a record, or the robot runs a measuring routine and proposes one | minutes | a new size of a shape you already handle |
| Learn it | demonstrations are collected and a policy is trained for shapes no recipe covers | a few hundred demonstrations and a rig to collect them | when unfamiliar shapes stop being rare |
The measuring routine is worth building, because the robot can fill in most of a record by itself. Segmentation against the table plane gives the rim diameter and the height. Closing carefully at the grasp height gives the width. Lifting gives the empty mass, and the full mass follows from the internal volume. The one number it cannot measure is the force that cracks the rim, because measuring that destroys a glass — so a new record starts with a deliberately gentle grip force, and a person tightens it by hand after watching a few runs.
Route three is where LeRobot comes in, and the learned methods document covers what it involves. It is the right answer when the library stops keeping up — a café with forty kinds of glassware, rather than a kitchen with three.
Across all three routes the pattern is the same, and it is the real answer to the question: a new size is nearly free, a new shape never is. Changing size changes numbers. Changing shape changes which surfaces can be held and which way up the thing can stand, and no amount of training data turns that into the same problem.
6. What will bite you#
- Broken glass is not a retry. It leaves shards, and an arm that will happily carry on moving through them. The response to a drop is to stop the cell and call a person. Decide that before writing the recovery logic, because it changes its shape.
- Water is worse than breakage. It spreads, it reaches the electronics, and it is a slip hazard for the people nearby. That is why there are two gates before the turn rather than one.
- A nearly-empty glass looks empty. A centimetre of water in the bottom is hard to see and easy to weigh, which is the whole argument for the second gate.
- One water threshold will not do. A full tea glass weighs less than an empty tall glass, so the gate has to come from the library rather than from a constant.
- Calibration drift eats your margin. You have 10 mm between rims. A 3 mm error between where the camera thinks the rack is and where it is has taken a third of it before the arm has moved.
- Wet glass is a different object. It slips at forces a dry one holds at. If the glasses come from a sink, every grip number has to be measured wet.
- The rack moves, and it fills. It is light plastic on a table. Fix it down or find it every cycle — and fill the pegs in an order that keeps the occupied ones away from the next one for as long as possible.
- Simulation will not give you the grip numbers. Rigid fingers on a thin rigid shell is close to the worst case for a physics engine. Use the simulator for reaching, planning, the turn and the clearances; get the grip from a real glass.
7. What to measure#
Robot demonstrations are easy to make look good, so decide what you are counting before you start. These six numbers are enough.
| Number | How to count it |
|---|---|
| Success rate | glasses standing ten seconds after release, over attempts |
| Full glasses correctly left alone | over full glasses presented — a miss here is a wet floor |
| Breakages per thousand attempts | the number that decides whether this can ever ship |
| Unfamiliar glasses correctly refused | over unfamiliar glasses presented — the measure of section 5 |
| Cycle time | from reaching for the glass to the arm clear of the rack |
| Failures caught before release | as a share of all failures — whether the checks are working |
The last row is the one people leave out and the one to watch most closely. A system that notices it has the glass wrong and puts it back down is in a different class from one that finds out by hearing it break, even at the same success rate.
Log every attempt: the camera frames, the force trace, the gripper width, and the planned and actual poses. Without those, a failure two weeks from now is a story rather than a bug.
8. Where to go next#
- The overview, and its grid of which method suits which task, is where these choices came from.
- Programmed methods covers the planning, the behaviour trees and the force control used in versions 1 to 3.
- Learned methods covers language as a planner for version 4, and learning from demonstrations for the third route in section 5.
- The learning path has five projects to build in simulation; projects 1 and 2 between them cover most of what this task needs.
- Tools and libraries is the fuller version of section 4.
9. Quesitons#
- How can we we train the model on where to pick the glass from? We can use some images and add some markets on where to pick it from, can that be used, if yes how?
- How can we train the gripping of the glass from the handle?
- Can we use reinforced learning here? if yes where?
First Project/Version#
-
we have table and there we have x number of different types of glasses. we also have the rack or glass stand where glasses are put down so that their open face looks down. we will have 6 slots in the rack for now.
-
we also know rack dimension like height of that stand and difference between two parallel glasses
-
we have 5 types of glasses and we need to place the upside down on rack. -- simple straight glass like rocks or vodka glass -- wine glass -- kind of glass with simple handle -- milkshake glass -- irish glass
-
as all type of glass are of different shapes so we have some values predefined like weight of empty glass, then their opening thickness, height etc.
-
what we need is simply pick up the glass and place it on the rack safely. and we need some type of confirmation that it has completed task one by one.
-
one last thing is we should not disturb the glasses which has some water in it. we need to pick only empty glasses.
-
total number and which glass were present, we are not sure.
-
we also dont know the location where rack is present on the table.
-
first we need initial camera images (table camera)
-
also here we need to capture the images of rack as well to find either we have enough empty slots or not. there we can use yolo i think would be better.
-
it will just tell us which slots are empty and which are already booked.
-
there we can use yolo segmentation it will do both things one is giving outline and other is classifiyng the type of glass.
-
the other appraoch could be using of SAM for outline and then use some other ML model for classifcation
-
after we get these details. we need to perform first action which is grupper going and picking up the glass
-
to find exact position where it need to pick. we can use learned policy for that glass only. because the policy knows all glass dimensions and its dynamics
-
then it would be easier for the arm to perform this thing on that type of glass.
-
after picking the glass from that location to upward there we check the weight logic as well by force sensors of the joints
-
after that upside thing is movelt part and placing also
-
but in placing part, after arm done with it, we need to confirm whether it acutally get successful or not by camera image to confirm if that slot has been booked.