Chapter 4 · Section 7 of 8 · Case study

Case study: standing an empty glass upside down on a drying rack

25 min read28 of 41 in Frameworks & Manipulation
On this page

This document takes one household job and says how it would actually be built: which part of it is a planner, which part is a model, which part is a force loop, and which named tool does each one.

The job is this. Glasses are standing on a table. The robot has to pick up the empty ones, turn each through 180 degrees, and stand it mouth-down on a drying rack — the ordinary kitchen kind with a flat base and vertical pegs standing up from it. The glasses that still have water in them must be left alone.

It sounds trivial and the moving part of it is trivial. Everything hard about it comes from the object. The glass is transparent, so a depth camera cannot see it. It has to be turned completely over, which most arms cannot do from an arbitrary starting pose. It breaks if squeezed. And a glass with water in it must be spotted before the turn, because after the turn the water is on the floor.

This is for somebody who has read the overview and wants the method choices in it applied to one concrete problem. It is a design document, not code. The glossary has the longer explanation of any term.

Contents#

  1. The task
  2. What makes it hard
  3. The four versions
  4. Which part uses what
  5. When an unfamiliar glass turns up
  6. What will bite you
  7. What to measure
  8. Where to go next

1. The task#

The sequence, in the order it happens:

  1. Find the glasses on the table, and find the rack.
  2. Decide which ones are empty.
  3. Choose a free peg.
  4. Grasp an empty glass and lift it clear.
  5. Confirm it is empty.
  6. Turn it 180 degrees, so the mouth points down.
  7. Lower it over the peg until the rim reaches the rack base.
  8. Let go.
  9. Check that it is standing.

Every number below comes from one example setup, which is also what the pictures are drawn to. Measure your own before using any of them.

ThingNumber
Glass rim, outside and inside70 mm and 64 mm
Glass height, wall thickness120 mm, 3 mm
Glass mass, empty and full218 g and 563 g
Peg diameter, peg height12 mm, 100 mm
Pegs on the rack, spacing6 in two rows of three, 90 mm apart

The arm is an ordinary six-axis arm bolted to the table, with a repeatability of around a tenth of a millimetre. Saying that early saves wasted effort: nothing here fails because the arm cannot hit a position accurately enough.

An attempt succeeds when the glass is mouth-down on a peg, still standing ten seconds later, with no neighbour moved and nothing broken. A full glass left untouched is also a success. Water anywhere is a failure, however the run ended.


2. What makes it hard#

Five things, each solved by a different part of the system.

The depth camera cannot see it#

A depth camera, meaning one that reports a distance for every pixel, sends light out and measures what comes back. Most of that light goes straight through a glass, and what does not is bent sideways by the curved wall. The depth picture therefore has a hole exactly where the glass is.

Why a depth camera returns a hole where the glass is
Why a depth camera returns a hole where the glass is

The ordinary camera picture still shows the glass, in its edges and its highlights. So the outline comes from a segmentation model run on that picture, and depth is demoted to two jobs it can still do: measuring the table plane, and confirming a result the model has already produced.

A glass with water in it must be caught before the turn#

There are two signals and they are worth using in order. Looking for a water line in the picture is cheap and can be done before the arm touches anything. Weighing is certain but only available after the lift, because the arm is the scales.

Both checks for water come before the turn
Both checks for water come before the turn

Weighing means reading the payload from the joint torques, or from a wrist force sensor, with the glass lifted clear. The margin is large — 218 g against 563 g — so the threshold is not delicate, and it is set low enough that a glass with a mouthful left in it is still refused.

The turn has to be planned backwards#

Turning the glass over is 180 degrees about a horizontal line. The last joint of most arms has a limited range, commonly plus or minus 175 degrees, and 180 degrees of turn does not fit into that if you start in the middle of it.

Where in the wrist's range the turn starts decides whether it finishes
Where in the wrist's range the turn starts decides whether it finishes

The fix is to turn the wrist back before closing the fingers. That means the grasp has to already know the release orientation, which is the thing most people get wrong the first time: they write a grasp, then a turn, and find the turn impossible.

The tolerance is not where you would guess#

Getting a 64 mm glass over a 12 mm peg allows 26 mm of sideways error. The real constraint is the rack filling up.

The peg is easy to hit; the glasses already on the rack are not
The peg is easy to hit; the glasses already on the rack are not

Pegs 90 mm apart and glasses 70 mm across leave 10 mm each side. Because the glass is tall, orientation spends that budget faster than position does: a 5-degree tilt swings the far end 10.5 mm sideways. So the descent must be vertical, and holding the glass upright matters more than placing it exactly.

Squeezing it wrong breaks it, once#

There is a band of grip force that works. Below it the glass slips; above it the rim cracks. The upper edge belongs to the glass and cannot be moved. The lower edge can, by using fingers with more friction, and that is the only lever there is.

The grip force window, and what widens it
The grip force window, and what widens it

Wet glass, which is what comes out of a sink, pushes the lower edge up and narrows the band further.


3. The four versions#

Build it four times, the way the learning path does. Each version removes one assumption from the one before, and each is a system you can run and measure on its own. Do not start the next version until the current one fails for a reason you can state out loud.

Version 1: everything known in advance#

Assumes: the glass always starts on a marked spot, every glass is the same, a person puts out only empty ones, and the rack was measured once and has not moved.

What it is: a fixed sequence, with two of its steps controlled by force rather than by position.

for peg in pegs:                          # measured once, in fill order
    move_to(pickup, wrist = -90°)         # pre-turned, so the flip will fit
    close_gripper(force = grip_force)
    if gripper_width > 68 mm: stop("no glass")
    lift(); turn_wrist(to = +90°)         # the glass is now mouth down
    move_above(peg, clearance = 130 mm)
    descend_until(vertical_force > 2 N)   # the rim has met the base
    open_slowly(); retreat(); check_standing()

Stack: ROS 2, the Robot Operating System, with URDF, the Unified Robot Description Format, to describe the arm; MoveIt 2 for the moves; ros2_control with an admittance controller for the descent.

Buys you: a complete working loop in days, a cycle time you can measure, and a rig to test everything that follows. Cannot do: anything if the glass moves, the rack moves, or somebody puts out a full one.

Version 2: the glass can be anywhere, and it might be full#

Adds perception. An overhead camera finds every glass and classifies each one as empty or full; a wrist camera checks the last stretch of the approach; the lift confirms the weight before the turn.

There is more than one way to find the glass, and which one is right depends on how fixed your set of objects is. Read the table as a shortlist, with the recommendation underneath it.

OptionWhat it gives youWhen to pick it
YOLO segmentation — You Only Look Once, a fast detector that also outlines what it findsoutlines of a fixed set of classes, plus the empty-or-full label in the same passthe default, once you have labelled pictures of your own glasses
SAM 2 — the Segment Anything Modeloutlines of objects it was never trained onprototyping, and labelling the pictures YOLO will be trained on
Grounding DINO with SAM 2finds things from a written phrase, such as "drinking glass"when the set of objects is open, or before you have any labels
ClearGrasp, TransCGfills in the hole a transparent object leaves in the depth picturewhen you need the 3D shape and not just the outline

Start with SAM 2, use it to label a few hundred pictures, then ship YOLO. SAM 2 is class-agnostic and heavy; YOLO is fast and gives you the empty-or-full class in the same pass, which is the class you actually need. What that costs you is a labelling job on your own table, and a model that knows nothing outside the classes you labelled.

Stack, on top of version 1: YOLO segmentation with two classes, empty glass and full glass; Open3D to fit a cylinder to each outline against the measured table plane; an AprilTag fiducial marker — a flat printed pattern a camera can locate exactly — on the rack base; BehaviorTree.CPP for the sequencing, because this is the version where recovery branches start to multiply.

Then it has to touch the glass. Perception only says where the glass is. Four of the nine steps in section 1 — grasping, confirming it is empty, coming down onto the peg, and letting go — are settled by force rather than by geometry, and the last step is settled by looking again. That is the difference between a sequence that runs and a sequence that works.

Where every force decision gets its number from
Where every force decision gets its number from

Read the table as one row per moment: what it decides, and what gives you the number.

MomentWhat it decidesWhat provides it
Closing on the glasshow hard to squeeze, so it neither slips nor cracksgripper_controllers in ros2_controllers, commanded as a force rather than as a width
While carrying itwhether it is slippingthe same controller's reported finger width, which keeps closing if the glass is sliding through
Just after the liftwhether it is emptyforce_torque_sensor_broadcaster, publishing the wrist wrench, with the gripper's own weight taken off
Coming down onto the pegwhen the rim has met the baseadmittance_controller, or cartesian_controllers, moving on the force instead of to a height
Before opening the fingerswhether the rack is now carrying the glassthe same wrench: if the load has not transferred, the glass is hung up, so lift away rather than let go
After retreatingwhether it is actually standingthe overhead camera and the same YOLO pass, asking whether there is a glass on that peg

Two things are worth taking from that table. The first is that none of it is a new framework: it is version 1's ros2_control with two more controllers loaded and one topic read, which is the practical reason the force work belongs here rather than in something written from scratch. The second is that only one of these costs money. You have to buy the force reading — a wrist force-torque sensor, or an arm that estimates it from joint currents well enough to trust — while the slip check is free, because every gripper already reports where its fingers are.

Buys you: the assumption that hurt most in version 1 is gone. The table can be loaded any way round, and full glasses are handled correctly. Costs you: training data and a labelling job, plus a component whose reasoning you cannot read.

Version 3: a recipe for each glass we know#

You know what your glasses look like, and version 2 tells you where they are. That is enough to write down what to do with each one, and writing it down beats learning it while the set is small and fixed. A glass library is a small file with one record per type, and a recipe is the version 1 sequence reading its numbers from that record instead of from constants.

Three types cover a normal kitchen. Read the table as one row per record, where every column is a number the recipe uses directly.

TypeRim, heightEmpty, fullWater gateMost tilt allowedWhere it goes
Tea glass55 mm, 90 mm128 g, 273 g146 g6.4°front row
Tumbler70 mm, 120 mm218 g, 563 g261 g4.8°either row
Tall glass75 mm, 160 mm305 g, 854 g373 g3.6°back row

Two of those columns are worked out rather than measured, and both are worth following. The tilt limit is the 10 mm neighbour budget from section 2 divided by the height, which is why the tall glass is the fussy one and the tea glass is forgiving. The water gate sits an eighth of the way up from empty to full, so that a mouthful of water is still caught.

Those gates are the clearest argument for keeping all of this per type rather than global.

Why the water gate needs one number per glass type
Why the water gate needs one number per glass type

Where to hold it

A record needs one more thing: where on the glass to take hold. It is not obliged to be the middle, and for these glasses the middle is the wrong answer. Four criteria argue about it, and three of them agree.

Four criteria up the height of a glass, and where they agree
Four criteria up the height of a glass, and where they agree

So the record gains four more fields — grasp height, grip width, approach and force — and the question is settled once, by a person with a ruler, rather than worked out at run time.

For a glass with no record, the grasp comes from the cylinder that perception already fitted: take pairs of opposing points with opposite surface normals, score them by friction and by how far they sit from the centre of mass, and keep the lowest one that clears the table. On a cylinder that is arithmetic rather than a search. The MoveIt 1 package for this, moveit_grasps, has no ROS 2 branch, so it is a short piece of your own code instead.

There are also grasp-proposal networks — Contact-GraspNet and GraspGen are the open ones — which take a point cloud and return ranked six-degree-of-freedom grasps. They are the right tool for a bin of objects you cannot list in advance, and they are worth knowing exist. They are not right here, for two reasons: a transparent glass gives no point cloud to feed them, and they rank grasps by whether the object will slip, having been trained on objects that cannot break. A grasp across the rim scores well.

However the grasp was chosen, check it rather than trust it. All three checks use sensors the task already has.

WhenWhat to look atWhat it tells you
As the fingers closeforce against finger widthforce rising at the width you predicted means the glass is where you thought it was
Just after the liftthe torque in the wrist wrenchan unexpected moment means you gripped higher, lower or further off the axis than intended
Before the turntilt 20 to 30 degrees slowlywhether the glass shifts in the fingers, while a shift is still recoverable

A failed check costs a regrasp. Not checking costs a glass.

Stack: no new frameworks at all. The library is a file, and the recipe is the version 1 sequence with its constants replaced by lookups. The perception from version 2 gains one job: say which of the three types it is looking at, which for these three is a question of size and is already answered by the cylinder fit.

Buys you: three glass types handled properly, with numbers you can read, change and argue about, and no training run anywhere. Costs you: it does not extend by itself. Every type is a record somebody wrote, and section 5 is about the fourth one.

Version 4: an instruction decides the goal#

If the job is only ever "put the glasses on the rack", a language model adds nothing and you should not fit one. There is a single goal, it never changes, and a sentence describing it is strictly worse than a constant.

It earns its place when the goal changes, and the natural way for that to happen here is sorting. Real instructions look like this:

  • "Put the tea glasses on the front row and the tall ones at the back."
  • "The wine glasses don't go on the rack — put them on the tray."
  • "Leave the tea glasses out, we're using them."
  • "The rack is full. Stack whatever is left on the tray."

Each of those changes which glasses are picked, in what order, and where they end up. None of them changes how a glass is held. That split is the whole design.

What the instruction decides, and what it never touches
What the instruction decides, and what it never touches

Stack: an open vision-language model such as Qwen3-VL reads the instruction and the overhead picture and writes a short programme that calls the version 3 recipes. That pattern — the model writes the plan, ordinary code runs it — is the one Code as Policies established, and the learned methods document covers the family properly.

Buys you: goals that change daily without a code change, and an operator who does not have to be a programmer. Costs you: a new kind of failure. The model will cheerfully produce a plan the cell cannot carry out — a peg that is taken, a glass type it has never seen, a tray that is not there — and nothing downstream will notice that the goal was wrong, only that some step failed. So a checker sits between the model and the recipes, in ordinary code, and refuses anything that does not match the library and the current state of the rack.


4. Which part uses what#

This is the whole shortlist in one place. Read it as one row per job: what we use, and the obvious alternative we are not using, with the reason.

JobWhat we useRather than
Simulate it firstGazeboMuJoCo — better contact, but simulated cameras and ROS in the loop matter more here
Describe the armURDFa bespoke model nothing else can read
Find the glassesYOLO segmentationcolour thresholding, which has nothing to work with on a transparent object
Label the training picturesSAM 2, or Grounding DINO with a phraseoutlining a few hundred glasses by hand
Tell empty from fullthe same model, two classes, then confirmed by weightvision alone, because the turn cannot be undone
Fill the depth hole, if neededClearGrasp, TransCGwaiting for the depth camera to get better
Fit the shape and the tableOpen3DPCL, the Point Cloud Library — capable, but heavier than this needs
Locate the rackan AprilTag marker, read by apriltag_rosFoundationPose, unless you cannot glue a marker on
Plan the motionMoveIt 2hand-written waypoints, which stop working the day the rack moves
Plan grasp, turn and place togetherMoveIt Task Constructorplanning each stage separately, then finding the grasp forbids the release
Drive the jointsros2_controlyour own control loop, where the hard part is the timing
Descend onto the rackthe admittance controller in ros2_controllers, or cartesian_controllerscommanding a height into a rigid base
Sequence and recoverBehaviorTree.CPPa state machine, which turns illegible once recovery branches multiply
Hold the per-glass numbersa file, one record per typeconstants spread through the code
Choose where to hold an unfamiliar glassantipodal sampling on the fitted cylinder, written yourselfContact-GraspNet, GraspGen — they need a point cloud a transparent glass does not give, and they rank by slip rather than by breakage
Turn an instruction into a goalQwen3-VL in the Code as Policies patterna menu of buttons, which cannot cover what people actually ask for
Learn a skill, if recipes run outLeRobotthe original ACT and diffusion policy repositories, which are quiet now
Record every attemptrosbag2, which ships with ROS 2log lines, which cannot show you the frame before the drop

The hardware choices are shorter, and only four of them matter.

PartWhat we useWhy
Grippertwo-finger parallel, with soft silicone padsthe pads raise friction, which widens the low side of the force band; suction fails on a curved, wet, inverted object
Where it holdsthe end that becomes the top after the turnthe fingers then stay clear of the pegs and the neighbouring rims during the descent
Camerasone overhead, one on the wristthe wrist camera removes the camera-to-arm calibration error, which is the error that actually sinks this task
Forcea wrist force sensor, or joint-torque estimatesit is needed twice: to weigh the glass, and to stop the descent on contact

Three of those carry a cost worth knowing before you commit. MoveIt Task Constructor is a noticeably steeper learning curve than plain MoveIt, and version 1 does not need it. BehaviorTree.CPP adds a second representation, written in XML, that a newcomer must learn before they can read your logic at all. And the admittance controller needs a force reading you trust plus an afternoon of stiffness tuning with nothing to show for it.


5. When an unfamiliar glass turns up#

Version 3 knows three glasses. This section is about the fourth, and it is the question that decides whether the system is a demonstration or a product.

Noticing it is the safety-relevant half#

The dangerous case is not a new glass. It is a new glass treated as an old one — a cylinder fitted to a wine glass returns an answer, and the answer is confident and wrong. Three cheap checks catch nearly all of it, and any one of them failing should stop the attempt:

  • the cylinder fit reports a poor fit against the outline;
  • the measured rim and height match no record in the library;
  • the weight after the lift matches neither the empty nor the full figure for the type the system thinks it is holding.

That third one is the quiet hero, because it runs after the grasp and before the turn, which is exactly where a wrong guess is still recoverable.

Then there are three routes, and you will use all of them#

Read the table as three responses to the same event, in the order you would reach for them.

RouteWhat happensWhat it costsWhen it is right
Refuse itthe glass is left on the table and flagged for a personnothingalways the default, and the only correct answer for a wine glass, which cannot go on a peg at all
Measure and addsomebody writes a record, or the robot runs a measuring routine and proposes oneminutesa new size of a shape you already handle
Learn itdemonstrations are collected and a policy is trained for shapes no recipe coversa few hundred demonstrations and a rig to collect themwhen unfamiliar shapes stop being rare

The measuring routine is worth building, because the robot can fill in most of a record by itself. Segmentation against the table plane gives the rim diameter and the height. Closing carefully at the grasp height gives the width. Lifting gives the empty mass, and the full mass follows from the internal volume. The one number it cannot measure is the force that cracks the rim, because measuring that destroys a glass — so a new record starts with a deliberately gentle grip force, and a person tightens it by hand after watching a few runs.

Route three is where LeRobot comes in, and the learned methods document covers what it involves. It is the right answer when the library stops keeping up — a café with forty kinds of glassware, rather than a kitchen with three.

Across all three routes the pattern is the same, and it is the real answer to the question: a new size is nearly free, a new shape never is. Changing size changes numbers. Changing shape changes which surfaces can be held and which way up the thing can stand, and no amount of training data turns that into the same problem.


6. What will bite you#

  • Broken glass is not a retry. It leaves shards, and an arm that will happily carry on moving through them. The response to a drop is to stop the cell and call a person. Decide that before writing the recovery logic, because it changes its shape.
  • Water is worse than breakage. It spreads, it reaches the electronics, and it is a slip hazard for the people nearby. That is why there are two gates before the turn rather than one.
  • A nearly-empty glass looks empty. A centimetre of water in the bottom is hard to see and easy to weigh, which is the whole argument for the second gate.
  • One water threshold will not do. A full tea glass weighs less than an empty tall glass, so the gate has to come from the library rather than from a constant.
  • Calibration drift eats your margin. You have 10 mm between rims. A 3 mm error between where the camera thinks the rack is and where it is has taken a third of it before the arm has moved.
  • Wet glass is a different object. It slips at forces a dry one holds at. If the glasses come from a sink, every grip number has to be measured wet.
  • The rack moves, and it fills. It is light plastic on a table. Fix it down or find it every cycle — and fill the pegs in an order that keeps the occupied ones away from the next one for as long as possible.
  • Simulation will not give you the grip numbers. Rigid fingers on a thin rigid shell is close to the worst case for a physics engine. Use the simulator for reaching, planning, the turn and the clearances; get the grip from a real glass.

7. What to measure#

Robot demonstrations are easy to make look good, so decide what you are counting before you start. These six numbers are enough.

NumberHow to count it
Success rateglasses standing ten seconds after release, over attempts
Full glasses correctly left aloneover full glasses presented — a miss here is a wet floor
Breakages per thousand attemptsthe number that decides whether this can ever ship
Unfamiliar glasses correctly refusedover unfamiliar glasses presented — the measure of section 5
Cycle timefrom reaching for the glass to the arm clear of the rack
Failures caught before releaseas a share of all failures — whether the checks are working

The last row is the one people leave out and the one to watch most closely. A system that notices it has the glass wrong and puts it back down is in a different class from one that finds out by hearing it break, even at the same success rate.

Log every attempt: the camera frames, the force trace, the gripper width, and the planned and actual poses. Without those, a failure two weeks from now is a story rather than a bug.


8. Where to go next#


9. Quesitons#

  1. How can we we train the model on where to pick the glass from? We can use some images and add some markets on where to pick it from, can that be used, if yes how?
  2. How can we train the gripping of the glass from the handle?
  3. Can we use reinforced learning here? if yes where?

First Project/Version#

  • we have table and there we have x number of different types of glasses. we also have the rack or glass stand where glasses are put down so that their open face looks down. we will have 6 slots in the rack for now.

  • we also know rack dimension like height of that stand and difference between two parallel glasses

  • we have 5 types of glasses and we need to place the upside down on rack. -- simple straight glass like rocks or vodka glass -- wine glass -- kind of glass with simple handle -- milkshake glass -- irish glass

  • as all type of glass are of different shapes so we have some values predefined like weight of empty glass, then their opening thickness, height etc.

  • what we need is simply pick up the glass and place it on the rack safely. and we need some type of confirmation that it has completed task one by one.

  • one last thing is we should not disturb the glasses which has some water in it. we need to pick only empty glasses.

  • total number and which glass were present, we are not sure.

  • we also dont know the location where rack is present on the table.

  • first we need initial camera images (table camera)

  • also here we need to capture the images of rack as well to find either we have enough empty slots or not. there we can use yolo i think would be better.

  • it will just tell us which slots are empty and which are already booked.

  • there we can use yolo segmentation it will do both things one is giving outline and other is classifiyng the type of glass.

  • the other appraoch could be using of SAM for outline and then use some other ML model for classifcation

  • after we get these details. we need to perform first action which is grupper going and picking up the glass

  • to find exact position where it need to pick. we can use learned policy for that glass only. because the policy knows all glass dimensions and its dynamics

  • then it would be easier for the arm to perform this thing on that type of glass.

  • after picking the glass from that location to upward there we check the weight logic as well by force sensors of the joints

  • after that upside thing is movelt part and placing also

  • but in placing part, after arm done with it, we need to confirm whether it acutally get successful or not by camera image to confirm if that slot has been booked.