BATH CLOSET WHAT THE DETECTOR SEES ?

In computer vision we're often limited to what pure vision models like YOLO or RF-DETR can read from the patterns in an image. Sometimes a region is obviously a bathroom because there's a toilet or a shower in it, or a kitchen because of the sink and the counters. But a master bedroom looks a lot like a guest bedroom. And very often the room comes with nothing but a label, a word written inside a rectangle. The model sees a rectangle.

That gap between “this is a rectangle with walls” and “this is the master bedroom, and that's its bathroom” is what we want to close. In this article we go through the main ways the industry tackles it, and where each one fails. None of them is enough on its own.

Two floor plans with every detected room shaded in a different color. A few rooms contain furniture or a label; most have no context at all.
Fig 01. Rooms detected across two plans. A few contain an object that gives them away, a few carry a label; most offer no context at all.

Approach 1 of 3: OCR and text extraction

This is the simplest case. When we have a PDF with searchable text, we can use common libraries to pull the text that falls inside the regions we selected. Combined with a good room detector, this already gives us a solid way to put names on the areas of a project.

The problem is that a lot of files don't have explicit text. Sometimes the drawing is rasterized, sometimes the text was exported as vectors, outlines that draw the letters without being a text object. In both cases there's nothing to extract, and this step simply doesn't work.

A floor plan detail where labels like Master Bedroom, Wardrobe and Stairs Down are drawn as vector strokes, each stroke painted a different color.
Fig 02. Every label on this sheet is line work. Each vector stroke is painted a different color; there's no text object to extract.

When the text layer is missing, we can still run actual OCR on the rendered image, with tools like Tesseract. In practice the accuracy is low. Architectural projects use all kinds of stylized fonts, and the mistakes those fonts cause are the dangerous kind. A 1 read as a 7 doesn't look like a failure. It looks like an answer, and it gets passed forward as context.

Closed OCR services, like Google Cloud Vision, handle these fonts much better. The trade is cost and latency: every region becomes a paid call and one more step in the pipeline. It works, but it's no longer the cheap and exact layer this step was supposed to be.

Approach 2 of 3: Object detection

We can also think about making this process more predictable by using other computer vision models, in particular object detectors. If we can identify furniture, sinks, toilets, appliances, we know more about the region they sit in and what it probably is, and we can link those objects to the room.

The catch is consistency when it's used by itself, and the upfront cost of building these models. Training takes data and time, and many of these objects have no standard way of being drawn in architectural plans. A bed or a table is often just a large rectangle, sometimes a circle, and what makes it a bed or a table is the context around it, not the shape itself. On its own, that hurts accuracy in exactly the cases where we need it. Combined with the other layers, it pays off, and that's how we use it.

Three crops of bedrooms from different plans, each drawing the bed differently: a bare rectangle with a headboard, a detailed bed with pillows, and a faint outline.
Fig 03. The same object, a bed, drawn three ways: a bare rectangle with a headboard, a detailed one with pillows, an outline that almost disappears. Out of context they're just shapes.

Approach 3 of 3: Sending masks to an LLM

When there's no text to extract and no object the detector can name, we can send the image to a language model that can read it. There are several ways to do that: the cropped mask, the mask drawn over the full image, the full image with the polygon highlighted, or just the crop. Which one works depends on the project and on how we deal with the image size limits of each model.

Large commercial plans are a real problem here. To fit a full sheet into the model's size limit we have to downscale it, and downscaling makes the small text much harder to read. When the reading gets hard, the model guesses, and a model that guesses is worse than a model that says nothing.

The same commercial floor plan crop shown twice: sharp as drawn on the left, and blurry after downscaling on the right, where labels and dimensions are hard to read.
Fig 04. The same crop as drawn (left) and after the downscale needed to fit the full sheet into a model input (right). The labels and dimensions get much harder to read.

Fonts are a problem too. Some of the fonts used in architectural drawings are hard for the model to read, and we've seen it confuse a 1 with a 7 more than once. On a room label that's annoying. On a dimension it changes the estimate.

Where the LLM pulls ahead

Resolution and fonts limit what the model can read. They don't limit what it can reason about, and that's where it does things the other two layers can't.

If we give it enough context, meaning a region larger than the polygon we're working on, it can infer things that aren't inside the polygon at all. It understands that Bathroom 1 belongs to Bedroom 1, and that Bathroom 2, which shares a wall with the same bedroom, has no direct relation to it. If we only send the polygon itself, we lose that. Geographically the two bathrooms are in the same place; only the surrounding layout tells them apart.

A residential floor plan where Bedroom 1 and Bedroom 3 each touch two rooms labeled only Bath, with plumbing walls highlighted in yellow.
Fig 05. Bedroom 1 touches two bathrooms, and the plan labels both of them just Bath. Bedroom 3 also touches two: one is the bedroom's, the other serves the hallway. None of that's written anywhere; only the surrounding layout tells.

We also stop depending on explicit labels. An object inside a region can say a lot about it: a desk means an office, a bed means a bedroom, a washer and dryer mean a laundry room. The model can make that kind of reading, and it fills in the rooms that have no label at all.

The same goes for questions that used to be hard for us to answer from geometry alone: is this area interior or exterior, is it covered or not. Those used to be real limits. They aren't anymore.

This kind of context also helps us understand jobs like remodels. When it's clear that the project touches only Bedroom 1 and only its bathroom, we can direct the takeoff to exactly that scope instead of quantifying the whole house.

Putting it together

In practice these aren't three alternatives. They're three layers.

Text extraction is cheap and exact, so it goes first whenever it works. Object detection comes next and gives us deterministic signals about what's in a room, when we have a model that can see it. The LLM sits on top of both, reading what the others couldn't and working out how the rooms relate to each other. Each layer covers the failure cases of the one before it.

Where the context goes

Classifying the region isn't the end of the pipeline. Every region we resolve becomes a fact in the project foundation, a persistent understanding of the project. When a later step needs to know which walls are plumbing walls, it doesn't reopen the drawing. It asks.

When the layers disagree, that's the system working. If the extracted text says closet and the detector sees a sink, we want to know before the estimate does, and an independent verification pass makes sure we do. Robustness comes from checking the readers against each other, not from stacking them.

Why this matters

An estimate is only as faithful as its reading of the project. The moment a region becomes Bathroom 2, attached to Bedroom 2, the walls around it become candidates for plumbing walls, usually framed with 2x6 studs instead of the regular 2x4. A porch changes how we size the flooring; a garage changes the slab and the fire separation. And so on through the whole takeoff.

This isn't theoretical for us. When we benchmarked general-purpose frontier agents on full residential takeoffs, they found most of the scope but couldn't turn it into defensible quantities. Coverage wasn't the bottleneck; grounding was.

The layers in this article have a major impact on the performance of Handoff-H1 and work as quality assurance for every takeoff it produces. You can read about the full system and its results in the H1 paper and in the release post on our blog.