ThesisOct 2025 – Oct 2026
Automated Floorplan Extraction and Refinement for Accessible Indoor Navigation
Master’s thesis: a pipeline that turns smartphone photos of evacuation floor plans into clean, editable maps for an indoor navigation system for people with visual and motor disabilities. RT-DETR-L reaches 0.894 mAP@50 and a human-in-the-loop editor replaces manual tracing with quick corrections.

Navindoor is a Sapienza project building indoor navigation for people with visual and motor disabilities. It needs accurate digital maps of university buildings, and the only source for most of them is the emergency evacuation plan printed on the corridor wall. Photographed with a phone, those plans are skewed, unevenly lit and covered in text, arrows and symbols. Tracing one building by hand can take days.
The thesis builds a system that does the tracing automatically and leaves a person to check and correct the result. That check is not optional: someone following spoken directions because they cannot see the building has no way to notice that the map sent them through a door that does not exist.

Data
There is no labelled dataset of Italian evacuation plans, so I built one from CubiCasa5K:
- a custom SVG parser that recovers absolute pixel coordinates by accumulating the full transformation matrices;
- a mapping from CubiCasa5K’s eighty-odd vector classes down to the five that Navindoor needs;
- dual-domain rasterisation: each plan is rendered once clean and once with scan-like noise, which doubles the dataset from about 5,000 to about 10,000 images at no annotation cost.
Two detectors, one controlled comparison
I trained YOLOv11-m (convolutional) and RT-DETR-L (a real-time detection transformer) under matched conditions on a single RTX 3090 Ti, at 1024 × 1024, with augmentation that keeps the plans orthogonal (no rotation, shear or perspective) and threefold oversampling of rare classes.
| Model | mAP@50 | mAP@50-95 | Missed walls | Missed doors |
|---|---|---|---|---|
| YOLOv11-m | 0.870 | 0.719 | 14% | 23% |
| RT-DETR-L | 0.894 | 0.764 | 9% | 18% |
The headline gap at IoU 0.5 comes mostly from an auxiliary class that is not exported, and on staircases YOLO is actually ahead (0.821 against 0.740). What made RT-DETR the better choice is what matters once boxes become polygons: it localises more precisely, misses fewer walls and doors, its validation losses stay flat where YOLO’s start drifting after epoch 35, and its F1 holds near the peak across confidence thresholds from about 0.3 to 0.8. On real phone photos of Sapienza plans it produced far fewer false walls from text and arrows.
From boxes to a map
Detections go through a geometric pipeline: overlapping boxes are merged by unary union, rooms are segmented with a distance-transform-guided flood fill, polygons are simplified with class-specific Douglas–Peucker tolerances and edges are snapped to right angles. A separate OCR step with PaddleOCR (PP-OCRv4) reads room codes from masked crops.
The editor
The system runs as a FastAPI backend with asynchronous GPU inference and a SvelteKit 5 editor where an operator fixes the result: SVG polygon editing, geometric snapping, semantic labels, door–room linking, undo/redo and export. The goal was one direct gesture per correction.

Limits
The controlled comparison runs on rendered residential plans, not on the photographs the problem is about, and the reduction in operator effort was clear during development but never formally measured. Self-supervised pretraining on unlabelled floor plans is the most promising next step to cut the labelled data needed for new building types.
Supervised by Prof. Emanuele Panizzi.