Research
How to find false negatives in robot human detection
Test robot human detection under occlusion and changing light. Define false negatives, compare model versions, and document simulation limits.
Research
Test robot human detection under occlusion and changing light. Define false negatives, compare model versions, and document simulation limits.
A bounding box looks reassuring. The person who never received one deserves a closer look.
In robot human detection, a false negative occurs when an eligible person receives no matching detection. For a robot perception team, missed detections create a difficult evaluation problem. Reviewing detected objects alone cannot reveal every person the system missed. You need a reference that identifies people independently of the detector's output.
Start with a clear evaluation interval, trusted labels, and controlled scene variations. Measure misses by condition. Then examine when they occur, how long they last, and what the result can establish.
Here is a practical way to organize that work.
“Detect people reliably” leaves several questions open. Specify the operating region, sensor view, person class, and evaluation interval first.
For image-based detection, document the rule that matches a prediction to a labeled person. Also record the confidence threshold and treatment of small or partly visible people. Use consistent rules when comparing model versions.
A person behind an opaque wall presents a different problem from a visible shoulder behind a cart. Keep fully hidden periods separate from periods where your requirement expects detection. Do not silently discard difficult examples after seeing the results.
If the system must account for unseen people, evaluate that behavior through a separate system requirement. A camera detector cannot establish that empty-looking space contains nobody.

Visible, missed, and fully hidden.
Labels must describe the scene without relying only on the detector under evaluation. Otherwise, an undetected person may disappear from both the output and the score.
For real recordings, review annotations, sensor timing, and ambiguous intervals. For synthetic recordings, inspect labels against the rendered sensor view. Check clipping, visibility, person identity, and timestamp alignment.
NVIDIA's Replicator documentation describes tools for sensor simulation, randomization, and annotated data generation. Those tools help create inputs. They do not establish that a test represents your deployment.
Choose the evaluation unit explicitly. A person-frame measures one labeled person in one frame. A sequence measures an encounter across time. They answer different questions, and repeated frames are not independent encounters.
Begin with a fixed scene and a known person trajectory. Hold the model, sensor settings, and scoring rules constant during a controlled comparison.
Then vary conditions that matter to the application:
Condition | Example comparison | What to keep visible in the record |
|---|---|---|
Occlusion | Clear view versus partial blockage by a cart | Which body regions remain visible, and when. |
Posture | Upright movement versus bending | Body position and the resulting image size. |
Lighting | Even lighting versus backlighting | Exposure settings and whether the sensor model represents them. |
Approach | Crossing versus moving toward the robot | Path, speed, distance, and evaluation interval. |
Scene context | Sparse background versus nearby clutter | Geometry and objects that changed. |
Appearance | Different clothing and body proportions | Asset provenance and the coverage limits of the set. |
First isolate useful comparisons. Then test combinations that may interact. Partial blockage and poor contrast together may expose a weakness that either condition alone misses.
Occlusion has a long history in robot safety research. A National Institute of Standards and Technology study examined occlusion monitoring with human tracking and laser scanners. That work supports treating visibility as an explicit concern. It does not establish performance for a modern camera model.

Controlled scenario comparisons.
For a defined set of eligible person-frames, report missed detections divided by eligible labeled person-frames. State the matching rule, exclusions, and denominator beside the result. Also report false positives, especially when comparing thresholds.
Then examine time. A short gap and a sustained miss can produce similar averages while creating different engineering problems. Useful sequence measures include time to first valid detection and the longest continuous miss interval.

Missed detections over time.
Preserve sequences with no valid detection. Do not remove them from a latency summary and then report only successful cases. Keep raw detections separate from any tracker output. A tracker may continue reporting a person during a detector gap.
Here is a deliberately invented reporting example. These numbers are not SafeWorld results or acceptance targets.
Suppose an eligible encounter fails when no valid detection occurs before its predefined deadline. Hold the person motion, model, and scoring rules fixed across these comparisons.
Condition | Failed encounters | Evaluated encounters | Observed fraction in this invented set |
|---|---|---|---|
Clear view, even lighting | 0 | 20 | 0% |
Partial blockage, even lighting | 3 | 20 | 15% |
Partial blockage, backlighting | 7 | 20 | 35% |
The combined result is ten failures across 60 encounters, or about 16.7%. That average hides the condition worth investigating first. Zero failures in 20 encounters also does not establish a zero failure probability.
Report the counts before drawing conclusions. Sample size, repeated assets, correlated sequences, and scenario selection limit what these fractions mean. Select a statistical method that matches the sampling design before estimating uncertainty.
Keep a fixed evaluation set outside model tuning. When a failure leads to a fix, add related held-out cases that test whether the improvement generalizes. Track the origin of assets and sequences to reduce leakage between training and evaluation.
Compare model versions on the same inputs first. If a threshold changes, report its effect on both misses and false positives. Review disagreements, invalid runs, and newly exposed failures.
Save enough information to repeat the comparison: inputs, labels, versions, thresholds, timestamps, and result definitions. More generated frames cannot compensate for a mislabeled or poorly scoped test.
Saved images cannot show what a different camera position would have seen. Keep the scene, human movement, and event definition available alongside the rendered recordings. Give each encounter a stable identifier.
When the camera changes, generate new sensor inputs and labels from those encounters. Check calibration, visibility, and eligible evaluation intervals again. Keep model changes separate from sensor changes where possible.
Compare results by encounter and condition, while recording each configuration. Retain held-out encounters that were not used to tune either configuration. This supports a reusable movement test set. It does not establish complete coverage of human behavior.
A detector result describes one component. Stopping behavior also depends on control logic, communication, actuation, and the robot's physical state. Test those requirements separately or within a clearly defined integrated evaluation.
Synthetic results also need comparison with representative real sensor data. Differences in exposure, noise, geometry, and human appearance can change the result. Report those limits before using a simulation score to support a deployment decision.
The immediate goal is a useful engineering answer: which conditions deserve investigation, and did a change improve them?
Use the detection evaluation worksheet to record these decisions. For requirement traceability, see the robot risk assessment example.
SafeWorld helps teams explore human scenarios and controlled variations around robots. Bring the missed-detection problem, your evaluation rules, and the input format your team can use. We can discuss how to test it and what evidence would make the result useful.