SafeWorld

Research

How to find false negatives in robot human detection

Test robot human detection under occlusion and changing light. Define false negatives, compare model versions, and document simulation limits.

A bounding box looks reassuring. The person who never received one deserves a closer look.

In robot human detection, a false negative occurs when an eligible person receives no matching detection. For a robot perception team, missed detections create a difficult evaluation problem. Reviewing detected objects alone cannot reveal every person the system missed. You need a reference that identifies people independently of the detector's output.

Start with a clear evaluation interval, trusted labels, and controlled scene variations. Measure misses by condition. Then examine when they occur, how long they last, and what the result can establish.

Here is a practical way to organize that work.

What counts as a false negative?

“Detect people reliably” leaves several questions open. Specify the operating region, sensor view, person class, and evaluation interval first.

For image-based detection, document the rule that matches a prediction to a labeled person. Also record the confidence threshold and treatment of small or partly visible people. Use consistent rules when comparing model versions.

A person behind an opaque wall presents a different problem from a visible shoulder behind a cart. Keep fully hidden periods separate from periods where your requirement expects detection. Do not silently discard difficult examples after seeing the results.

If the system must account for unseen people, evaluate that behavior through a separate system requirement. A camera detector cannot establish that empty-looking space contains nobody.

Visible, missed, and fully hidden.

Visible, missed, and fully hidden.

How do you label people the detector missed?

Labels must describe the scene without relying only on the detector under evaluation. Otherwise, an undetected person may disappear from both the output and the score.

For real recordings, review annotations, sensor timing, and ambiguous intervals. For synthetic recordings, inspect labels against the rendered sensor view. Check clipping, visibility, person identity, and timestamp alignment.

NVIDIA's Replicator documentation describes tools for sensor simulation, randomization, and annotated data generation. Those tools help create inputs. They do not establish that a test represents your deployment.

Choose the evaluation unit explicitly. A person-frame measures one labeled person in one frame. A sequence measures an encounter across time. They answer different questions, and repeated frames are not independent encounters.

How should you test occlusion and lighting?

Begin with a fixed scene and a known person trajectory. Hold the model, sensor settings, and scoring rules constant during a controlled comparison.

Then vary conditions that matter to the application:

Condition

Example comparison

What to keep visible in the record

Occlusion

Clear view versus partial blockage by a cart

Which body regions remain visible, and when.

Posture

Upright movement versus bending

Body position and the resulting image size.

Lighting

Even lighting versus backlighting

Exposure settings and whether the sensor model represents them.

Approach

Crossing versus moving toward the robot

Path, speed, distance, and evaluation interval.

Scene context

Sparse background versus nearby clutter

Geometry and objects that changed.

Appearance

Different clothing and body proportions

Asset provenance and the coverage limits of the set.

First isolate useful comparisons. Then test combinations that may interact. Partial blockage and poor contrast together may expose a weakness that either condition alone misses.

Occlusion has a long history in robot safety research. A National Institute of Standards and Technology study examined occlusion monitoring with human tracking and laser scanners. That work supports treating visibility as an explicit concern. It does not establish performance for a modern camera model.

Controlled scenario comparisons.

Controlled scenario comparisons.

Which human detection metrics should you report?

For a defined set of eligible person-frames, report missed detections divided by eligible labeled person-frames. State the matching rule, exclusions, and denominator beside the result. Also report false positives, especially when comparing thresholds.

Then examine time. A short gap and a sustained miss can produce similar averages while creating different engineering problems. Useful sequence measures include time to first valid detection and the longest continuous miss interval.

Two four-frame sequences compare long and brief detection gaps. Each photo aligns with a matching segment in the full-width gray status strip below.

Missed detections over time.

Preserve sequences with no valid detection. Do not remove them from a latency summary and then report only successful cases. Keep raw detections separate from any tracker output. A tracker may continue reporting a person during a detector gap.

Here is a deliberately invented reporting example. These numbers are not SafeWorld results or acceptance targets.

Suppose an eligible encounter fails when no valid detection occurs before its predefined deadline. Hold the person motion, model, and scoring rules fixed across these comparisons.

Condition

Failed encounters

Evaluated encounters

Observed fraction in this invented set

Clear view, even lighting

0

20

0%

Partial blockage, even lighting

3

20

15%

Partial blockage, backlighting

7

20

35%

The combined result is ten failures across 60 encounters, or about 16.7%. That average hides the condition worth investigating first. Zero failures in 20 encounters also does not establish a zero failure probability.

Report the counts before drawing conclusions. Sample size, repeated assets, correlated sequences, and scenario selection limit what these fractions mean. Select a statistical method that matches the sampling design before estimating uncertainty.

Did the model update introduce new misses?

Keep a fixed evaluation set outside model tuning. When a failure leads to a fix, add related held-out cases that test whether the improvement generalizes. Track the origin of assets and sequences to reduce leakage between training and evaluation.

Compare model versions on the same inputs first. If a threshold changes, report its effect on both misses and false positives. Review disagreements, invalid runs, and newly exposed failures.

Save enough information to repeat the comparison: inputs, labels, versions, thresholds, timestamps, and result definitions. More generated frames cannot compensate for a mislabeled or poorly scoped test.

What should stay fixed when cameras change?

Saved images cannot show what a different camera position would have seen. Keep the scene, human movement, and event definition available alongside the rendered recordings. Give each encounter a stable identifier.

When the camera changes, generate new sensor inputs and labels from those encounters. Check calibration, visibility, and eligible evaluation intervals again. Keep model changes separate from sensor changes where possible.

Compare results by encounter and condition, while recording each configuration. Retain held-out encounters that were not used to tune either configuration. This supports a reusable movement test set. It does not establish complete coverage of human behavior.

What can simulation tell you about robot perception?

A detector result describes one component. Stopping behavior also depends on control logic, communication, actuation, and the robot's physical state. Test those requirements separately or within a clearly defined integrated evaluation.

Synthetic results also need comparison with representative real sensor data. Differences in exposure, noise, geometry, and human appearance can change the result. Report those limits before using a simulation score to support a deployment decision.

The immediate goal is a useful engineering answer: which conditions deserve investigation, and did a change improve them?

Use the detection evaluation worksheet to record these decisions. For requirement traceability, see the robot risk assessment example.

SafeWorld helps teams explore human scenarios and controlled variations around robots. Bring the missed-detection problem, your evaluation rules, and the input format your team can use. We can discuss how to test it and what evidence would make the result useful.