What is the best fixation detector for webcam eye tracking?
Every fixation count, duration and dwell time depends on the detector that produced it, and the classic detectors were built for laboratory trackers recording 500 to 2000 samples a second. Which one works best on webcam data, at about 30? We tested them against an EyeLink 1000 and on real online participants, and found that none of them does well on its own.
Outline
Abstract
Fixations are the unit most eye-tracking analyses are built on: fixation counts, durations, dwell times and areas of interest all inherit whatever the detection algorithm does. Detection methods were developed for laboratory trackers sampling at 500 to 2000 Hz; a webcam delivers about 30 samples per second, with far more noise per sample. So which fixation detector is the best choice for webcam data?
We compared Labvanced's dispersion-based detectors with the velocity-based algorithm of Engbert and Kliegl, using an exact JavaScript port of Titus von der Malsburg's saccades R package that we verified against the original (no mismatches across 13,790 output rows). The comparison ran on three datasets: 23 participants recorded with the webcam and an EyeLink 1000 at the same time, 94 participants of a fixation and pursuit task, and 225 Prolific participants of a 45-target accuracy task.
The algorithms fail in opposite directions. Dispersion detection never merges two real fixations but fragments one steady look into several pieces: a median of 3 where there is 1, in 82 % of steady looks. Engbert and Kliegl, whose threshold is set from each participant's own velocity noise, rarely fragments, but on webcam data it merges short fixations and small saccades, and misses large target saccades in the noisiest participants. Position accuracy hardly depends on the algorithm, because the remaining webcam error is a constant offset rather than noise.
The best detector for webcam data is therefore neither of them, but a combination: we built a new detector that keeps dispersion detection's sensitivity to real saccades and adds a live merge step, which joins two consecutive fixations when their centers are closer than the measurement noise can explain. Against EyeLink it raises the one-to-one F1 score from 0.58 to 0.75. On 195 Prolific participants it doubles the share of steady looks reported as one fixation (22.5 % to 50 %), lengthens the fixation covering a look from 847 to 1490 ms, and improves its accuracy slightly but significantly (7.40 to 7.29 % of the screen diagonal, p = .0003). The price is latency: a fixation is reported only once the next one has started, about one second later, which matters for gaze-contingent designs. Merged Fixations is the default detector in Eye Tracking 2.0; the real-time and lagged detectors remain available.
More coming soon
The analysis is complete. The full write-up, with the method, the three datasets, the results detector by detector, the figures and our recommendations for each study design, will follow in this note, together with its data.