1. Biomechanical & Optical Foundations of Video-Oculography
Video-Oculography (VOG) is the non-invasive optical quantification of ocular position, angular gaze
displacement, and palpebral fissure dynamics. The human eyeball functions as an optical spheroid with an
anterior radius of curvature of approximately $7.8\,\text{mm}$ (the cornea) embedded in a posterior globe
with a radius of $12.0\,\text{mm}$ (the sclera).
In high-end clinical and psychophysical research, eye tracking traditionally employs Pupil Center Corneal
Reflection (P-CR) under controlled infrared illumination ($850\,\text{nm} - 940\,\text{nm}$). In P-CR
systems, near-infrared diodes create a virtual glint image on the anterior corneal convex mirror surface
(the first Purkinje image). By calculating the mathematical vector between the pupil centroid and the
first Purkinje glint, head translation movements can be decoupled geometrically from pure ocular rotation.
However, pervasive consumer eye tracking must operate under visible-light ambient illumination using
ubiquitous RGB webcams without dedicated infrared strobes. Under these unconstrained conditions, the
system must solve a complex biophysical optimization problem: separating dynamic iris pigment boundaries,
corneal specular reflections, eyelash occlusions, and scleral shadows purely from digital RGB pixel arrays
in real time.
2. Mathematical Formulations of Pupil Centroid Extraction
Following facial landmark isolation via deep neural architectures, this application extracts two
bilateral sub-arrays corresponding to the left and right ocular palpebral apertures. Within each
bounding matrix, pupil centroid localization is formulated through digital luminance thresholding and
first-order spatial moments.
ITU-R BT.601 Luminance Conversion
Because human cones exhibit non-uniform photopic spectral sensitivity ($V(\lambda)$, peaking at
$555\,\text{nm}$ in the green spectrum), raw RGB pixel vectors are transformed into monochromatic
luma values ($Y$) using standardized ITU-R BT.601 coefficients:
$$I_{\text{gray}}(x, y) = 0.299 \, R(x, y) + 0.587 \, G(x, y) + 0.114 \, B(x, y)$$
The melanin-rich pigment of the human pupil acts as an optical light trap, absorbing nearly all incident
visible light and yielding the lowest luminance values in the orbital region.
Thresholding and Spatial Centroid Moments
A calibrated darkness threshold $\theta_{\text{dark}}$ isolates candidate dark pixels comprising the
pupillary aperture:
$$S_{\text{dark}} = \{(x, y) \mid I_{\text{gray}}(x, y) < \theta_{\text{dark}}\}$$
The pupil centroid $\mathbf{c} = (\bar{x}, \bar{y})$ is computed as the first-order spatial center of
mass across the thresholded pixel cluster:
$$\bar{x} = \frac{1}{|S_{\text{dark}}|} \sum_{(x, y) \in S_{\text{dark}}} x, \quad \bar{y} = \frac{1}{|S_{\text{dark}}|} \sum_{(x, y) \in S_{\text{dark}}} y$$
When $|S_{\text{dark}}| = 0$, the algorithm registers an absence of pupil luminance, which corresponds
biophysically to complete eyelid closure (spontaneous blinking) or severe off-axis gaze occlusion.
Normalized Gaze Vector Coordinate Calculation & Adaptive Calibration
Once pupil centroids are isolated relative to the left and right eye bounding boxes of width $W_{\text{eye}}$
and height $H_{\text{eye}}$, raw horizontal and vertical displacement vectors are computed:
$$g_{x,\text{raw}} = \left( \frac{\bar{x} - x_{\text{eye\_origin}}}{W_{\text{eye}}} - 0.5 \right), \quad g_{y,\text{raw}} = \left( \frac{\bar{y} - y_{\text{eye\_origin}}}{H_{\text{eye}}} - 0.5 \right)$$
To eliminate individual orbital asymmetries, spectacles offsets, and camera mounting tilts, this system
applies an affine baseline calibration offset $(\delta_x, \delta_y)$ and an adaptive scalar gain $K_{\text{gain}}$:
$$g_x = \operatorname{clip}\left( K_{\text{gain}} \cdot (g_{x,\text{raw}} - \delta_x), -1.0, 1.0 \right)$$
$$g_y = \operatorname{clip}\left( K_{\text{gain}} \cdot (g_{y,\text{raw}} - \delta_y), -1.0, 1.0 \right)$$
This dynamic calibration expands raw micro-displacements into a full, responsive visual tracking range
spanning the entire display interface.
3. Convolutional Neural Networks for Real-Time Facial Alignment
Prior to pupil segmentation, the system must localize the face within the video frame with millisecond
precision. Traditional computer vision relied on Viola-Jones Haar-like cascade classifiers or active shape
models (ASM), which suffer severe accuracy collapse under non-frontal head poses, facial tilt, and dynamic
lighting.
This application utilizes Google's BlazeFace framework (Bazarevsky et al., 2019), a
lightweight Single Shot Multibox Detector (SSD) optimized specifically for mobile GPUs and browser-native
WebGL execution. BlazeFace replaces standard convolutional filters with depthwise separable
convolutions:
$$\frac{\text{Computational Cost}_{\text{Depthwise}}}{\text{Computational Cost}_{\text{Standard}}} = \frac{1}{N} + \frac{1}{D_K^2}$$
where $N$ is the number of output channels and $D_K$ is the spatial kernel dimension ($3 \times 3$). This
reduces multiply-accumulate operations by more than $85\%$, achieving inference latencies below
$10\,\text{ms}$ on standard desktop hardware.
BlazeFace directly outputs six 2D facial keypoints alongside the face bounding box: right eye center,
left eye center, nose tip, oral center, right ear canal, and left ear canal. These keypoints establish
an anatomical anchor frame that stabilizes eye bounding boxes against translational head movement.
4. Sources of Optical Noise, Refractive Artifacts & Calibration
While client-side VOG provides an accessible biometric tool, several physical limitations impact
accuracy:
- Spectacle Glare and Reflections: Corrective lenses introduce geometric distortion
and bright specular reflections that disrupt gray-level thresholding. The system provides separated
left and right threshold sliders ($\theta_{\text{dark\_L}}, \theta_{\text{dark\_R}}$) to allow
individual eye calibration under asymmetric lighting.
- Eyelash Shadowing & Ptosis: Downward gaze or partial eyelid droop (ptosis) casts
dermal shadows across the superior iris, pulling the computed centroid upward. Adjusting the Eye
Height Ratio restricts the vertical search window to minimize palpebral margin interference.
- The Vestibulo-Ocular Reflex (VOR): When a human rotates their head while fixating on
a stationary target, the vestibular labyrinth drives equal and opposite compensatory eye movements.
Without 3D head pose tracking, head rotations produce apparent gaze shifts that must be corrected by
relating eye centroids to the fixed inter-ocular baseline distance.
5. Bio-Acoustic Sonification & Clinical Applications
To provide immediate sensory biofeedback, this simulator integrates real-time Web Audio synthesis:
- Blink Detection Audio Chimes: Spontaneous blinks (typically occurring $12 - 18$ times
per minute with durations of $100 - 400\,\text{ms}$) trigger rapid frequency-modulated audio pings
(left blink at $520\,\text{Hz}$; right blink at $480\,\text{Hz}$). This audio biofeedback aids
diagnosticians in tracking cognitive workload, fatigue, and dry-eye syndrome.
- Gaze-Velocity Tracking Tones: Gaze transitions emit subtle acoustic sweeps, providing
auditory verification of fixation transitions for hands-free assistive communication dashboards.
How to Use This Interactive Laboratory
- Automated Sandbox Demonstration: By default, the simulator operates in an
automated mathematical sandbox mode. Observe the simulated wireframe face, moving pupil centroids,
and real-time gaze telemetry without needing a webcam.
- Connect Live Camera: Click "ACTIVATE LIVE WEBCAM" in the control
panel or on the canvas overlay. Grant browser permissions to begin tracking your own eyes locally.
- Calibrate Bounding Geometries: Use the Eye Bounding Width and Height
sliders to frame your eye sockets snugly within the clean glowing rectangular regions.
- Fine-Tune Darkness Thresholds: If your pupils are not detected, adjust the
Left/Right Darkness Threshold sliders until the high-contrast neon cyan rings and gold crosshairs
lock cleanly onto your pupils.
- Calibrate Center Gaze: Look directly at the camera lens and click "CALIBRATE
CENTER GAZE". This zeroes your primary line of sight to $(0.00, 0.00)$ and enables
smooth, responsive look-direction tracking with on-screen gaze vectors.
Related Interactive Laboratories on BioniChaos
Open Access License: This interactive educational module is released under
CC BY-NC 4.0 (Attribution-NonCommercial)
for non-commercial research, academic study, and clinical education.
Commercial & Enterprise Licensing: For white-labeling, proprietary LMS/course
embedding, hardware dashboard telemetry integration, or custom feature engineering, secure a
commercial license at
BioniCloud.com or contact
Dr. Yuri Beno.