1. Three observations that a vision system must explain1. 视觉系统需要解释的三个现象

When you leave a bright street and enter a dark room, objects become visible gradually. At dusk, a flower may still have a clear outline after its color becomes difficult to judge. While reading, the word you fixate is sharp, but neighboring words lose detail. These effects arise at different stages: photon capture, retinal adaptation, spatial sampling and neural pooling.从明亮的街道走入暗室,物体的轮廓会逐渐显现。黄昏时,一朵花的轮廓可能仍然清楚,颜色却已难以辨认。阅读时,注视的词很清晰,旁边的词则难以逐字看清。这些现象涉及不同环节:光子捕获、视网膜适应、空间采样与神经信号汇聚。

01 · Sensitivity01 · 灵敏度Rod pathways support dim-light detection. Pooling signals can improve sensitivity while sacrificing localization.视杆通路支持暗光探测。汇聚多个位置的信号可提高灵敏度,也会牺牲定位精度。
02 · Acuity02 · 清晰度The foveal center has densely packed cones and fine-grained connections. High spatial detail is concentrated near fixation.中央凹中心具有密集的视锥和精细的连接,高空间分辨率集中在注视点附近。
03 · Contrast03 · 对比Many retinal outputs depend on a stimulus relative to its surroundings and recent history, not on intensity alone.许多视网膜输出取决于刺激相对周围区域及近期输入的差异,而非仅由绝对强度决定。

The sensitivity–resolution trade-off is physical as well as biological. If independent measurements are pooled, signal counts add and random fluctuations partly average out; the output then loses the identity of the individual positions. Conversely, smaller sampling regions preserve location but collect fewer photons in the same exposure. Biological vision distributes these resources unevenly across the retina.灵敏度与分辨率之间的取舍同时具有物理和生物基础。独立测量相加时,信号计数累积,随机涨落得到部分平均,但输出不再保留各个位置的身份。反之,较小的采样区域保留了位置信息,同样曝光时间内收集的光子却更少。生物视觉将这些资源不均匀地分配在视网膜上。

Purves et al.: rod/cone distribution and foveal acuityPurves 等:视杆、视锥分布与中央凹视敏度

2. From an image on the retina to a cellular circuit2. 从视网膜上的像到细胞回路

The cornea provides most of the eye’s refractive power; changing the shape of the lens adjusts focus for different distances. The iris changes the pupil diameter and therefore the light admitted. An optical image forms on the retina, but the optic nerve does not transmit a rectangular array of pixel values: its axons belong to retinal ganglion cells, whose responses have already been transformed by local circuits.角膜提供眼球大部分屈光力,晶状体通过形状变化调节不同距离的焦点。虹膜改变瞳孔直径,从而调节入光量。光学图像落在视网膜上,但视神经传输的并不是一张矩形像素表:其中的轴突来自视网膜神经节细胞,其响应已经过局部回路的处理。

National Eye Institute: cornea, lens, retina and optic nerve美国国家眼科研究所:角膜、晶状体、视网膜与视神经

The eye, in section眼球的空间结构
Oblique eye cutaway showing the lens behind the iris, zonular fibers, layered eye wall and optic nerve / 眼球斜剖面:虹膜后方的晶状体、悬韧带、分层眼球壁与视神经
03Lens晶状体
The biconvex lens lies behind the iris. Zonular fibers attach near its equator; accommodation changes its curvature.双凸晶状体位于虹膜后方。悬韧带在赤道部附近附着,调节通过改变晶状体曲率完成。
01
From an optical image to a neural circuit从光学图像到神经回路
A · The eyeA · 眼球
1234hν≈ 24 mm
  1. 1Cornea: refraction角膜:折射
  2. 2Lens: accommodation晶状体:调焦
  3. 3Retina: transduction视网膜:换能
  4. 4Optic nerve: axons视神经:轴突
B · A peripheral retinal sectionB · 周边视网膜切面
hνRPEONLOPLINLIPLGCL12345
  1. 1Photoreceptors感光细胞
  2. 2Bipolar cells双极细胞
  3. 3Ganglion cells神经节细胞
  4. 4Horizontal cell水平细胞
  5. 5Amacrine cell无长突细胞
Tissue thickness and cell counts are schematic; rays indicate convergence, not a full optical trace. Light and neural signals travel in opposite directions through the layer stack. RPE supports the outer segments; ONL/INL/GCL contain nuclei or somata, OPL/IPL contain synaptic connections.组织厚度与细胞数量为示意;光线仅表示会聚,不是完整的光学追迹。光与神经信号在层间的主要传播方向相反。RPE 支持外节;ONL/INL/GCL 容纳细胞核或胞体,OPL/IPL 容纳突触连接。

In most of the retina, incoming light crosses inner neural layers before reaching the photoreceptor outer segments next to the retinal pigment epithelium (RPE). The principal signal route then runs in the opposite direction: photoreceptor → bipolar cell → ganglion cell. Horizontal cells connect across the outer retina; amacrine cells modify signaling in the inner retina. Their lateral interactions contribute to spatial contrast, temporal filtering and other computations.在视网膜的大部分区域,入射光先穿过内层神经组织,才到达色素上皮(RPE)附近的感光细胞外节。主要信号通路随后朝相反方向运行:感光细胞 → 双极细胞 → 神经节细胞。水平细胞在外层提供横向联系,无长突细胞在内层调节信号;这些横向作用参与空间对比、时间滤波等计算。

The central fovea is specialized: overlying layers are displaced sideways and cone sampling becomes dense. The optic disc is a separate location where ganglion-cell axons exit; its lack of photoreceptors creates the anatomical blind spot. RPE cells absorb stray light, recycle visual pigment and remove shed outer-segment material. Müller glia help maintain the ionic and metabolic environment. A functional retina requires both the signaling circuit and this support tissue.中央凹具有特殊结构:覆盖其上的内层向两侧移开,视锥采样更加密集。视盘则是另一处位置,神经节细胞轴突由此离开;由于缺少感光细胞,它对应生理盲点。RPE 吸收杂散光、参与视觉色素循环并清除脱落外节。Müller 胶质细胞协助维持离子与代谢环境。视网膜的正常工作同时依赖信号回路和支持组织。

Webvision: retinal layers, cell types and foveal specializationWebvision:视网膜分层、细胞类型与中央凹结构

Webvision: retinal pigment epithelium and photoreceptor maintenanceWebvision:色素上皮与感光细胞的维护

3. The physical budget: photons, blur and sampling3. 物理约束:光子、模糊与采样

A dim scene supplies fewer photons during a given integration interval. Under a Poisson counting model, the variance equals the mean: a mean count of 25 has a standard deviation of 5, while a mean of 100 has a standard deviation of 10. The absolute fluctuation grows, but its fraction of the signal falls. Quadrupling collected photons therefore doubles the shot-noise-limited signal-to-noise ratio.暗场景在固定积分时间内提供的光子更少。在泊松计数模型下,方差等于均值:平均计数 25 对应标准差 5,平均计数 100 对应标准差 10。绝对涨落增加了,但它在信号中所占的比例下降。因此,收集光子数增加到四倍,散粒噪声限制下的信噪比才增加到两倍。

For a monochromatic camera-sensor example, mean photoelectron count is set by quantum efficiency η, irradiance E at the pixel, sensitive area A, integration time T and photon energy hc/λ. The equation below makes the exposure trade-off explicit. Photoreceptors instead initiate biochemical responses to absorbed photons; their amplification and neural noise are additional stages, not camera electrons with different names.以单色照明下的相机像素为例,平均光电子数由量子效率 η、像素处辐照度 E、感光面积 A、积分时间 T 与光子能量 hc/λ 决定。下式直接给出曝光取舍。生物感光细胞则由被吸收的光子启动生化响应;其放大与神经噪声属于后续过程,不能直接等同于相机中的电子计数。

Increasing electronic gain scales the signal and the noise already present; it does not increase the number of captured photons. Extending T does collect more photons, but a moving image travels farther during the exposure. For an image velocity v measured in pixels per second, the approximate motion-blur length is vT pixels. Pooling adjacent samples is another way to gain sensitivity, at the cost of spatial detail.提高电子增益会同时放大已有的信号与噪声,并不增加实际捕获的光子。延长 T 可以收集更多光子,却使运动图像在曝光期间移动得更远。若像面速度 v 的单位是像素/秒,则运动模糊长度约为 vT 个像素。合并邻近样本也能提高灵敏度,代价是空间细节减少。

EMVA 1288: quantitative characterization of image sensorsEMVA 1288:图像传感器的定量表征

02
Two limits before neural computation神经计算之前的两道限制
A · Counting noiseA · 计数噪声
010020030040001020SNRμ

4× detected photons → 2× SNR探测光子数 ×4 → 信噪比 ×2

B · DiffractionB · 衍射
D = 2 mm1.15′D = 4 mm0.58′λ = 550 nm

2× aperture → ½ diffraction angle孔径 ×2 → 衍射角 ÷2

Calculated ideal limits, not measurements of a person or a camera. μ is the mean detected count; real systems also have aberrations, sampling limits and other noise sources.计算得到的理想极限,并非对人眼或相机的测量。μ 是平均探测数量;真实系统还受像差、采样与其他噪声限制。

Light is also a wave. Even a perfectly focused point forms a diffraction pattern through a finite circular aperture. At wavelength λ and aperture diameter D, the Rayleigh angle is approximately 1.22λ/D. For λ = 550 nm and D = 2 mm this is 3.36 × 10⁻⁴ rad, or about 1.15 arcminutes. Doubling D halves this ideal angle, but real eyes also have aberrations, scattering and nonuniform receptor sampling.光也具有波动性。即使完全合焦,点光源经过有限圆孔后仍形成衍射斑。波长为 λ、孔径直径为 D 时,瑞利角约为 1.22λ/D。取 λ = 550 nm、D = 2 mm,得到 3.36 × 10⁻⁴ rad,约 1.15 角分。将 D 加倍可使这一理想角度减半,但真实眼睛还存在像差、散射以及不均匀的感光细胞采样。

An engineering image model separates these stages: first convolve scene irradiance with a point-spread function, then integrate over each pixel area, then sample in space and time, and finally add the relevant sensor noise. More pixels cannot recover spatial frequencies already lost in the optical blur. Sampling too coarsely can alias surviving detail into false patterns.工程成像模型应分开描述这些步骤:先让场景辐照度与点扩散函数卷积,再对每个像素面积积分,随后在空间与时间上采样,最后加入相应的传感器噪声。增加像素不能恢复已被光学模糊滤掉的空间频率;采样过疏还会把残留的高频细节混叠成错误纹理。

OpenStax University Physics: circular apertures and the Rayleigh criterionOpenStax 大学物理:圆孔衍射与瑞利判据

4. Why light hyperpolarizes a photoreceptor4. 为什么光使感光细胞超极化

A vertebrate photoreceptor is electrically active in darkness. cGMP keeps cyclic-nucleotide-gated (CNG) channels in its outer-segment membrane open, allowing an inward cation current. Light absorption changes the visual pigment’s conformation, activates transducin and then PDE6, and accelerates cGMP hydrolysis. Fewer CNG channels remain open. The reduced inward current makes the membrane potential more negative: hyperpolarization.脊椎动物的感光细胞在黑暗中并非没有电活动。cGMP 使外节膜上的环核苷酸门控(CNG)通道保持开放,产生内向阳离子电流。光吸收引起视觉色素构象变化,激活转导蛋白及 PDE6,加快 cGMP 水解。保持开放的 CNG 通道减少,内向电流随之下降,膜电位变得更负,这就是超极化。

Purves et al.: the phototransduction cascadePurves 等:光转导的分子级联

03
Light reduces the photoreceptor’s inward current光使感光细胞的内向电流减少
A · Rod-cell compartmentsA · 视杆细胞区室
1234hν
  1. 1Discs: visual pigment膜盘:视觉色素
  2. 2Inner segment: metabolism内节:代谢支持
  3. 3Nucleus细胞核
  4. 4Ribbon synapse带状突触
B · A center light incrementB · 中心光照增强
Photon absorbed by visual pigment视觉色素吸收光子
Transducin → PDE6 → cGMP ↓转导蛋白 → PDE6 → cGMP ↓
CNG channels close → inward current ↓CNG 通道关闭 → 内向电流减少
Hyperpolarization → glutamate ↓超极化 → 谷氨酸释放减少
ON ↑Sign-inverting synapse反号突触
OFF ↓Sign-preserving synapse保号突触
The molecular cascade applies to vertebrate rods and cones; the ON/OFF branch shown is the direct cone–bipolar pathway. Rod output predominantly uses the rod bipolar–AII amacrine route under dim conditions. Arrows denote changes relative to the preceding state, not absolute firing rates.分子级联适用于脊椎动物的视杆与视锥;图中 ON/OFF 分支表示视锥—双极细胞的直接连接。暗光下视杆输出主要经视杆双极细胞与 AII 无长突细胞中继。箭头表示相对先前状态的变化,不是绝对放电率。

This voltage change reduces glutamate release from the photoreceptor terminal. ON cone bipolar cells use a sign-inverting mGluR6 pathway: less glutamate permits depolarization through its downstream mechanism. OFF cone bipolar cells use ionotropic glutamate receptors and preserve the sign: less glutamate reduces their excitation. One light increment can therefore drive ON and OFF channels in opposite directions. “Light means more activity everywhere” is the wrong circuit model.膜电位变化使感光细胞末端释放的谷氨酸减少。ON 型视锥双极细胞采用反号的 mGluR6 通路:谷氨酸减少后,其下游机制允许细胞去极化。OFF 型视锥双极细胞采用离子型谷氨酸受体,保留信号符号:谷氨酸减少使兴奋性输入降低。因此,同一次增亮会使 ON 与 OFF 通道朝不同方向变化;不能把回路理解为“光更强,所有细胞都更活跃”。

Webvision: ON/OFF bipolar signaling and rod pathwaysWebvision:ON/OFF 双极细胞信号与视杆通路

Photoreceptor and many bipolar responses are graded. Ganglion-cell output uses action potentials: information is carried by spike times, firing rates and the population pattern, not by proportionally scaling spike height. A leaky integrate-and-fire approximation treats the membrane as a capacitor with a leak conductance, driven by synaptic current. Crossing a threshold emits an event and resets the model voltage.感光细胞以及许多双极细胞的响应是渐变的。神经节细胞输出采用动作电位:信息由脉冲时刻、放电率与群体模式携带,而不是通过等比例增高脉冲幅度表达。泄漏积分发放近似将膜视为带漏电导的电容,由突触电流驱动;膜电位越过阈值后产生事件,并复位模型电位。

The membrane time constant τ = Cₘ/gL determines how quickly the model integrates and forgets its input. This is a useful circuit-level approximation; it omits the voltage-gated conductances that generate the biological spike waveform. Dark and light adaptation additionally involve pigment availability, calcium-dependent feedback and circuit changes. They cannot be reproduced by changing pupil size alone.膜时间常数 τ = Cₘ/gL 决定模型积累和遗忘输入的速度。这是电路层面的有效近似,但没有描述生成真实动作电位波形的电压门控电导。明暗适应还涉及视觉色素的可用性、钙依赖反馈以及回路变化,不能仅靠改变瞳孔大小来复现。

Gerstner et al.: membrane circuits and integrate-and-fire modelsGerstner 等:膜等效电路与积分发放模型

Webvision: recovery and adaptation in rod/cone phototransductionWebvision:视杆与视锥光转导中的恢复与适应

The LIF dynamics and threshold-crossing derivation develops the membrane model in more detail.LIF 动力学与越阈时刻推导进一步展开了这一膜电位模型。

5. Color is a spectral measurement, not a wavelength label5. 颜色是光谱测量,不是波长标签

The light reaching the eye combines the illuminant spectrum with the surface’s spectral reflectance and viewing geometry. Human S, M and L cones have overlapping sensitivities. A single cone response is ambiguous: fewer photons near its sensitivity peak and more photons away from the peak can produce the same response. Comparing different cone classes provides spectral information unavailable from any one class.到达眼睛的光由照明光谱、物体的光谱反射率以及观察几何共同决定。人的 S、M、L 三类视锥具有重叠的敏感性。单个视锥的响应存在歧义:在敏感峰附近接收到较少光子,与远离峰值处接收到较多光子,可能产生相同响应。比较不同类型的视锥,才能获得单一类型无法提供的光谱信息。

Here Φ(λ) is the incident photon spectrum integrated over the collection area and time; sᵢ includes the wavelength-dependent probability of contributing to that cone’s response. Mapping an entire spectrum to three responses discards information. Different spectra can therefore match in cone response under specified viewing conditions: metamerism. Rod-dominated dim-light vision lacks the same three-cone comparison, which helps explain the loss of reliable color discrimination at low light.这里 Φ(λ) 是在收集面积和时间上积分后的入射光子光谱,sᵢ 包含不同波长对该类视锥响应的贡献概率。把整条光谱映射为三个响应必然丢失信息,因此不同光谱在特定观察条件下可能产生相同的视锥响应,即同色异谱。以视杆为主的暗光视觉缺少同样的三类视锥比较,这有助于解释低照度下可靠辨色能力的下降。

04
Color is inferred from a population response颜色来自多类细胞的联合响应
A · Overlapping sensitivitiesA · 重叠的光谱敏感性
400500600700SMLnm
B · A many-to-one measurementB · 多对一的测量
∫
SML
A spectrum → three responses一条光谱 → 三个响应
Different spectra can give the same triplet.不同光谱可以产生相同的三元响应。
The curves illustrate overlap only; their peak positions and widths are not a calibration dataset. A cone counts absorbed photons without attaching a wavelength label to each photon.曲线仅说明光谱重叠,峰值与带宽不能作为标定数据。视锥对吸收的光子产生响应,不为每个光子附加波长标签。

OpenStax Behavioral Neuroscience: photoreceptors and color codingOpenStax 行为神经科学:感光细胞与颜色编码

A camera’s RGB filters are not the human S/M/L sensitivities. Color reproduction requires calibrated transformations and an assumed or estimated illuminant. White balance corrects channel gains for illumination; it does not reconstruct an arbitrary lost spectrum. For physically meaningful averaging, blur and exposure operations, work in a linear-light space before applying a display transfer function. Averaging gamma-encoded values generally gives a different result from averaging the underlying light.相机的 RGB 滤光片并不等于人的 S/M/L 敏感性。颜色再现需要标定变换,并依赖假定或估计的照明。白平衡校正照明引起的通道增益差异,但不能重建任意丢失的光谱。要对光进行有物理意义的平均、模糊和曝光运算,应先在线性光空间计算,再施加显示传递函数。直接平均经过伽马编码的数值,通常不等于平均它们所代表的光。

W3C CSS Color 4: linear-light color mixingW3C CSS Color 4:线性光空间中的颜色混合

6. From center–surround circuitry to an edge filter6. 从中心—周围回路到边缘滤波

For many ON-center ganglion cells, a bright spot in the center increases firing, while light in the surround suppresses that increase. A large uniformly lit area can excite both components, partly canceling their effects. The spatial arrangement makes a boundary between light and dark regions especially informative. Horizontal-cell and inner-retinal mechanisms both contribute; one circular filter cannot represent every ganglion-cell type or all brightness perception.对许多 ON 中心神经节细胞,中心亮斑使放电增加,周围受光则会抑制这种增加。大范围均匀照明同时激活两个分量,它们的效应部分抵消。这种空间组织使明暗区域之间的边界成为重要信息。水平细胞与内层视网膜机制均有贡献;一个圆形滤波器不能代表所有神经节细胞类型,也不能解释全部亮度知觉。

Webvision: receptive-field centers, surrounds and ganglion-cell responsesWebvision:感受野中心、周围与神经节细胞响应

A useful engineering approximation subtracts two normalized Gaussian averages. The narrow kernel samples a local center; the wider kernel estimates its neighborhood. For a one-dimensional signal, let W[k] = G₁[k] − G₃[k], where each sampled kernel is normalized after truncation. This guarantees a zero-sum filter even with a finite implementation.一种实用的工程近似是将两个归一化高斯平均相减。窄核采样局部中心,宽核估计邻域。对于一维信号,令 W[k] = G₁[k] − G₃[k],每个离散核都在截断后重新归一化。这样,有限长度实现仍能保证滤波核总和为零。

Apply W to z[x] = ln(I[x]/Iref), with strictly positive intensities and a fixed positive reference. A uniform exposure multiplier a then adds the constant ln(a), which the zero-sum filter removes. This exact property belongs to this log-domain model; it is not a claim that every retinal neuron is perfectly exposure invariant. Clipping, zeros, sensor noise and nonlinear adaptation break the idealization.令 z[x] = ln(I[x]/Iref),其中强度严格为正、参考值固定且为正,再施加 W。全局曝光乘数 a 会在对数域增加常数 ln(a),而零和滤波器会去掉这一常数。这是该对数域模型的严格性质,并不意味着所有视网膜神经元都完全不受曝光影响。饱和截断、零值、传感器噪声与非线性适应都会破坏这一理想化条件。

05
An edge becomes two signed responses一条边界被分解为正负响应
A · Center minus surroundA · 中心减周围
-9090.40k
G₁G₃W = G₁ − G₃
B · Filtered log intensityB · 对数强度的滤波结果
2080I(x)0OFFONq(x)x
Computed from the discrete kernels and replicated boundary samples used in the C++ program below. Each Gaussian is normalized after truncation to ±9 pixels. The dashed curve is W. OFF = max(−q, 0); ON = max(q, 0).按下方 C++ 程序的离散核与边界复制方式计算。两个高斯核截断到 ±9 像素后分别归一化,虚线表示 W。OFF = max(−q, 0),ON = max(q, 0)。

The program uses a 20-to-80 intensity step and replicated boundary samples. It checks two consequences of the model: a constant field gives zero response, and multiplying every intensity by three preserves q. ON and OFF below are nonnegative filter outputs, not measured firing rates.程序使用从 20 到 80 的强度阶跃,并以复制边界样本处理越界。它检查模型的两个结论:恒定场响应为零;全部强度乘以三后 q 不变。下方 ON 与 OFF 是非负滤波输出,不是实测放电率。

#include <algorithm>
#include <cmath>
#include <iomanip>
#include <iostream>
#include <vector>

std::vector<double> gaussian(double sigma, int radius) {
    std::vector<double> g(2 * radius + 1);
    double sum = 0;
    for (int k = -radius; k <= radius; ++k) {
        g[k + radius] = std::exp(-0.5 * k * k / (sigma * sigma));
        sum += g[k + radius];
    }
    for (double& value : g) value /= sum;
    return g;
}

std::vector<double> response(const std::vector<double>& image) {
    constexpr int radius = 9;
    const auto center = gaussian(1, radius);
    const auto surround = gaussian(3, radius);
    std::vector<double> out(image.size(), 0);
    for (int x = 0; x < static_cast<int>(image.size()); ++x) {
        for (int k = -radius; k <= radius; ++k) {
            const int j = std::clamp(x - k, 0,
                                    static_cast<int>(image.size()) - 1);
            out[x] += (center[k + radius] - surround[k + radius])
                      * std::log(image[j]); // Iref = 1; image[j] > 0
        }
    }
    return out;
}

int main() {
    std::vector<double> image(61, 20);
    std::fill(image.begin() + 30, image.end(), 80);
    const auto q = response(image);
    for (double& value : image) value *= 3;
    const auto brighter = response(image);
    const auto flat = response(std::vector<double>(61, 50));
    double exposure_error = 0, flat_error = 0;
    for (std::size_t i = 0; i < q.size(); ++i) {
        exposure_error = std::max(exposure_error,
                                 std::abs(q[i] - brighter[i]));
        flat_error = std::max(flat_error, std::abs(flat[i]));
    }
    std::cout << std::fixed << std::setprecision(6)
              << "dark-side OFF = " << std::max(-q[29], 0.0) << '\n'
              << "bright-side ON = " << std::max(q[30], 0.0) << '\n'
              << std::boolalpha
              << "flat field ~ 0: " << (flat_error < 1e-12) << '\n'
              << "exposure invariance: " << (exposure_error < 1e-12) << '\n';
    return flat_error < 1e-12 && exposure_error < 1e-12
           && q[29] < 0 && q[30] > 0 ? 0 : 1;
}

This direct implementation makes the operator explicit. For two-dimensional images, Gaussian separability reduces a K × K convolution to horizontal and vertical K-tap passes; subtract the two blurred images afterward. Boundary treatment and kernel normalization must remain consistent when comparing implementations. Benchmark throughput only after checking the numerical result.这一直接实现把运算写得清楚。对于二维图像,高斯核可分离性将 K × K 卷积降为水平、垂直两次 K 点滤波,随后再将两张模糊图相减。比较实现时,边界处理与核归一化必须一致。先验证数值结果,再做吞吐量基准测试。

7. From temporal contrast to event cameras7. 从时间对比到事件相机

Retinal output contains multiple temporal response types, including sustained and transient channels. Adaptation changes sensitivity according to recent input; eye movements also change the retinal image of an otherwise static scene. Engineering can borrow the principle of reporting change without reproducing the entire biological circuit. A dynamic vision sensor implements a particularly concrete version at each pixel.视网膜输出包含持续型与瞬态型等多种时间响应。适应使敏感性随近期输入调整;即使场景静止,眼动也会改变视网膜上的像。工程上可以借鉴“报告变化”的原则,而不必复制整套生物回路。动态视觉传感器在每个像素上实现了一个明确的版本。

An ideal event pixel compares log intensity L(t) with a stored reference. Once their difference reaches +C or −C, it emits an event with position, timestamp and polarity, then advances its reference by that signed threshold. Equal log thresholds correspond to equal ratios in opposite directions, not equal percentage increases and decreases. With C = ln(1.2), the positive step is +20%, while the negative step is about −16.7%.理想事件像素将对数强度 L(t) 与保存的参考值比较。差值达到 +C 或 −C 时,它发出带位置、时间戳和极性的事件,并将参考值沿相应方向更新一个阈值。对称的对数阈值对应相反方向的等比例变化,而不是相同百分比的增减。C = ln(1.2) 时,正变化为 +20%,负变化约为 −16.7%。

06
Transmit changes at the time they occur在变化发生时传输变化
A · One ideal event pixelA · 一个理想事件像素
0C2C0204060ln I − ln I₀ms
Log intensity对数强度Reference参考值
B · A threshold, not a frame clockB · 阈值触发,不是帧时钟
C = ln 1.2+20% from the reference → ON相对参考值 +20% → ON
exp(−C) = 1 / 1.2About −16.7% → OFF约 −16.7% → OFF
(x, y, t, p)Address · time · polarity位置 · 时刻 · 极性
Analytical example: I rises exponentially from 100 to 144 between 10–20 ms, remains constant until 40 ms, then falls to 100 by 50 ms. The ideal events occur at 15, 20, 45 and 50 ms; the reference moves by ±C after each event.解析示例:I 在 10–20 ms 从 100 指数上升至 144,保持到 40 ms,再于 50 ms 降至 100。理想事件发生在 15、20、45、50 ms,每次事件后参考值移动 ±C。

The following samples include every threshold-crossing time in the plotted trajectory. The output is 15 ms ON, 20 ms ON, 45 ms OFF and 50 ms OFF. With more coarsely sampled data, this loop can recover event counts per interval but not the original crossing times; a real event sensor detects changes asynchronously.下列样本包含图中轨迹的每个越阈时刻,输出依次为 15 ms ON、20 ms ON、45 ms OFF、50 ms OFF。若输入采样更稀疏,这个循环只能恢复每个采样间隔内的事件数量,无法恢复原始越阈时刻;真实事件传感器是异步检测变化的。

#include <array>
#include <cmath>
#include <iostream>

int main() {
    const double C = std::log(1.2);
    const std::array<int, 8> time_ms{0, 10, 15, 20, 40, 45, 50, 60};
    const std::array<double, 8> log_ratio{0, 0, C, 2*C, 2*C, C, 0, 0};
    double reference = log_ratio.front();
    int count = 0;
    for (std::size_t i = 1; i < time_ms.size(); ++i) {
        // The tolerance only handles floating-point roundoff.
        while (std::abs(log_ratio[i] - reference) >= C - 1e-12) {
            const int polarity = log_ratio[i] > reference ? 1 : -1;
            reference += polarity * C;
            std::cout << time_ms[i] << " ms "
                      << (polarity > 0 ? "ON" : "OFF") << '\n';
            ++count;
        }
    }
    return count == 4 ? 0 : 1;
}

A constant scene ideally produces no events only when its image is also stationary and sensor noise is absent. Motion or illumination flicker can generate a dense stream. Threshold mismatch, background activity and bandwidth limits matter in hardware. Event polarity resembles an ON/OFF distinction, but an event camera’s threshold circuit is not a retinal ganglion cell or the LIF membrane model above.只有像面也保持静止、且没有传感器噪声时,恒定场景才会在理想情况下不产生事件。运动或照明闪烁可以带来密集事件流。硬件中的阈值不一致、背景活动与带宽限制都不可忽略。事件极性具有 ON/OFF 的区分,但事件相机的阈值电路并不等于视网膜神经节细胞,也不等于前面的 LIF 膜模型。

Gallego et al., Event-based Vision: A Survey — sensor model and limitationsGallego 等,《Event-based Vision: A Survey》——传感器模型与限制

8. From retinal signals to an engineering system8. 从视网膜信号到工程系统

Retinal preprocessing is not the end of vision. Along the principal geniculostriate pathway, ganglion axons project through the optic tract to the lateral geniculate nucleus and then to primary visual cortex. Nasal retinal fibers cross at the optic chiasm, so the right visual hemifield is represented in the left cerebral hemisphere and vice versa. This organization groups visual space rather than assigning one hemisphere to each eye.视网膜预处理并不是视觉的终点。在主要的膝状体—纹状皮层通路中,神经节轴突经视束投向外侧膝状体,再到初级视觉皮层。鼻侧视网膜纤维在视交叉处交叉,因此右半视野主要对应左半球,反之亦然。这种组织按视觉空间分工,而不是把一只眼睛分配给一个半球。

Cortical processing combines orientation, binocular disparity, motion, context and feedback across multiple areas. A center–surround filter detects a local contrast pattern; it does not identify the object responsible. Likewise, an event stream is a measurement representation, not a recognition result. A robotic system still needs geometry, tracking or learned inference and a decision rule tied to its task.皮层处理在多个区域中结合方向、双眼视差、运动、上下文与反馈。中心—周围滤波器检测的是局部对比模式,并不会识别产生该模式的物体。同样,事件流是一种测量表示,而不是识别结果。机器人系统仍需几何估计、跟踪或学习推断,并需要与任务相匹配的决策规则。

UTHealth Neuroscience Online: visual fields, LGN and cortical processingUTHealth Neuroscience Online:视野、外侧膝状体与皮层处理

The biological mechanisms above suggest concrete design choices. Each choice must be evaluated against the information it preserves and the failure it introduces.上述生物机制可以转化为具体的设计选择。每种选择都应根据保留了哪些信息、又引入了哪些失效方式来评估。

Constraint约束Implementation实现What to measure测量什么
Photon-limited detection光子受限的探测Longer integration or spatial pooling before thresholding延长积分,或先空间汇聚再设阈值SNR and detection error, together with blur and localization error同时测量信噪比、探测错误、模糊与定位误差
Local contrast局部对比Normalized center–surround filtering at relevant scales在相关尺度上做归一化中心—周围滤波Flat-field response, exposure sensitivity, noise amplification and edge localization恒定场响应、曝光敏感性、噪声放大与边缘定位
Color reproduction颜色再现Linear-light processing, white balance and calibrated color transforms线性光处理、白平衡与标定颜色变换Color error across illuminants; clipping and metameric failures不同照明下的色差、饱和截断与同色异谱失配
Fast motion快速运动Timestamped event streams with motion-aware aggregation带时间戳的事件流,配合考虑运动的聚合End-to-end latency, tracking error and peak event throughput端到端时延、跟踪误差与峰值事件吞吐量
Limited processing budget有限的处理预算Fine sampling near the task target, coarser context elsewhere任务目标附近精细采样,其余区域保留较粗上下文Task error versus compute, plus missed peripheral targets任务误差与计算量,以及周边目标漏检