01 / REPRESENTATION SIMILARITY
Adjacent Linear CKA Structure preserved between consecutive layers.
Centered Kernel Alignment · same samples · centered features
Subtracting the mean removes absolute position, so CKA compares the internal geometry of the representations.
X = Z ℓ − 1 z ˉ ℓ ⊤ , Y = Z ℓ + 1 − 1 z ˉ ℓ + 1 ⊤ X=Z_\ell-\mathbf1\bar z_\ell^\top,\qquad Y=Z_{\ell+1}-\mathbf1\bar z_{\ell+1}^\top X = Z ℓ − 1 z ˉ ℓ ⊤ , Y = Z ℓ + 1 − 1 z ˉ ℓ + 1 ⊤ Sharp drop → inspect geometric restructuring. A drop alone does not establish a phase transition.
Notation & invariance Z ℓ ∈ R n × d ℓ , z ˉ ℓ = 1 n ∑ i = 1 n z i ( ℓ ) , ∥ M ∥ F = ∑ i , j M i j 2 Z_\ell\in\mathbb R^{n\times d_\ell},\quad \bar z_\ell=\frac1n\sum_{i=1}^n z_i^{(\ell)},\quad \|M\|_F=\sqrt{\sum_{i,j}M_{ij}^2} Z ℓ ∈ R n × d ℓ , z ˉ ℓ = n 1 i = 1 ∑ n z i ( ℓ ) , ∥ M ∥ F = i , j ∑ M ij 2 Rows: samples in the same order. Defined when the denominator is nonzero.
CKA ( a X Q , b Y R ) = CKA ( X , Y ) , a , b ≠ 0 , Q ⊤ Q = R ⊤ R = I \operatorname{CKA}(aXQ,bYR)=\operatorname{CKA}(X,Y),\quad a,b\ne0,\quad Q^\top Q=R^\top R=I CKA ( a X Q , bY R ) = CKA ( X , Y ) , a , b = 0 , Q ⊤ Q = R ⊤ R = I Invariant to orthogonal rotations and uniform scaling.
Kornblith et al. (2019)
Experiment implementation
02 / CLASS SEPARATION
Raw-vector Fisher ratio F ℓ = between-class separation within-class dispersion F_\ell=\frac{\text{between-class separation}}{\text{within-class dispersion}} F ℓ = within-class dispersion between-class separation Why raw vectors? Full-dimensional geometry · No projection before measurement
For example: d ℓ = 768 d_\ell=768 d ℓ = 768 → measure in all 768 dimensions.
z i ( ℓ ) → PCA / UMAP / t-SNE p i ∈ R 2 or 3 (visualization only) z_i^{(\ell)}\xrightarrow{\text{PCA / UMAP / t-SNE}}p_i\in\mathbb R^{2\text{ or }3}\quad\text{(visualization only)} z i ( ℓ ) PCA / UMAP / t-SNE p i ∈ R 2 or 3 (visualization only) A projected plot can hide or distort high-dimensional geometry.
Direction + magnitude · Before unit normalization
z ^ i = z i ∥ z i ∥ 2 , z j = α z i ( α > 0 ) ⇒ z ^ j = z ^ i \hat z_i=\frac{z_i}{\|z_i\|_2},\qquad z_j=\alpha z_i\ (\alpha>0)\ \Rightarrow\ \hat z_j=\hat z_i z ^ i = ∥ z i ∥ 2 z i , z j = α z i ( α > 0 ) ⇒ z ^ j = z ^ i Normalization removes length differences; raw Fisher retains them. Cosine metrics complement it by focusing on direction.
Full definition C / D: cat / dog samples · N: class size · μ: centroid
μ C = 1 N C ∑ i ∈ C z i , μ D = 1 N D ∑ i ∈ D z i \mu_C=\frac1{N_C}\sum_{i\in C}z_i,\qquad \mu_D=\frac1{N_D}\sum_{i\in D}z_i μ C = N C 1 i ∈ C ∑ z i , μ D = N D 1 i ∈ D ∑ z i F ( a z ) = F ( z ) ( a ≠ 0 ) , F ∈ [ 0 , ∞ ) F(az)=F(z)\quad(a\ne0),\qquad F\in[0,\infty) F ( a z ) = F ( z ) ( a = 0 ) , F ∈ [ 0 , ∞ ) Uniform scaling cancels; feature-specific scaling can change F. These statements assume a positive denominator; the code floors it at 10 − 20 10^{-20} 1 0 − 20 .
Higher F does not by itself guarantee better classification accuracy.
Experiment implementation
03 / LOCAL LABEL MIXING
Local Label Entropy (LLE) Neighborhood purity · Label mixing · Local class organization
Find each sample’s k nearest neighbors by cosine distance in the full representation space. Exclude the sample itself.
d cos ( z i , z j ) = 1 − z i ⊤ z j ∥ z i ∥ 2 ∥ z j ∥ 2 , N k ( i ) = kNN j ≠ i ( z i ) d_{\cos}(z_i,z_j)=1-\frac{z_i^\top z_j}{\|z_i\|_2\|z_j\|_2},\qquad \mathcal N_k(i)=\operatorname{kNN}_{j\ne i}(z_i) d c o s ( z i , z j ) = 1 − ∥ z i ∥ 2 ∥ z j ∥ 2 z i ⊤ z j , N k ( i ) = kNN j = i ( z i ) p i , c = 1 k ∑ j ∈ N k ( i ) 1 [ y j = c ] , c ∈ { c a t , d o g } p_{i,c}=\frac1k\sum_{j\in\mathcal N_k(i)}\mathbf1[y_j=c],\quad c\in\{\mathrm{cat},\mathrm{dog}\} p i , c = k 1 j ∈ N k ( i ) ∑ 1 [ y j = c ] , c ∈ { cat , dog } Lower LLE → less local label mixing. High entropy can also occur near an organized class boundary; low entropy alone does not guarantee correct classification.
Neighborhood scale & baseline k: neighbors per sample · n: total samples · ℓ: layer · 0 log 2 0 : = 0 0\log_2 0:=0 0 log 2 0 := 0
k ↑ ⇒ broader neighborhood , 0 ≤ H i ( k ) ≤ 1 k\uparrow\ \Rightarrow\ \text{broader neighborhood},\qquad 0\le H_i^{(k)}\le1 k ↑ ⇒ broader neighborhood , 0 ≤ H i ( k ) ≤ 1 Compare layers using the same samples and k. Shuffling labels keeps the geometry fixed and provides a reference for local organization.
O r d e r ( k ) = 1 − L L E ( k ) E s h u f f l e [ L L E ( k ) ] \mathrm{Order}(k)=1-\frac{\mathrm{LLE}(k)}{\mathbb E_{\mathrm{shuffle}}[\mathrm{LLE}(k)]} Order ( k ) = 1 − E shuffle [ LLE ( k )] LLE ( k ) Defined for a positive shuffled mean; negative values indicate more mixing than the shuffled baseline.
Experiment implementation · DenseNet implementation