LAWSON DONG / FIELD NOTES
Leave me a message
Back to Research

REPRESENTATION ANALYSIS / METHODS

The measuring tools

01 / REPRESENTATION SIMILARITY

Adjacent Linear CKA

Structure preserved between consecutive layers.

Centered Kernel Alignment · same samples · centered features

Subtracting the mean removes absolute position, so CKA compares the internal geometry of the representations.

X=Zℓ−1zˉℓ⊤,Y=Zℓ+1−1zˉℓ+1⊤X=Z_\ell-\mathbf1\bar z_\ell^\top,\qquad Y=Z_{\ell+1}-\mathbf1\bar z_{\ell+1}^\topX=Zℓ​−1zˉℓ⊤​,Y=Zℓ+1​−1zˉℓ+1⊤​
CKA⁡(X,Y)=∥X⊤Y∥F2∥X⊤X∥F ∥Y⊤Y∥F\operatorname{CKA}(X,Y)=\frac{\|X^\top Y\|_F^2}{\|X^\top X\|_F\,\|Y^\top Y\|_F}CKA(X,Y)=∥X⊤X∥F​∥Y⊤Y∥F​∥X⊤Y∥F2​​

Closer to 1

Similar structure

Closer to 0

Less similar structure

Sharp drop → inspect geometric restructuring. A drop alone does not establish a phase transition.

Notation & invarianceZℓ∈Rn×dℓ,zˉℓ=1n∑i=1nzi(ℓ),∥M∥F=∑i,jMij2Z_\ell\in\mathbb R^{n\times d_\ell},\quad \bar z_\ell=\frac1n\sum_{i=1}^n z_i^{(\ell)},\quad \|M\|_F=\sqrt{\sum_{i,j}M_{ij}^2}Zℓ​∈Rn×dℓ​,zˉℓ​=n1​i=1∑n​zi(ℓ)​,∥M∥F​=i,j∑​Mij2​​

Rows: samples in the same order. Defined when the denominator is nonzero.

CKA⁡(aXQ,bYR)=CKA⁡(X,Y),a,b≠0,Q⊤Q=R⊤R=I\operatorname{CKA}(aXQ,bYR)=\operatorname{CKA}(X,Y),\quad a,b\ne0,\quad Q^\top Q=R^\top R=ICKA(aXQ,bYR)=CKA(X,Y),a,b=0,Q⊤Q=R⊤R=I

Invariant to orthogonal rotations and uniform scaling.

Kornblith et al. (2019)

Experiment implementation

02 / CLASS SEPARATION

Raw-vector Fisher ratio

Fℓ=between-class separationwithin-class dispersionF_\ell=\frac{\text{between-class separation}}{\text{within-class dispersion}}Fℓ​=within-class dispersionbetween-class separation​

Why raw vectors?

Full-dimensional geometry · No projection before measurement

zi(ℓ)∈Rdℓ→measure directlyFℓz_i^{(\ell)}\in\mathbb R^{d_\ell}\quad\xrightarrow{\text{measure directly}}\quad F_\ellzi(ℓ)​∈Rdℓ​measure directly​Fℓ​

For example: dℓ=768d_\ell=768dℓ​=768 → measure in all 768 dimensions.

zi(ℓ)→PCA / UMAP / t-SNEpi∈R2 or 3(visualization only)z_i^{(\ell)}\xrightarrow{\text{PCA / UMAP / t-SNE}}p_i\in\mathbb R^{2\text{ or }3}\quad\text{(visualization only)}zi(ℓ)​PCA / UMAP / t-SNE​pi​∈R2 or 3(visualization only)

A projected plot can hide or distort high-dimensional geometry.

Direction + magnitude · Before unit normalization

z^i=zi∥zi∥2,zj=αzi (α>0) ⇒ z^j=z^i\hat z_i=\frac{z_i}{\|z_i\|_2},\qquad z_j=\alpha z_i\ (\alpha>0)\ \Rightarrow\ \hat z_j=\hat z_iz^i​=∥zi​∥2​zi​​,zj​=αzi​ (α>0) ⇒ z^j​=z^i​

Normalization removes length differences; raw Fisher retains them. Cosine metrics complement it by focusing on direction.

Higher F

Greater separation relative to spread

Lower F

Less separation relative to spread

Full definition

C / D: cat / dog samples · N: class size · μ: centroid

μC=1NC∑i∈Czi,μD=1ND∑i∈Dzi\mu_C=\frac1{N_C}\sum_{i\in C}z_i,\qquad \mu_D=\frac1{N_D}\sum_{i\in D}z_iμC​=NC​1​i∈C∑​zi​,μD​=ND​1​i∈D∑​zi​
F=∥μC−μD∥221NC∑i∈C∥zi−μC∥22+1ND∑i∈D∥zi−μD∥22F=\frac{\|\mu_C-\mu_D\|_2^2}{\frac1{N_C}\sum_{i\in C}\|z_i-\mu_C\|_2^2+\frac1{N_D}\sum_{i\in D}\|z_i-\mu_D\|_2^2}F=NC​1​∑i∈C​∥zi​−μC​∥22​+ND​1​∑i∈D​∥zi​−μD​∥22​∥μC​−μD​∥22​​
F(az)=F(z)(a≠0),F∈[0,∞)F(az)=F(z)\quad(a\ne0),\qquad F\in[0,\infty)F(az)=F(z)(a=0),F∈[0,∞)

Uniform scaling cancels; feature-specific scaling can change F. These statements assume a positive denominator; the code floors it at 10−2010^{-20}10−20.

Higher F does not by itself guarantee better classification accuracy.

Experiment implementation

03 / LOCAL LABEL MIXING

Local Label Entropy (LLE)

Neighborhood purity · Label mixing · Local class organization

Find each sample’s k nearest neighbors by cosine distance in the full representation space. Exclude the sample itself.

dcos⁡(zi,zj)=1−zi⊤zj∥zi∥2∥zj∥2,Nk(i)=kNN⁡j≠i(zi)d_{\cos}(z_i,z_j)=1-\frac{z_i^\top z_j}{\|z_i\|_2\|z_j\|_2},\qquad \mathcal N_k(i)=\operatorname{kNN}_{j\ne i}(z_i)dcos​(zi​,zj​)=1−∥zi​∥2​∥zj​∥2​zi⊤​zj​​,Nk​(i)=kNNj=i​(zi​)pi,c=1k∑j∈Nk(i)1[yj=c],c∈{cat,dog}p_{i,c}=\frac1k\sum_{j\in\mathcal N_k(i)}\mathbf1[y_j=c],\quad c\in\{\mathrm{cat},\mathrm{dog}\}pi,c​=k1​j∈Nk​(i)∑​1[yj​=c],c∈{cat,dog}
Hi(k)=−∑c∈{cat,dog}pi,clog⁡2pi,c,LLEℓ(k)=1n∑i=1nHi(k)H_i^{(k)}=-\sum_{c\in\{\mathrm{cat},\mathrm{dog}\}}p_{i,c}\log_2p_{i,c},\qquad \mathrm{LLE}_\ell(k)=\frac1n\sum_{i=1}^n H_i^{(k)}Hi(k)​=−c∈{cat,dog}∑​pi,c​log2​pi,c​,LLEℓ​(k)=n1​i=1∑n​Hi(k)​

0 bits

One label · pure neighborhood

(pi,cat,pi,dog)=(1,0) or (0,1)(p_{i,\mathrm{cat}},p_{i,\mathrm{dog}})=(1,0)\text{ or }(0,1)(pi,cat​,pi,dog​)=(1,0) or (0,1)

1 bit

Equal mix · maximum entropy

pi,cat=pi,dog=12p_{i,\mathrm{cat}}=p_{i,\mathrm{dog}}=\tfrac12pi,cat​=pi,dog​=21​

Lower LLE → less local label mixing. High entropy can also occur near an organized class boundary; low entropy alone does not guarantee correct classification.

Neighborhood scale & baseline

k: neighbors per sample · n: total samples · ℓ: layer · 0log⁡20:=00\log_2 0:=00log2​0:=0

k↑ ⇒ broader neighborhood,0≤Hi(k)≤1k\uparrow\ \Rightarrow\ \text{broader neighborhood},\qquad 0\le H_i^{(k)}\le1k↑ ⇒ broader neighborhood,0≤Hi(k)​≤1

Compare layers using the same samples and k. Shuffling labels keeps the geometry fixed and provides a reference for local organization.

Order(k)=1−LLE(k)Eshuffle[LLE(k)]\mathrm{Order}(k)=1-\frac{\mathrm{LLE}(k)}{\mathbb E_{\mathrm{shuffle}}[\mathrm{LLE}(k)]}Order(k)=1−Eshuffle​[LLE(k)]LLE(k)​

Defined for a positive shuffled mean; negative values indicate more mixing than the shuffled baseline.

Experiment implementation · DenseNet implementation

AI STOPPED
Astra

Ready for our next conversation.

Nemi

Ready for our next conversation.

0/50 · Local conversation