A Brief Introduction to Dimensionality Reduction: PCA and t-SNE
In machine learning, having too many features can lead to several problems, such as:
- Overfitting
- Slower processing speeds
- Difficulty visualizing data with more than three features
This is why dimensionality reduction becomes necessary. In practice, when dealing with hundreds or thousands of features, manually picking them is clearly not a sensible approach. Below, we introduce two widely used dimensionality reduction techniques in machine learning.
PCA (Principal Component Analysis)
Before introducing PCA, let’s first define our objective:
Transform a sample from an n-dimensional feature space into a k-dimensional feature space, where k < n.
Here are the main steps of PCA:
- Standardize the data
- Construct the covariance matrix
- Use Singular Value Decomposition (SVD) to obtain the eigenvectors and eigenvalues
- Sort the eigenvalues in descending order and select the top eigenvalues and their corresponding eigenvectors
- Project (map) the original data onto the eigenvectors to obtain the new set of features
The most crucial component of PCA is Singular Value Decomposition. Therefore, let’s take a closer look at SVD in the next section.
Intuitive Understanding of Singular Value Decomposition
In matrix factorization, Singular Value Decomposition is a well-known method. In high school mathematics, the most common application of matrix factorization is solving systems of equations (such as LU decomposition). We can gain an intuitive understanding from the SVD formula:
![]()
Where is an matrix, and are orthogonal matrices, and is the singular value matrix. The singular values correspond to the eigenvalues of matrix . In PCA, these are also referred to as principal components, representing the importance of the preserved information. They are arranged in descending order along the diagonal, forming a diagonal matrix.
So what does correspond to here? Naturally, it corresponds to our features. However, it is important to note that here, we usually compute using the covariance matrix. Remember that the data must be standardized before performing SVD.

Covariance matrix
Because the covariance matrix is often denoted by Sigma, make sure not to confuse it with above. To reduce dimensions, we can multiply the first columns of by the corresponding singular values in to obtain the new features. From a geometric perspective:
![]()
Geometrically, this operation actually projects onto the first vectors of .

The black line represents the eigenvector, and its length corresponds to the eigenvalue.

The blue dots represent the original positions of the data, while the red dots represent their projected positions on the eigenvector. With this, we have successfully reduced 2D data down to 1D.
Naturally, we can also reduce from 3D to 2D:


Applications of PCA
During dimensionality reduction, we want to retain the most critical features and discard the less important ones.
For example, when recognizing a person, the most important identifying features might be the eyes, nose, mouth, etc., whereas features like skin tone or hair can be discarded. In fact, PCA is commonly used for dimensionality reduction in facial recognition (Eigenfaces).

This provides an intuitive overview of SVD. Due to length constraints, we cannot dive deeply into every detail. If you are interested in SVD, feel free to check out Wikipedia.
t-SNE
PCA is an intuitive and effective dimensionality reduction technique. However, as seen when converting from 3D to 2D, some clusters can end up completely jumbled together.
PCA is a linear dimensionality reduction method. If the relationships between features are nonlinear, using PCA may lead to underfitting.
t-SNE is another dimensionality reduction technique, but it uses a more sophisticated approach to model relationships between high-dimensional and low-dimensional spaces. t-SNE approximates high-dimensional data using the probability density function of a Gaussian distribution, while using a Student’s t-distribution to approximate the low-dimensional data. It then computes similarities using Kullback-Leibler (KL) divergence and optimizes via gradient descent (or stochastic gradient descent).
Probability Density Function of the Gaussian Distribution

Where is a random variable, is variance, and is the mean.
Thus, the original high-dimensional data can be expressed as:

And the low-dimensional data can be represented using the probability density function of a Student’s t-distribution (with 1 degree of freedom):

Where represents data points in high-dimensional space, and represents data points in low-dimensional space. and denote their respective probability distributions.
Why use a Student’s t-distribution to approximate low-dimensional data? Mainly because projecting down to lower dimensions inevitably causes significant information loss. A t-distribution prevents the projection from being overly affected by outliers.
When sample sizes are small, the t-distribution models the population distribution better and is less sensitive to outliers.
Probability density functions of the Student’s t-distribution and Gaussian distribution
Similarity Between Two Distributions
To calculate the similarity between two distributions, KL divergence (Kullback-Leibler Divergence) is commonly used, also known as Relative Entropy.

t-SNE uses perplexity (Perp) as a hyperparameter.

The original paper states that perplexity is typically chosen between 5 and 50.
Cost Function
Calculating Cost using KL divergence:

Computing the gradient yields:

Finally, gradient descent (or stochastic gradient descent) is used to find the minimum.
Practical Experiment: Testing on MNIST
The test dataset can be downloaded here. First, let’s look at the result when reducing to 2D using PCA.
PCA
PCA Dimensionality Reduction
As you can see, after reducing down to 2D, the data is jumbled together into a single mass with virtually indistinguishable clusters. This happens because PCA’s linear projection loses too much information along the way.
t-SNE
Next, let’s test with t-SNE:
Dimensionality Reduction with t-SNE
This is the result using t-SNE. Even after dimensionality reduction, the data remains cleanly clustered. The difference between the two (PCA vs. t-SNE) is strikingly evident across these two figures.
Summary
Subsequently, several algorithms were proposed to enhance the performance of t-SNE; for details, see Accelerating t-sne using tree-based algorithms. Most popular data analysis languages and libraries have implemented it, including scikit-learn, R, MATLAB, and others.
However, because t-SNE is a nonlinear dimensionality reduction method, its execution time is significantly longer than that of PCA.
- When there are too many features, using PCA might cause the reduced features to underfit; in such cases, consider using t-SNE instead.
- t-SNE requires noticeably more computational time.
- The paper also describes several optimization techniques (such as how to choose perplexity). Since I haven’t finished reading it all yet, I will update this post with more details in the future.
References
- Van der Maaten L, Hinton G. Visualizing data using t-SNE
- Accelerating t-sne using tree-based algorithms Accelerating t-SNE computation using various tree-based algorithms
- 線代啟示錄 - 奇異值分解
This article was also published on Medium
Related Posts
- When a Measure Becomes a Target: From the Window Tax to Pull Request Counts I once wrote a script to tally how many PRs I contributed in a quarter, how many reviews I left, and how many tickets I closed, hoping to use numbers to prove my output to my manager. My manager simply remarked that performance isn't just about output. Years later, I finally understood—when a measure becomes a target, it ceases to be a good measure. From the British window tax and the Hanoi rat bounty to evaluating developers by PR counts today, the underlying mechanism is exactly the same.
- Using Cloudflare Images for Image Storage and Transformation Putting an image on a webpage is the simplest task in frontend development. But doing it properly—including resizing, generating multiple formats, and withstanding heavy traffic—is actually an entire end-to-end solution. Eventually, I offloaded everything to Cloudflare Images, keeping only a single original image.
- Stop Using AWS Access Keys Access Keys are an easily overlooked security risk in AWS. By pairing OIDC with IAM Roles, GitHub Actions can securely operate AWS resources without storing any secrets.
- Database Primary Keys: AUTO_INCREMENT, UUID, and UUIDv7 Backend developers often face the choice of primary keys: should you use auto-increment or UUID? What about collisions? How does UUIDv7 compare to created_at + index in performance? Here are the design decisions and benchmark results from testing 20 million rows.