# A Data-Centric Approach to Image Classification

### **THE PROBLEM**

The problem included classifying the shown scenes into six different categories : **buildings, forest, glacier, mountain, sea, and street**.

But the challenge was that **we could not change the ResNet-18 baseline model**. Instead, we had to use a **data-centric approach** to improve the accuracy of the model.

So in this article, I am going to explain how I used **3LC** and its interface to inspect the data, understand where the model was struggling, and improve the accuracy without changing the model itself.

### The Baseline

Firstly, we started by training the model using the already provided and labeled data.

The initial validation accuracy was **67.66%**.

Now, the task was not to change the model or use a different architecture. The main task was to improve the data in such a way that the existing model could make better predictions.

This is where 3LC became useful.

### Inspecting the Dataset Using 3LC

Using 3LC, I could inspect the dataset and also see different information about how the model was performing on the images.

The **2D maps and 3D maps** helped in visualizing the data.

We could also use different filters such as:

*   Confidence
    
*   Defined labels
    
*   Undefined labels
    
*   Predicted labels
    
*   Sample weights
    

Using combinations of these filters, we could find the samples that might actually be useful for improving the model.

The idea was not to randomly select and label more images.

Instead, the goal was to **strategically find the images that could provide useful information to the model**.

![](https://cdn.hashnode.com/uploads/covers/6a2848ecbb5f3b99aa15a90a/ac9ad518-fd31-4cab-a6dc-a076ab1d267c.png align="center")

### Selecting High-Confidence Undefined Samples

One of the first things we looked at was the **undefined samples where the model was highly confident in its prediction**.

These images did not have a confirmed label yet, but the model was already very confident about what it thought the image represented.

For example, if an image was undefined but the model predicted **mountain with 95% confidence**, that sample could be inspected and, if the prediction was correct, labeled as a mountain.

This allowed us to add useful samples to the training data instead of randomly labeling images.

## Using the 3D Embedding Map

Another useful feature was the **3D embedding map**.

Embeddings are basically numerical representations of the images created by the model. They help us understand how the model is representing different images and where similar images are located.

Using the 3D map, we could see different clusters of images.

The **lasso option** was also useful here because we could select a particular cluster and inspect the images inside it.

This helped us look more closely at areas where different classes were close to each other.

![](https://cdn.hashnode.com/uploads/covers/6a2848ecbb5f3b99aa15a90a/4787b531-f3c0-40b3-9fae-114e91f8dfe7.png align="center")

## Mountain and Glacier Overlap

One example that I found interesting was the overlap between **mountain and glacier** samples.

Some mountain and glacier images were located very close to each other in the embedding space that was because model was facing difficulty in differentiating between a snow-covered mountain and a glacier.

So instead of ignoring these overlapping areas, we could select those clusters using the lasso tool and inspect the images.

If we found samples that were incorrectly labeled or useful undefined samples, we could manually assign the correct labels.

By correctly labeling samples from these difficult regions, we could give the model better data to learn from.

### Our Data-Centric Strategy

The main strategy was therefore **not randomly labeling more data**.

We were strategically:

1.  Finding high-confidence undefined samples
    
2.  Inspecting their predictions
    
3.  Looking at the embedding maps
    
4.  Finding overlapping or interesting clusters
    
5.  Manually checking the images
    
6.  Correcting or assigning useful labels
    
7.  Retraining the model
    
8.  Checking the results again
    

After retraining, we could again inspect the model and find more areas that needed attention.

So the process became an iterative cycle:

**Inspect → Label → Retrain → Analyze → Repeat**

## Improving the Accuracy

After multiple rounds of making changes to the data, retraining the model, and analyzing the results, the validation accuracy gradually improved.

The final development run achieved **76.55% validation accuracy**, which was an improvement of **8.89 percentage points** over the initial 67.66% baseline.

And the important part was that the **ResNet-18 model itself was not changed**.

The improvement came from working on the data and using the information provided by the model to understand which samples could help improve the training process.

![](https://cdn.hashnode.com/uploads/covers/6a2848ecbb5f3b99aa15a90a/c3d4fbe4-81de-438c-9dfb-b094fbe6999a.png align="center")
