www.neuralnetworkdesigner.net  ·  Email Classification Project

📧 Email Classification Spam / Malicious Detection

End-to-end deep learning pipeline for classifying emails as Ham (legitimate) or Spam / Malicious.

📊 Dataset kaggle.com/datasets/isuranga/spammalicious-detection 100,000 emails · 30 features

1. Dataset Overview

The Spammalicious Detection dataset contains 100,000 labeled emails collected in 2024, with 30 features spanning content, metadata, security, and behavioral signals. The target variable is binary: 0 = Ham / Legitimate, 1 = Spam / Malicious.

Feature Categories

Content raw_text, subject, body_plain, body_html, language
Metadata from_address, from_domain, to_addresses, date, message_id
Security spf_result, dkim_result, dmarc_result, x_spam_score, received_origin_ip
Behavioral has_attachments, attachment_types, num_urls, num_phone_numbers

2. Model Design

A sequential feedforward neural network was designed for text-based email classification. The model takes 29 encoded features as input and passes them through two dense hidden layers with ReLU activation, a dropout layer for regularization, and a softmax output layer.

Input shape = (29,)
Dense units=128 · relu
Dropout rate=0.3
Dense units=64 · relu
Output units=2 · softmax

Note: The design canvas also supports advanced layers such as LSTM, GRU, Bidirectional, BatchNorm, and various pooling layers for more complex architectures.

3. Training Process

Data Splitting

Training Parameters

Optimizer: SGD Learning Rate: 0.010000 Batch Size: 512 Validation Split: 0.20 Epochs: 10 Save Best Model Checkpoint Early Stopping Reduce LR on Plateau

4. Evaluation Results

Test Loss
0.0100
Test Accuracy
0.9987
Final Train Loss
0.0183
Final Val Accuracy
0.9988

Training Progress (per epoch)

Epoch Loss Val Loss Accuracy Val Accuracy
10.44590.23540.86380.9858
20.17340.09750.98490.9935
40.05050.03550.99480.9971
50.03080.04020.99080.9976
60.02510.01830.99080.9978
70.01830.01030.99760.9981
90.01830.01180.99760.9988
📉 Loss over epochs
📈 Accuracy over epochs

Classification Report

Class Precision Recall F1-score Spam 0.998 0.997 0.997 Ham 0.997 0.998 0.997
Train Loss: 0.0183 Val Loss: 0.0118 Train Accuracy: 0.9976 Val Accuracy: 0.9988

5. Summary

✅ Achieved 99.87% test accuracy with consistently low loss and high precision/recall across both classes. The model generalizes well, with validation metrics closely tracking training performance.

The pipeline is ready for deployment in real-world spam filtering, email security research, and multilingual spam pattern detection.