Mike! ADALINE
Mike, I’ve been working on a new algorithm — ADALINE (ADAptive LInear NEuron) — with LLM assistance. I’ve put this post together to demonstrate what I’ve learned so far. ADALINE seems like a good bridge between the algorithms we’ve looked at before and the more complex neural networks used today.
In this post I discuss:
- ADALINE, and what it is.
- The Iris dataset, and how we can teach a machine to tell flowers apart.
- Standardization and why it’s important for ADALINE to converge.
- Functions that can be composed to implement ADALINE.
- How we build the model and use it.
ADALINE Introduction
Think of ADALINE as a smart weight-calculating machine. Its job is to look at several pieces of information (like a flower’s petal length and width) and decide which category that object belongs to.
ADALINE performs a 3-step cycle:
- It assigns a “weight” (importance) to every input. It multiplies the inputs by these weights, adds them up, and produces a raw score.
- It compares its raw score to the actual answer. If the answer was “1” and the model guessed “0.4,” it calculates that gap (the error).
- It uses that error to nudge the weights slightly so that next time it sees the data, its guess is closer to the truth.
It repeats this thousands of times until the error is as small as possible.
ADALINE was revolutionary for:
- Noise cancellation. One of its most famous uses was in telephone lines, filtering out echoes and background static in real time.
- Predictive filtering. Predicting where a signal (like a radio wave) is headed based on where it has been.
- Simple pattern recognition. Sorting data into two distinct groups — Accept vs. Reject, Species A vs. Species B.
ADALINE also has some blind spots:
- It can only solve problems where a straight line can separate the data. If your data categories are tangled like a marble cake, ADALINE can’t curve its logic to separate them.
- It’s designed to choose between two options. It struggles once you ask it to choose between three or more categories at once.
- It requires standardized input. If one input uses huge numbers (thousands) and another uses tiny numbers (decimals), ADALINE fails to learn efficiently.
Iris Dataset
I use a standard machine learning dataset that has been around a long time, the Iris dataset. It contains 150 samples of iris flowers divided equally among three species: Setosa, Versicolor, and Virginica. For each flower, researchers measured four physical characteristics in centimeters: sepal length, sepal width, petal length, and petal width.
from sklearn import datasets import numpy as np import matplotlib.pyplot as plt # 1. Load the Iris dataset iris = datasets.load_iris() # 2. Extract features: Sepal Length and Petal Length for the first 100 samples # (This gives us 50 Setosa and 50 Versicolor samples.) X = iris.data[:100, [0, 2]] y = iris.target[:100] # 3. Convert target labels to 1 (Versicolor) and -1 (Setosa) y_adaline = np.where(y == 0, -1, 1) # 4. Standardize (crucial for ADALINE convergence) X_std = np.copy(X) X_std[:, 0] = (X[:, 0] - X[:, 0].mean()) / X[:, 0].std() X_std[:, 1] = (X[:, 1] - X[:, 1].mean()) / X[:, 1].std()
Standardization here is performed manually for each feature (column) so they’re on the same scale. For each feature — Sepal Length (X[:, 0]) and Petal Length (X[:, 1]) — the code calculates:
- The mean (
.mean()): the average value of that specific measurement. - The standard deviation (
.std()): a measure of how much the numbers vary or spread out from that average.
It then applies this formula to every data point in the column:
X_std = (Original Value - Mean) / Standard Deviation
Subtracting the mean centers the data. Dividing by the standard deviation squashes or stretches the data so most values fall between -1 and 1.
After these lines run the features are no longer measured in centimeters; they’re measured in units of standard deviation. A value of 1.5 simply means “this petal is 1.5 units larger than the average petal in this group,” which makes it easy for ADALINE to treat Sepal Length and Petal Length with equal importance.
print("--- Raw Data Statistics ---")
print(f"Mean of Sepal Length: {X[:, 0].mean():.2f}")
print(f"Std of Sepal Length: {X[:, 0].std():.2f}")
print("\n--- Standardized Data Statistics ---")
print(f"Mean (centered at 0): {X_std[:, 0].mean():.2f}")
print(f"Std (scaled to 1): {X_std[:, 0].std():.2f}")
print("\nFirst 5 standardized samples [Sepal, Petal]:")
print(X_std[:5])
Implementation Functions
Here are the functions I’ve been working with to implement ADALINE.
def train_adaline(X, y, eta=0.01, epochs=50):
"""Orchestrate the training process.
X: list of input vectors
y: list of target values
eta: learning rate
epochs: number of passes over the dataset
"""
n_features = len(X[0])
weights = [0.0] * n_features
bias = 0.0
cost_history = []
for i in range(epochs):
epoch_cost = 0
for xi, target in zip(X, y):
weights, bias, cost = update_weights(xi, target, weights, bias, eta)
epoch_cost += cost
avg_cost = epoch_cost / len(y)
cost_history.append(avg_cost)
return weights, bias, cost_history
def linear_activation(inputs, weights, bias):
"""Calculate the continuous net input (z).
The heart of the neuron. Multiplies each input by its corresponding
weight and adds a bias. In ADALINE the raw value (z) is used to
calculate the error, which allows for smoother learning.
"""
return sum(x * w for x, w in zip(inputs, weights)) + bias
def update_weights(inputs, target, weights, bias, learning_rate):
"""The ADALINE Delta Rule (gradient descent).
Subtract the output (raw signal) from the target, then move the
weights in the direction that reduces that error. Return error**2
(sum of squared errors) — minimizing this cost is the primary goal
of training.
"""
output = linear_activation(inputs, weights, bias)
error = (target - output)
new_weights = [w + learning_rate * error * x for x, w in zip(inputs, weights)]
new_bias = bias + learning_rate * error
cost = error ** 2
return new_weights, new_bias, cost
def step_function(z):
"""The final classifier: returns 1 if z >= 0, else -1.
While the model learns using continuous numbers, this quantizer
converts the signal into a class label.
"""
return 1 if z >= 0.0 else -1
def predict(inputs, weights, bias):
"""The complete pipeline: linear signal -> step function."""
z = linear_activation(inputs, weights, bias)
return step_function(z)
Toy example
# [Hours Studied, Sleep Hours] - Standardized
X = [[1.5, 0.8], [1.2, 1.1], [0.1, -0.5], [-1.1, -1.0], [-0.5, 0.2]]
y = [1, 1, -1, -1, -1]
w, b, history = train_adaline(X, y, eta=0.01, epochs=20)
print(f"Final Weights: {w}, Bias: {b}")
print(f"Initial Error: {history[0]:.4f} -> Final Error: {history[-1]:.4f}")
Decision Boundary Plot (toy data)
import matplotlib.pyplot as plt
import numpy as np
X_train = np.array([[1.5, 0.8], [1.2, 1.1], [0.1, -0.5],
[-1.1, -1.0], [-0.5, 0.2]])
y_train = np.array([1, 1, -1, -1, -1])
w, b, history = train_adaline(X_train, y_train, eta=0.01, epochs=50)
plt.figure(figsize=(8, 6))
plt.scatter(X_train[y_train == 1, 0], X_train[y_train == 1, 1],
color='green', marker='o', label='Pass (1)')
plt.scatter(X_train[y_train == -1, 0], X_train[y_train == -1, 1],
color='red', marker='x', label='Fail (-1)')
# Decision boundary: w[0]*x1 + w[1]*x2 + b = 0
x1_min, x1_max = X_train[:, 0].min() - 1, X_train[:, 0].max() + 1
x1_values = np.array([x1_min, x1_max])
x2_values = -(w[0] * x1_values + b) / w[1]
plt.plot(x1_values, x2_values, color='black', linestyle='--',
label='Decision Boundary')
plt.xlabel('Hours Studied (Standardized)')
plt.ylabel('Sleep Hours (Standardized)')
plt.title('ADALINE: Passing vs. Failing Students')
plt.legend()
plt.grid(True, alpha=0.3)
plt.show()
Iris decision boundary
import matplotlib.pyplot as plt
import numpy as np
w, b, history = train_adaline(X_std, y_adaline, eta=0.01, epochs=20)
plt.figure(figsize=(8, 6))
plt.scatter(X_std[y_adaline == -1, 0], X_std[y_adaline == -1, 1],
color='red', marker='o', label='Setosa (Target -1)')
plt.scatter(X_std[y_adaline == 1, 0], X_std[y_adaline == 1, 1],
color='blue', marker='x', label='Versicolor (Target 1)')
x1_min, x1_max = X_std[:, 0].min() - 0.5, X_std[:, 0].max() + 0.5
x1_values = np.array([x1_min, x1_max])
x2_values = -(w[0] * x1_values + b) / w[1]
plt.plot(x1_values, x2_values, color='black', linestyle='--',
linewidth=2, label='ADALINE Boundary')
plt.xlabel('Sepal Length [Standardized]')
plt.ylabel('Petal Length [Standardized]')
plt.title('ADALINE: Separating Iris Setosa from Versicolor')
plt.legend(loc='upper left')
plt.grid(True, alpha=0.3)
plt.show()
Training run and predictions
final_w, final_b, cost_history = train_adaline(X_std, y_adaline,
eta=0.01, epochs=20)
print("Model training complete.")
print(f"Final Weights: {final_w}")
print(f"Final Bias: {final_b:.4f}")
test_samples = [X_std[0], X_std[75]] # Sample 0 (Setosa) and Sample 75 (Versicolor)
actual_labels = [y_adaline[0], y_adaline[75]]
print("\n--- Testing our composable functions ---")
for i, x_test in enumerate(test_samples):
prediction = predict(x_test, final_w, final_b)
label_map = {1: "Versicolor", -1: "Setosa"}
print(f"Flower {i+1} Features: {x_test}")
print(f"Predicted: {label_map[prediction]} | Actual: {label_map[actual_labels[i]]}")
print("-" * 30)