Showing posts with label neural net. Show all posts
Showing posts with label neural net. Show all posts

Thursday, February 12, 2009

NeuroEvolution of Augmenting Topologies

During my second quarter as a computer science grad I took a course on neural networks. As mentioned in a previous post, I had previously taken a course in pattern recognition. Since I already knew quite a bit about neural nets I decided to research the topic of NeuroEvolution of Augmenting Topologies (NEAT) instead of work on the scheduled class project (comparing the performance of feed-forward, recurrent, and radial basis function neural nets).

NEAT is a genetic algorithm for evolving artificial neural networks that protects innovation by dividing network types into species based on similar characteristics. During crossover new networks inherit nodes and node connections from their parent networks to form a new network structure. Node genes and connection genes are labeled with innovation numbers so that as those genes can be tracked as they are passed on to future generations. After crossover new networks undergo mutation where new nodes or new connections are added to the network and connections weights are adjusted by a random process.

Unlike traditional learning where connection weights are adjusted via a backpropagation technique, NEAT selects members from the spawned population, called genomes, for breeding. The evolved members are composed of nodes with sigmoidal activation functions that behave in the same way as traditional multi-layer perceptron networks.

Using jNEAT, a Java implementation of NEAT, I attempted to solve 2 different classification problems. The first problem was to classify 150 iris samples as 1 of 3 breeds based on 4 features. The second problem was to classify 178 wine samples as originating from 1 of 3 batches of wine based on 13 features. After conducting a number of different experiments using various configurations settings, I was able to generate a network with 96% correct classification for the iris test dataset and a network with 85% correct classification for the wine test dataset.

About a year later I went on to apply the concept of speciation to a genetic algorithm in my masters thesis.

Resources:

Wednesday, October 29, 2008

Handwritten Digit Recognizer

During my first quarter as a computer science grad I took an incredibly enlightening course in pattern recognition. Teams of students were assigned the task of implementing a handwritten digit recognizer. Like humans, computers need to determine the content of handwritten information before it can be used in a meaningful way. This is accomplished through a form of optical character recognition (OCR).

The postal service accepts packages and envelopes with handwritten addresses which must be read and interpreted in order to sort mail and send each item to its intended destination. It's both impractical and expensive to have humans sort large volumes of mail so automated computer systems are often used instead. The systems often consist of cameras which take pictures of the addresses and feed the images into program for processing.

My team decided to implement a convolutional neural network similar to Lecun's LeNet-5.
  • The first layer is the input layer and consists of one neuron per pixel in a 29x29 padded version of the sample image.
  • The second layer applies 6 feature maps to the input layer. Each feature map is a randomly distributed 5x5 convolutional kernel.
  • The third layer applies 50 feature maps to all 6 of the previous feature maps after sub-sampling. Again, each feature map is a randomly distributed 5x5 convolutional kernel. These 2 layers are referred to as a trainable feature extractor.
  • The fourth and fifth layers are referred to as a trainable feature classifier. These 2 layers are fully connected and compose a universal classifier.
A convolutional neural network exploits the spatial structure of digits and attempts to train weights to identify spatial differences between digits. We calculated classification error rates using standard backpropagation, which played a major role in training the system.

After running the network with 100 hidden nodes for 5 epochs over 60000 MNIST training samples and 10000 test samples, our network misclassified 851 samples and had a 91.4% success rate. Not bad.

While this project focused on recognizing handwritten digits, the concepts and algorithms covered can be easily extended to apply to all alphanumeric characters.

Resources: