The video shows the word spotting results of camera captured images on a pages of the George Washington collection. The codebook is an agglomerative clustering which uses the Shannon Entropy as the split criteria to generate the codewords. The training images are generated from synthetic text (i.e. true type fonts).
The visual signatures in the example are generated with HOG32 descriptors (8 orientation bins with a 2x2 spatial grid), codewords are weighted using a LLC with 3 neighbors and descriptors are approximated with a Hierarchical k-Means for fast encoding.
No source code available.