Sunday, June 22, 2008

2008 Canadian Robot Vision Conference

This is my report on the Canadian Intelligent Systems Collaborative (AI/GI/CRV/IS) 2008 Conference. This was five conferences in one event. My interest was in the CRV (Computer and Robot Vision) conference held by the Canadian Image Processing and Pattern Recognition Society (CIPPRS).
The buildings in which conference was held were being reconstructed and the noise was distracting. But, with the inconvenience tolerated, the conference was very enlightening for a beginner like me.
The conference was essentially three days, with a keynote each morning and papers or talks throughout each day.
The only keynote I heard was the first, which was by Peter Carbone from Nortel. This was interesting to me (with my history in telecommunications), but off-target for the AI people who were the majority of conference attendees. Mr. Carbone made several predictions that I think will be significant in telecom: widespread broadband wireless penetration by 2010, telephony using SOA with mashup potential, and 100 MBs real-time encryption capability.
Following the keynote the CAIAC Precarn Intelligent Systems Challenge was announced. This offers a $10K prize to the student submitting the best method of detecting ships meeting at sea using satellite and radar tracking data. I think the poor data makes this a challenging problem.

The highlight of the conference for me was a talk by Dr. Steven Zucker from Yale. He was the only presenter who seemed interested in doing AI and computer vision to emulate biology, as I am. I think artificial systems should understand what they're working on; be a part of their world, as biological systems are. A biological goal Zucker identified is guiding animal movement, such as monkeys jumping to tree branches. This objective is the same as Arathorn's example of goats jumping to rocky ledges. Zucker's talk was mainly on stereo vision. He confirmed my assessment that Canny edge detectors suck. He implemented a nice curve detector based on tangents. He showed how he used spacial and orientation disparity to get a better matching of the image pair. A nice point was about self-referential calibration: a system that can move can identify its own parts (e.g. in a mirror) as the ones that move when it moves them.
I missed the talk by Dr. James Crowley, to my regret. I gathered that his points included that intelligence requires embodiment and autonomy. This confirms my subscription to the philosophy of Spinoza, who states that the mind is the entire body. Any organism's mental reality would not be what it is without all of the sensory input and motor feedback provided by the body.
The talk by Dr. Greg Dudek about his AQUA robot was interesting because of the focus and completeness of the project. It's another very specialized machine, although you can program its actions by a visual language. Apparently they discarded a visual system that recognized human hand gestures.

I attended all of the CRV paper presentations. These seemed to be arranged in ascending order of complexity and accomplishment. I was surprised that I could understand much of the work. Some of the papers were not amazing to me at all. Some were incremental improvements on previous work. Most were applications of existing work. This may be a survey of the state of the art, or it may just be a sampling of people who are trying to get attention (who didn't go to other conferences). I'm not going to summarize all of the papers - just give criticisms of the ones I found useful.
'An Efficient Region-Based Background Subtraction Technique' and 'Ray-based Color Image Segmentation' presented image segmentation optimizations based in iterative deduction. This is a good and intuitive idea and I think I can implement it using layers of neural networks. The ray-based segmentation idea was clever, but had problems finding all segments and was slow. I still don't know if colour-based segmentation is natural.
The methods used in 'A Cue to Shading - Elongations Near Intensity Maxima' to differentiate shadows from textures got me confused, but Gipsman's point that knowledge of shading detection is still primitive surprised me - I think analysis of shading would be fundamental to determining shape and orientation of 3D objects. I agree with her that feedback from higher layers will be essential. But I think the feedback will loop: the shape of the shadow will help in recognizing the object and the shape of the object will help in recognizing the shadow.
'Fast Normal Map Acquisition Using an LCD Screen Emitting Gradient Patterns' presents an innovative method for lighting objects to get 3D information. An interesting point is their use of the polarized LCD light and a filter to remove specular reflection. I found later that the human eye can differentiate linear from non-linearly polarized light (see Haidinger's Brush). Perhaps the brain can use this information in determining where the light is really coming from?
'Realtime visualization of monocular data for 3D reconstruction' was a treat for me because it relates so well to my planned measurement with a camera project. To me, this paper is like an instruction book on how to model 3D space from a single camera. I must look into its Simultaneous Localization and Mapping (SLAM) methods and other tricks. Monocular is cheap, stereo is more accurate? Again, the system doesn't have a clue what its looking at, but it may be a good start for a more complex system. I must analyze to more depth.
'Object Class Recognition using Quadrangles' is a general-purpose implementation of edge-based object recognition which also considers colour uniform regions. On top of this the authors implemented a structural descriptor (the paper describes quadrangles only, but the speaker described use of ellipse descriptors in their newer work) and template-based spacial relationship matching system much simpler than that used by Sinisa Todorovic's self-learning segment-based system (but less capable). 'Geometrical Primitives for the Classification of Images Containing Structural Cartographic Objects' is another system base on edge/ region/ structural descriptor, but focused on the single problem of finding roads, bridges and such in satellite imagery. The software seems to be more capable, handling higher level primitives such as blobs, polygons, arcs and junctions. It uses AdaBoost binary classification. Results were good except for detecting bridges. I suggest looking for the bridge shadows.
Most of the motion tracking papers used very simple recognition techniques or did not describe them. '3D Human Motion Tracking Using Dynamic Probabilistic Latent Semantic Analysis' presents a highly mathematical approach that seems to be another form of template matching. It works well but it will take quite an effort for me to understand it. 'Visual-Model Based Spatial Tracking in the Presence of Occlusions' presents a pre-processing trick to mask occlusions from a template/visual-model based 3D tracking system. While the system is highly performant, using the GPU, it is highly specialized to a single object. 'Automatically Detecting and Tracking People Walking Through Transparent Door with Vision' tracks Harris corners through time. It can be taught to subtract expected movements from new ones, by simple geometric trajectory comparison. This can serve many applications, but the use of just corners means it can't tell you what is moving through the scene. But could specializations like this be used as keys to brain behaviour? e.g. does the brain just use the moving corners of a door to perceive it? 'Invariant Classification of Gait Types' classifies body movements by comparison to a database of shape contexts derived from template silhouettes. This is an efficient and accurate method used in handwriting recognition, and I think I'll look into it more, because the bin concept applied to pattern matching lends itself to implementation using neural networks.
'Active Vision for Door Localization and Door Opening using Playbot' is another specialization - for doorframes and handles. Its advance is active vision. The robot solves its position geometrically after it detects the door by using a pre-programmed door size - meaning it will only work with one size of door. The active vision part is that the robot takes pictures at multiple angles and multiple positions and solves its position using the camera angles and the detected door edges and corners, then calculates a move to a new position. 'Automatic Pyramidal Intensity-based Laser Scan Matcher for 3D Modeling of Large Scale Unstructured Environments' tackles an incredibly hard problem of mosaicing adjacent spherical laser images without feature, location or rotation information by matching their overlapping depth values. This is useful for other mosaicing problems, but I don't think its complicated methods will be required in most computer vision applications which will have less interval between images and can rely on feature detection and sense of place. '6D Vision Goes Fisheye for Intersection Assistance' shows that using fisheye lenses provides a wider angle of view with only small hits on their low processing time requirements and relatively wide accuracy requirements in a real-time stereo mobile object tracking application. 'Challenges of Vision for Real-Time Sensor Based Control' explains how additional sensor input to an extended Kalman filter can be used to supplement poor video data caused by bad camera angles.

One thing I brought away from this conference is that although there is a large amount of existing work and many new efforts in the computer vision field, the presented applications are not trying to understand or duplicate biology. They're using mathematical methods to solve specific problems. Well perhaps the concepts can be implemented in neural networks. And the solutions are so specific! I guess it'll be a long time until there is general purpose vision. And not surprisingly so, because that will require general purpose concept representation. Too bad I didn't hear the AI papers too.
Since this was my first academic conference, I learned that what to look for in papers is what is new, or what can be adapted to my purpose. Attending has motivated me to get an IEEE membership so I can access more research papers. Poster presentations seem pretty valueless to me. Either they don't present enough information or I am forced to stand while reading an entire paper.
Another thing is that there is a lot of existing technology out there that can be used to solve problems. A counterpoint to this and a kind of semi-corollary to the first point is that a lot of the existing technology is highly focused, inaccurate and slow, so there is still a lot of research and development needed.
From a business point of view, I got no leads on paying work. Some people at the conference believe that contracting in this field can be viable. But I think I'll have to prove I'm capable by example before anyone will hire me. Since most researchers only solve special cases, another opportunity is to complete a project to make it useful in lots of situations.

Sunday, April 13, 2008

Web Site Update

Nimajin website
I've been busy. Not making any money, but busy. When the jobs aren't pouring in it's time to concentrate on marketing. So I'm prospecting and networking. And I've updated the Nimajin website. It looks much more professional now. I've added a resume and portfolio so people can learn more about me. I'm not real happy with the site yet - it concentrates on my past and not my future. So it'll change again.
The future is not real clear to me now. I like to be associated with science and communications. My intention to concentrate on media processing and content recognition is still attractive to me, but I'm not finding much interest from others. I think it's mostly because of my lack of communication. I'll work on that. Some friends say I've got to jump on the web services bandwagon to make a living. Improving people's advertising or bookkeeping is very useful, but not as incredibly fascinating to me as making a machine able to draw a floor plan of my house from photographs or making it able to fly me through town from surveillance camera input. I know what it'll take to do these and I'll keep working on them, but it'd be great to have some sponsorship so I could afford some more time to work on them.
I think this happens to a lot of people. There's not enough commercial value to our ambitions of making machines more intelligent, so over time we have to abandon our efforts and go where the money is. I'm still holding on!
That isn't really where I wanted to go with this post, but heck I'll post anything. My goal is to perkily point out how great the new web site is and how ready I am to go an extra mile, learn more new technology and build anything you can think of. What I really love is making people happy.

Monday, January 28, 2008

Tune Searching

You know when you remember a bit of melody, but can't place the song? I've got an idea to create a website to accept input of musical notation and find out what the song is and return information, MP3's and sheet music. So all we need is a simple notation input, a database of notation for all songs in history, rights to use the notation, a fast and smart matching algorithm, a web front end, and we've got a music search site!

Input:
  • Maybe make the PC's qwerty keyboard act like a musical keyboard
  • Capture from a midi device
  • Put a musical keyboard on the screen & let them click on it
  • Accept input from microphone and decode the essentials to notation
  • Show their input on screen in some form of notation
  • Allow playback so they can interactively adjust it until it sounds right to them
  • Include some way to adjust the tempo (slider)
  • Include some way to adjust the spacing (timing of each note), like stretching it or a slider
  • Include something to adjust the timbre / instrument they hear on playback

Database:
  • Once the input is in, convert it to some kind of text or binary
  • Encode the database in the same format and search
  • We'll get snippets as input and it'll be timed wrong, out of key and notes will be wrong
  • We'll want closest matches

A brief look for 'music search' returns nothing. All music search is by artist, song name etc, even for sheet music. It seems somebody tried something like what I want years ago at ThemeFinder - I don't understand what input they're asking for. I found out about Abc notation, which stores notation as text, and has software that can play midi and generate musical notation. More looking. 'tunesearch' gives Richard Robinson's Tunebook Search and JC's ABC Tune Match at trillian.mit.edu. They all seem related to ThemeFinder.

Then I found Musipedia. Musipedia is almost what I was planning. It lets you input by keyboard, by drawing notes, tapping in the rythm or by humming, singing or whistling. All the stuff I thought of except the editing. And ugly and kludgy. I tried it. I suck at keyboards and didn't take the time to put in Pachelbel's Canon well enough that I recognized it on playback. It was - almost. Of course it didn't match that to anything I knew. When I whistled it, the closest thing it found that I knew was The Doors' 'People Are Strange'! It didn't find Elvis' 'Hound dog' by me tapping either. I think it's database is limited. The database is like a wiki though - anybody can add more tunes. This is a great idea. Maybe it's search is the problem. I'm disappointed. Musipedia is based on a prize-winning retrieval mechanism by Rainer Typke. I bought his book.

Musipedia included a Google ad from Midomi. At Midomi you to sing into your microphone and it finds a match for what you sing. In contrast to Musipedia, this site is slick, and it works. It helped me get my microphone level correct. When I sang 'I'm Leaving On A Jet Plane' (after doing some talking first) it knew it. It uses more than one rendition of a song to make a match. It encourages people to sing songs that they know and it uses these in its database. It seems to be getting lots of karaoke people.

I don't know if these guys are making money because they're both covered with Google ads. Musipedia could use some polish and Midomi needs instrumental input. So should I persue my idea?

Thursday, January 17, 2008

The Future is Coming!

Software UI is coming to look and act more like Star Trek control consoles.

The iPhone's UI looks like Star Trek graphics. Check out the black background and brightly coloured icons, all behind a glossy glass surface.
Office 2007 is starting to look like that and all of the WPF examples too.


But the really big deal is that Star Trek consoles are touch screens. They are operated by tapping on them. The iPhone operates by tapping, pinching, dragging, flicking, etc. The Microsoft Surface is so like a Star Trek console. It's a big, bright touch console. We are there!


I'm real happy about this. I've always loved the way the consoles looked in Star Trek. I want to be flicking chunks of code around in one of those one day. It makes me wonder whether all of the computer developers love Star Trek consoles too, or if the Star Trek set designers just correctly imagined the future.

Tuesday, January 15, 2008

Vista Spooler.xml

I was working away on my less than year old Vista PC with a 250 Gig hard drive when Windows tells me the disk is getting full! Windows is all like 'disk cleanup' and 'remove some programs'. Serious! Was it the VS2008 I just installed? No.
I used Silurian DiskSpaceChart to find out where all my disk was going. That's a typical disk usage pie chart thing that puts itself in the right menu. You can drill down through the large folders to find the large file. It's OK. It was the first thing I found for Vista.
So what I found was something writing continuously to C:/Windows/system32/spool/spooler.xml, at like 300 MB per minute! I tried to find out who was writing the file. Resource Monitor helpfully identified 'system'. Thanks. I didn't have SysInternals Process Monitor installed, and I didn't want to try while the machine was so sick, so I couldn't get any details on who was doing all that writing.
I rebooted into safe mode and deleted the file.
When I booted back the file was 33k and stayed that way. I think the file is a printer log and the problem is either my HP printer drivers for the 3390 or MS XPS. I recall running WireShark and seeing every machine with those 3390 drivers polling the printer status over the LAN every millisecond.

Sunday, December 16, 2007

Why Straight Lines?

Why should a visual recognition system expect straight lines? The arrangement of photoreceptors in the retina has no straight lines. Very little in nature that may be seen has straight lines. Primitive people do not build angular objects. Yet, I think straightness came from building things. What could be the first straight thing? Baskets? Arrows? Lines drawn in clay? Not even knife edges were straight. Maybe the horizon - out over the ocean. There is no survival value in building a recognition system within an animal upon a framework of straight lines. Except maybe for orientation to the horizon.
Why did Hubel and Wiesel concentrate on straight lines?
The things that are important to recognize are curved. In all directions. Food like berries, fruits, grasses and animals are all curved. The path upon which you walk. The trees over your head.
So can a recognition system be built upon collections of curves?

Monday, December 10, 2007

Artificial Intelligence: Answer the Questions

The point of AI is to get answers to questions.

In music:

  • What is that instrument?
  • Who is that artist?
  • What else have they done?
  • What is similar?
  • Has this been done before?
  • How is that sound made?
  • What are the words?
  • What is its name?
  • Where was this performed?

In images and video:

  • What is that thing?
  • Who made it?
  • What else have they made?
  • What is similar?
  • Where can I get one?
  • Who is that person?
  • What is it made from?
  • How old is it?
  • Who made it?
  • Can you show me more details?
  • How did they do that?

About things:

  • What are the parts?
  • How much does it cost?
  • Where can I get one?
  • How does it work?
  • Who made it?
  • Who's working on its development?

Then there are the 'what if' scenarios: Substitute this for that - what happens?

AI is so complex and deep that it is very easy to get lost in the implementation details. The details need to get worked out, but the point is to answer questions.