Friday, June 30, 2006

A Thousand Words or So

by Paul B. Wiener

All definitions of information may have only one thing in common: they assume it is recognizable. One definition I especially like says: it is a measure of how surprising something is. True, that comes from a protein engineering glossary at the Biomedical Centre in Uppsala, Sweden, but that doesn’t make it any less true. I have a special affection for surprising information because most of it can’t be searched for, and when it’s found, much of it can’t be described. Often it’s not even recognized until its effects are felt, like childhood. This makes someone like me, a librarian with a limited memory and a healthy skepticism about truth, feel a little less guilty for not always being able to answer the questions people ask me. Oh, I can find answers easily enough. Answers are the easy part of information. I just have to rephrase their questions sometimes, to make the answers work. How do you teach people to enjoy being surprised: by giving them what they expect?

One way is by visiting web sites with names that provide no clues about their content. And when the site appears, what’s visible also isn’t much of a clue. Not right away.

One such site is Dredge, produced by a group of electronic media students in our own Department of Art. (And fittingly, you won’t find it listed anywhere by using the anemic search engine on Stony Brook’s Home Page either, though I’m sure that wasn’t the students’ intention.) Virtually nothing on Dredge’s front page tells you what it is – a display site for artistic productions and experiments. Instead, you have to look for the hyperlinks across a large dark, teasing space until you decide that those things down there must be the links. The search process is prompted by the page design. Your “visual brain” is forced to overrule your “textual brain.” Many clueless sites are like Dredge, visually stunning, often textless pages made by people in the arts - photographers, cartoonists, graphics freaks, website designers, Adobe acrobats, hashers, neo-cartographers, poets, topologists - even scientists and farsighted young entrepreneurs like Alex Tew, who invented the Million Dollar Home Page. These people love the challenge of picturing those famous thousand words before you can even think of them. Think of the information gained as similar to what we learn from our dreams, even though they don’t have subtitles.

Another quirky graphic artist’s page that lets you decide how to find its informational content is Leif Parsons’ Page. And another is Ian Timourian’s Mandalabrot.net, which focuses on fractals, visual remixes, generative design, and many of the other new forms of art made possible by computer technology. The site Visual Complexity studies the visual display of information by, well, showing it off. It stuns you with its opening page that presents hundreds of unexplained proprietary information design templates. Bit offers some verbal encouragement too: “Functional visualizations are more than innovative statistical analyses and computational algorithms. They must make sense to the user and require a visual language system that uses colour, shape, line, hierarchy and composition to communicate clearly and appropriately, much like the alphabetic and character-based languages used worldwide between humans.”

Even political forums on the web can score points without using words, animating information to appeal to the newer kinds of “information literacy”. Many web sites and blogs, like GPrime.net, Molecular Expressions, and An Atlas of Cyberspaces, use text to introduce links to the latest text-free games, special effects photography, cartography, optical illusions, digital video and flash animations. And let’s not forget the latest craze, the ever-teasing YouTube. These sites offer learning experiences that can be visually instructive far more quickly than they can be explained – or justified - in words.

Information as most librarians and scholars know it uses symbols – not only words, but numbers, formulae, marks – as well as color and sound, to communicate and document experience. More importantly, most librarians use language to describe the symbols. Symbols called “words” tag ideas, facts, events, people, experiences, memories, feelings, observations. We use them to organize various attributes and similarities. The world thus described is sometimes called “recorded history,” sometimes “science,” and sometimes “reality,” and sometimes “searchable.” What do we do about the information that cannot be so described - the recorded stuff that can only be perceived non-verbally? Do we translate it into words? How many translations (copies, messengers, media, generations, reproductions) will records survive before they lose “authenticity”, whatever that is? This problem long intrigued the philosopher Walter Benjamin, who applied it to works of art. But he wrote it before the internet existed. One answer seems to be: some records survive translation better than other records, and better than most works of art.

There’s a fascinating website that draws attention to the paradoxical fact that art reproduced on the web, because it is “lighted from within,” is sometimes more beautiful than the original thing. Before you leave, take a look at Bibliodyssey, a blog about the beauty of book illustration, old and new. Strange concept - book illustration - isn’t it? Who needs it, especially today, when screens illuminate words everywhere? Illustrating books seemed normal enough once, but here you realize that you are celebrating “books” by reading a blog, (there were no “blogs” three years ago), on a computer screen no doubt using a Windows-based GUI (coded in letters, numbers and signals), and looking at a digital image – one which presumably can never degrade, (since a digital fact weighs no more than an idea and can be dog-eared indefinitely) unless electrons themselves disappear into time….Here’s a lovely image of a drawing (click!) that someone scanned from a 500-year-old manuscript, whose maker once had it painted there by an artist – at this moment the “real one” probably sits unknown and untouchable on a shelf in an ancient library in Rouen, a library morphing into a museum. If most information isn’t art, is it still subject to the kind of degradation that copying and translating from any medium produces? And if the image (or the book, or the movie: remember Fahrenheit 451?) remains in my memory long after the printed and digitized one disappears, will it still be authentic?

Wednesday, June 07, 2006

Truth without Consequences

Google has introduced one of those new search applications that’s just the kind of thing that makes many librarians distrust Googled information. It’s called Google Trends and it provides context-free information that seems to have no bearing outside the arcane world of search itself. But no one ever accused Google of humility. This new engine is a seductive toy that at this stage promises much more than it delivers. Here’s how it works: you searcha word or short (parenthesized) phrase, put a comma after it, then select another (and another three, if you like) to compare it to. Google Trends then tells you how relatively frequently those terms were searched over a few years (on a graph with no scales), as well as what countries or cities it was searched in, and in what language (based on the web version of Google). Some frequency highlights are correlated to currents events, but other than that no explanations are provided, or even suggested, as to why something is trending in any direction. Or on how many searches the trend is based on – 866 or 435,205. Some students may welcome this, for who can doubt that a reason exists for a trend? And a halfway decent student can find reasons much faster than he can find facts.

What are the reasons for interesting trends? A search comparing the terms “stony brook,” “sunysb,” “stonybrook” and “stony brook university” shows us that the single word stonybrook is the second search term of choice used in the US, and that mostly English speakers use it, that India is by far the place where the second highest number of searches came from, that Chinese was easily the second most used Google site for this search, that the phrase Stony Brook is used everywhere much more than the other terms seeking “Stonybrookness”, and that outside the campus community, almost no one ever searches with the term sunysb. You can spend hours discovering similar fascinating, puzzling and potentially meaningless trends by comparing all kind of things, places, names, even numbers. For example, try comparing “0, zero and nothing.” Or “da Vinci, da Vinci Code, Dan Brown and Leonardo.” Which of these five concepts do you think is most searched: “truth, fact, fiction, information and myth?”

۞ I forget just how I came upon this next site, The Athanasius Kircher Image Gallery, at the Stanford University Library. Probably it was randomly from a blog listing. Here are some literally fabulous rarely-seen illustrations of rarely-seen phenomena created by the great, albeit unknown, 17th century German Jesuit polymath. What, you never heard of him? You’ll need to download the DjVu plug-in to see his work. But why were they digitized? The site has impeccable bona fides and links to some fascinating pages about the man, like The Athanasius Kircher Project at Stanford University, The Correspondence of Athanasius Kircher, The Societa Italiana di Storia della Scienza in Florence, Project Director Michael John Gorman, a lecturer in the history of science (at one time) at Stanford

and now a scholar living in Dublin, Ireland, the program of a 2001 Colloquium on Kircher, a gloss on an Exhibition of his Baroque Encyclopedia, an article about Kircher scholarship from the Chronicle, and Wikipedia’s inevitable page on the man. And of course there’s The Proceedings of the Athanasius Kircher Society , which renders the obscurity behind Kircher’s genius into blog-like clarity.

About the only thing I haven’t been able to find is something that explains how his genius impacted the world - though he did make an appearance in an Umberto Eco novel and corresponded with over 750 people. Most of the web sites about Kircher seem full of superficial details. I suppose it takes a certain kind of genius to publish and illustrate dozens of books on many major and arcane subjects in his lifetime, but where’s the rest of him? I haven’t been able to see any of his actual correspondence online (server problems), but there’s no doubt that the images displayed in the Gallery are exotic, deliberately impossible, funny, skillfully drawn and suggestive. It is good to know that we can find images of his Musurgical Ark or his Tarantula and the Musical Antidote to its Poison, even if we aren’t told what they mean. Much of Kircher’s quoted text is either untranslated, is vague, is given no context, is written in strange symbols, or refers to sources that only a grant could verify. In fact, it all would make more sense if we knew it was a hoax, but it isn’t. Still, you won’t find these images on a proprietary database. Who would expect you to pay to see his works? No one. Kircher’s fascinating inaccessability is being paradoxically extended by something libraries will always do well: preserving the virginity of information for its marriage to scholarship.

۞ Statistics don’t lie, they just never tell the whole truth. What reference librarian hasn’t had a student come up and say something like “I need statistics on how many unmarried Japanese men under 40, over 5’8” and under 150 pounds, living abroad and earning less than $50,000 a year missed the train to work because they were shaving on Tuesday?” While there are librarians who attempt to look for answers to such questions, using the usual resources, I rarely do. Instead of telling them the correct answer - 7,387 (J. Ethnic Shaving Sociometrics) - which they never believe, I try to explain that there simply can’t be statistics for such fact . What I can’t say is that I refuse to spend months of searching, analyzing, computing, cross-linking, interpreting and translating to find out that there is no answer. But they wouldn’t accept that either. Why should they, when everyone knows there are fantastic sites offering statistics about baseball, the census, libraries, pregnancy, literacy, jobs, mortality rates, crime, food and movies? Many excellent sites gather tens of thousands of statistical studies and let you believe you can search them. Go ahead, make their day. Have you ever looked at StateMaster or Statistical Universe? If the numbers are there, someone has probably crunched them. But that’s the catch: where do the numbers come from?

It simply doesn’t occur to most people that no one sits around at a desk all day at the center of the world, keeps statistical track of everything happening in the universe at a given moment and relates it to everything else that happens. Only one being does that. For the rest of us, there’s not enough time to track what occurs in time. Maybe Google is working on it. One of the strange lessons of information science is that there are statistics for everything under the sun but what you are looking for. Absorb this lesson, and statistics become less intimidating, almost recreational. Hank Aaron has the all-time home run record? Yes, and he also grounded into double plays – 328 – more than any other player in history – except Cal Ripken, Jr. We know this because there is such a thing as a double play, innings, outs, and rules in a game called baseball. When consequences rule life, statistics will always come to bat. There is such a thing as an infant death, as box office receipts, as employment rates, as book circulation, as starvation. And then there’s everything else. But what about those expatriated Japanese men who shave every Tuesday? Surely they exist, surely they matter. But where are the statistics? Wasn’t anybody watching these guys? When all the information in the world in all languages from books, articles, indexes, statistics, reports, libraries, institutes, laboratories and think tanks - is linked, will librarians be able to answer such questions? Of course not.

Paul B. Wiener

1 June 2006

Tuesday, May 02, 2006

Reality Checks

Reality can be an interesting source of information. Or is it the other way round? I use the web all the time to find out about the real world. Don’t you? And what I find is “information.” If it isn’t, what is it? I use reality all the time too (in the form of a computer, a chair and the electrical grid) to find ------information. Defining information can seem like a silly word game, self-indulgent, like defining reality or god or love. It’s what undergrads do late at night after smoking a few beers. Well, call me sophomoric, but I’ve always wondered: how do you separate information from reality? Should you? Perhaps reality is to information what god is to religion.

Baba Ram Dass’ famous book title gave us a clue about how people can be real: Be Here Now. Reality, he means, is the attention we paid to what exists where we are in the present moment. Is that clear? It sounded pretty good in 1971, when LSD was still a major search engine. But what kind of criteria are these now for evaluating the “reality” of information on a website? Can information be as real as a person? Does a website have to refer to information that is “present,” or is being generated while you’re looking for it? Does it need to be verifiable or observable in an independent source, or can it be real simply because it was generated by a clever algorithm, and turned by software into a graphic display archived on the web? Is there such a thing as good information? When is information a “primary source” and when is it the content contained in that source? It might be interesting to examine a few websites with these questions in mind. Have you ever played “Where’s Waldo?” Find the reality in these websites.

The Voynich Manuscript site provides a great deal of information about this famous artifact, which currently resides at Yale. The images abound, along with the history. The manuscript is distinguished by the fact that no one understands a word of it, or can trace its complete provenance. All agree it is strikingly beautiful. Its language resembles a few other alphabets and some of its drawings resemble real plants from real places. The Voynich Manuscript, named after the antiquarian dealer who acquired it in 1912 from an undisclosed location in Europe, practically shouts for attention. It looks well put together. Possibly it was composed in an artificial language that had a meaning. It may be an herbal and cosmological codex. Its mystery has outlasted several deciphering hoaxes and well-reasoned attribution of authorship to Roger Bacon, the medieval scientist. But really, no one knows what it is. It is a real artifact with a real history. It has attracted real scholarship, and there is little doubt that it’s been around at least since the 15th century. But it may be meaningless all the same, or permanently unknowable. The web may extend its life forever. Observe a bibliographic freak: a famous, historical primary source waiting for its first citation.

But anyone can create a manuscript. Try creating a person who’s still alive after more than five human lifespans, yet who doesn’t even exist: Hamlet, the most celebrated literary nonentity of modern times. HamletWorks is a site devoted to this young man, about whom as much has been written as about his creator. Hamletworks uses a deliberately textual approach rather than a performance-oriented one, but then, Hamlet always was a man who needed to explain things. In addition to a line-by-line hyperlinked analysis of the play, Hamletworks, largely the outgrowth of work done by Professor Bernice W. Kliman, contains the full text of many different editions of the play, a link to the Century Dictionary and Cyclopedia, the complete run of Hamlet Studies, several concordances, critical and historical essays, illustrated versions of the play, The Enfolded Hamlet, and an extensive bibliography. All this for someone who didn’t exist. Do literary critics extract reality from textual information, or just use it as a gateway?

Making of America Chopin Early Editions African American Sheet Music

Information doesn’t begin as information. Thought has to be expressed, then formatted and recorded. If information is a recipe for action and discovery, its ingredients can be words, numbers, images, sounds, formulae, signals, the medium itself. Doesn’t all information come formatted? Sometimes information first appears as a primary source: being in the presence of the medium that first represented it lends it a special authority. It’s still quite thrilling to see actual words penned by Jane Austen in the original notebook in a display case at the Morgan Library; or to see Chopin’s notation; or the old illustrated song sheets written by the children of slaves – a century-old equivalent of a music download. Presumably, it will be just as thrilling for your child to see the same thing in a digitized version. And in 200 years? Maybe one of your teeth will contain the Library of Congress.

Librarians and researchers alike need to remember that many digitized documents are not available in proprietary databases, yet most library or catalog search engines don’t enable you to search among all available digitized resources at once, whether they were produced as marketable database, or produced as a labor of love. How many proprietary databases are put together with the imagination and painstaking care shown in the sites above, or in Library of Congress’ free American Memory collection? Libraries are starting to put websites in their catalog, but the descriptive terms are still problematic, and until it becomes routine, you’ll have to be resourceful.

What a welcome contrast to the Voynich Manuscript: digitized primary documents with origins no one doubts. They look so authentic, even though it’s now easy to create your own electronic documents – letters, news reports, handwriting, photos - in formats that look 100 or 300 years old. What reference librarian has not been tormented by an innocent request for primary source material? It should be easy to supply a patron with many, and the sites above make it so, for certain students. But wait. Isn’t most of what’s in a library a secondary source? Should a professor, or her students, ever settle for a digital version of a primary source? I’m not sure. Isn’t history authentic even if we didn’t witness it? Or are the physical qualities of the documentation as important as memory in verifying that what was recorded actually happened? In academia scholarship is meant to trump digitized evidence, but it’s still allowed to use it. If you treat a document like a photograph, and retouch the image of its text, who’s to say that that new Constitutional Amendment they’re talking about won’t be Photoshopped overnight! People love to manipulate images: it’s fun, and usually legal. Visit this site and see for yourself..

All the labor and love that go into most digitizing projects are wasted, however, if a teacher insists that a primary source for research can never be retrieved from digital holdings. Recently, I was told of a professor who insisted that her student photocopy an article from a journal, rather than download a simple text version. If the student had made copies from an online source of the same article, but digitized in its original format, the joke would have been on the professor. But then, can we be certain a manufactured image couldn’t be reproduced as a full page of text in a print journal? As usual, Walter Benjamin had it all figured out 70 years ago, in The Work of Art in the Age of Mechanical Reproduction. Today, the ALA has a useful website that tries to make this all much clearer than I can, Using Primary Sources on the Web.

And now let’s change directions again. Here are two sites that let you create your own information, and then find a reality that matches it. That’s easy enough to do with a normal text-based search engine, where you can put in any text or number or string and fish for matches, but here it’s a little different. With Retrievr you can create shapes and images in different colors, using your mouse, within a small blank white rectangle, and then ask the search engine to match the general pattern and color against Flickrs huge databank of photos. Admittedly, compared to an infinity of possible designs it’s a small database to choose from, but it’s a start. What’s worth pondering is this: what possible use could finding this information serve? It’s easy to sense that there must be one. Tunespotting lets you do a very similar thing with sounds, using note and rhythmic patterns, and matches your input to a large but also limited database. Using your keyboard or mouse you can enter approximate musical note, phrase and rhythm sequences and the program matches it against similar general patterns of known compositions (like “Happy Birthday” or “The Moonlight Sonata”) that have already been encoded. It’s not terribly accurate at present, but think of the possibilities! Using very fuzzy search terms, both programs provide excellent examples of how fine and subtle are the elements that constitute an individual composition. And both allow you to see if reality measures up to information.

Paul B. Wiener

1 May 06

Thursday, April 13, 2006

SCREEN SAVERS

Where does information come from? From the world, from the senses, from the mind? If sometimes it's self-evident, at other times it has no fixed form. So how do we know it when we see it? There is a running debate among some physicists about whether information can escape from a black hole. Current opinion says that it can. As quoted in Wikipedia, "Quantum perturbations of the event horizon might allow information to escape from a black hole." I have no idea what this means, but it is a theoretical fact, and evidently it occupies a privileged place in the mind of Stephen Hawking. To me, most information is like this: always on the move and about to be true, shape shifting, something to be declared, compelling our attention as it barely escapes the black hole of reality. Information, when we can remember it, gives us a way to describe the black hole we can't see. Reality itself has no meaning: it just is. Information mediates between states of being. This little column will show you some of the sites that have changed the nature of attention. With consciousness our screen on the world, information is our screensaver.

How does information get on the web? Well, people usually put it there, of course. Providers of information to the academic community are usually responsible professionals - writers, professors, journalists, publishers, educators, statisticians, scientists, government workers and others whose business is displaying, marketing and distributing information to those who need to believe that it comes from official, scholarly, peer-reviewed, verifiable, secured sources.

But information isn't generated only by researchers, nor only by people. Habit and method puts information out there too. Time, history, inertia, software, algorithms, computers, and search engines do their share. So do faith, creativity, egomania. Some information seems simply to emerge, like fractals. What musicians, novelists, architects and photographers create - gods, heros and villains - also becomes information. The brilliant, twisted, or mute obsessions of pilgrims and hermits, victims, soldiers, and inmates also become information. Selves gather unto selves, electricity makes waves, voices chat from collective throats and sound like nothing heard before. Virtual reality becomes more than a metaphor. Sometimes whimsy, accident and fanaticism put information there. No one knows what it will lead to. What does it have to do with librarians?

A quiet battle rages among academic librarians between those who believe that the information needs of patrons should come from databases on the "invisible web" and those who believe that authority should be drawn from the entire web. Of course it's easier to use the invisible web: it's a much smaller part of the world wide web. The trick is in getting to know and trust it: it takes time and intelligence. Proprietary databases are excellent for those who have password or proxy access to them, and the good luck to know where to find them; the general web is what's left, and is created by and for the masses. The better we get to know it, the more we shall be amazed by it.

Most of a library's databases consist of copyrighted "articles" and reports that must be purchased or leased, and are written by credentialed experts. Composed of indexed research and data submitted by experts and scholars to editor-peers, the "information" is retrieved and displayed in formats modeled on the print publications they replace. You pay a database vendor for lending you its authority, not for its convenience or ingenuity. The long arm of print culture often drags them slowly across the landscape of proprietary protocols. The free web, on the other hand, can seem like a lawless, borderless place where anyone can stake his authority on format and rhetoric, Google is king, experts are self-described and keywords like photons illuminate cyberspace generating relevance.

Now let's look at some of the internet-based websites that contain the shapes information weaves across the web. In their creativity, subject matter, graphic design, user interface, and application, websites like these can often provide far more than is asked or expected of the conventional panoply of library databases. But they do little good if they languish in neglect. Take a moment and look at them. They are important, independent information sources with a life and a magic of their own. Some of them have taken light-years to reach us. To base information literacy solely on knowing proprietary databases is like learning only the nouns of a language. Such a tongue-tied approach to our search for knowledge is bound to be disappointing.

1. International Archives for the (Hammond B-3) Jazz Organ

Try to imagine anyone other than a music fanatic from Germany putting together a database like this, annotating over 10,000 recordings made by jazz organist on this famously definitive instrument, making it searchable by many parameters, and offering it as a downloadable archive. Links to everything else in the Hammond jazz world - concerts, videos, festivals, new recordings, polls, history, discussion forums - are also given. Why are these labors of love always free?

2. Whichbook.net

Having a hard time finding a good novel to read, and you don't like your husband's taste? Have a computer connected to 150 responsible human readers do it for you! This engine allows you to find new titles by selecting the degree of appeal in a book of such qualities as safety, violence, sex, seriousness, sadness, optimism, race, age, setting, and thematic description. Only books published after 1995 are listed, but the century's young. After you input your preferences, titles are summarized briefly. If you live in the UK, you can even find out where to borrow it.

3. Feature Films from the Internet Archive

Yes, at no cost you really can view or download over 640 full-length feature films on your computer, like Salt of the Earth, Steamboat Bill, Jr., The Quiet One, Malice in the Palace, The Last Woman on Earth and New York City Ghetto Fish Market 1903. When you've seen them all, there are still another 30,000 unusual short films you can explore. Having a powerful computer and connection helps.

4. Speech Accent Archive

Ever wonder why American actors never speak with a foreign accent? Hear how the same paragraph in English sounds when spoken by hundreds of foreign-born speakers, including those whose first language is Punjabi, Hunanese, Afrikaans, French, Sindi, Dinka, Serbian, Pidgin English, Mortlockese, Tagalog, Swedish and Yoruba.

5. Avibase

Why spend weeks looking for an Adelaide's warbler, a three-toed jacamar or even an Ivory-Billed Woodpecker when a few mouse clicks will fill your virtual checklist? Here is a searchable metasite in 13 languages with links to every imaginable aspect of birds and birding worldwide: descriptions, maps, checklists, sounds, photos, videos, research centers, conservation, travel, trip reports, bird name etymology, books. But do birds fly in cyberspace?

6. Live Air Traffic Map from JFK

See a dynamic (if slightly time-delayed) map of flights departing from and arriving at JFK Airport. Gain new respect for air traffic controllers. Individual flights are marked, altitude and speed are given, aerial views and major roads are shown from a 5-to 80-mile radius. With this software, some other airports can be seen as well. Should this site exist?

7. Exploring The Wasteland

Where were you when you first heard the droning recital of T.S.Eliot's The Wasteland? In class, or asleep - or both? Future students never need lose sleep over this assignment again. This site provides a detailed, automated line-by-line critical, textual and historical commentary of the famous poem, and was created by a software engineer who, though he has no degree in English, didn't major in it, doesn't teach it, and took only one course in English Lit, has a thing for Eliot, and is proud of it. See under: labor of love.