IBM's NICA Archive Sports New Search Software

Posted
By: Jim Rosenberg IBM's NICA archive now relies on Convera Corp.'s RetrievalWare search-and-categorization technology to manage text and page files. NICA became CCI Europe's archive and succeeded The Associated Press Preserver. RetrievalWare enables users to index and search text files, HTML, XML, more than 200 proprietary document formats, and relational database tables.

CCI MediaStore's (then the product's name) first user, the San Antonio Express-News also is among the first to use RetrievalWare with NICA, since upgrading to version 5 of the archive in January.

Wanting to move from a homemade photo archive to one able to handle everything from assignment tracking to storage and retrieval (a requisite for its pagination-system choice), the paper found CCI working on a solution in 1998. But by the time its pagination was implemented, said Chief Technology Officer Nina Brooks, CCI had abandoned its archive work in favor of NICA.

The Express-News stayed with CCI, seeking production-system compatibility while "keep[ing] our archive away from production servers," said Brooks. It allowed the paper design input, a good price, and a port to the Sun and Oracle software it already used. After about 12 months, the paper could archive everything and move metadata, after publication, to master images in NICA via a CCI-built gateway.

"Then things went awry between CCI and IBM," Brooks recalled. So, to archive all components, her paper chose IBM, which dropped its Text Miner software in favor of Convera's page-searching product, according to IBM Product Manager Stefano Stinchi.

The paper now archives readable pages and even joins its Admarc metadata to stored ads. Faster text-only searching will come later because librarians are satisfied with the existing DataTimes system.

Morris Digital Library Systems bought RetrievalWare by itself in 2000 from predecessor company Excalibur Technologies to handle text and images of 1.3 million pages dating back to 1786. It uses Convera FileRoom to upload pages to the database. "It gave us ... a pretty robust system that wouldn't take a week to do your search," said the Morris unit's business-development and technology manager, Mark Chapin.

Its pattern-recognition capability is valuable, he said, because scanning of 40-year-old microfilms of 180-year-old pages produced many optical-character-recognition errors. Concept searches -- distinguishing, say, nature (cloud or river banks) from finance (investment or piggy banks) -- also is possible, but "we didn't implement that," said Chapin. But it may prove useful, he added, since the database was redesigned to speed display of large numbers of search results, which pattern recognition can elicit.

Comments

No comments on this item Please log in to comment by clicking here


Scroll the Latest Job Opportunities From The Media Job Board