By: Jim Rosenberg Technology Could Be Used For Digital Delivery
Mediterraneans know a thing or two about preserving the past. At
Nexpo 2001, the newest ways to save, show, and spread the printed
word came from lands that gave us the Psalmist and the Gospels of
the oldest best-seller and first beneficiary of movable type.
Systems developed in Israel by Olive Software Inc. and in Greece
by Lambrakis Press combine advantages of print presentation,
digital files, and Web browser accessibility.
Using its archiving and daily home-delivery technology, Olive
aims to enable publishers to "leverage printed-edition assets"
without relying on an Internet production staff, "allowing them
to make money online for the first time" -- and from only a
fraction of their Web sites' total number of visitors, said
Marketing and Sales Vice President Alon Men.
Whereas Denver-based Olive is a venture-capital-backed software
developer, family-controlled Lambrakis is a multiple-media
company that grew out of newspaper publishing in 1920s Athens.
Publicly traded since 1998, the company also financially backs
new-technology startups. Also, three years of in-house
development have concentrated on archiving, and it seems more
restrained in outlook. "It's probably not the right time yet,"
said Nickolas Gouraros, digital technologies department director
at Lambrakis Press Archives. He expects "the market will open" in
four to five years, with relevant technologies improving by that
time.
Nevertheless, Lambrakis has a working product -- e-preserve
newspapers -- and a digitizing service that it says was selected
to scan hundreds of thousands of newspaper pages dating back to
1890. Already it has digitally preserved a million of its own
titles' pages, from which it has indexed 400,000 articles, most
from 1922 to 1944, some through 1985.
Though both firms had their first Nexpo booths last month in New
Orleans, it was neither's first crack at newspapers. Olive's
founder went to the same show in the mid-1990s under another
firm, Iota, as a guest in the booth of Hyphen. After those Greek-
named firms respectively left the market and went out of
business, real Greek geeks began "working from scratch" along
similar lines. In May, Lambrakis took e-preserve newspapers to
the Newstec show in Brighton, England.
From scanned paper copies or their microfilms, the companies make
browsable, readable, and searchable digital files that preserve
the look of the original pages. Not only may stories, art, and
ads be viewed as they first appeared on a page, but the on-screen
image of that page may be used to conduct a search, read a story
that jumps to another page, or navigate an edition. Pages are
identified, scanned, and input into the systems through software
that recognizes such things as headlines and bylines to
distinguish and digitally clip out stories and other page
components.
Olive offers two products that share core technology. Its
ActivePaper Daily recognizes and stores, in extensible markup
language (XML), all structural components on a page made up for
print, with further tags applied to preserve every possible
element. "It's like having a PostScript in XML," said Men, so
that a paper can post its "printed edition online in its original
format." Olive says its SmartScroll technology efficiently
displays news pages on 15-inch monitors.
Readers would be able to view the on-screen newspaper pages
wherever they have access to a Web browser. For each subscriber
choosing delivery to a digital doorstep, a publisher dispenses
with almost all post-pagination costs of producing, packaging,
and distributing a copy and can still include the subscription in
total paid circulation. It also may represent added value for
existing paper-and-ink subscribers and for advertisers, said Men.
Subscribers' "papers" arrive fast and always dry -- ready to be
printed as needed. A ticker atop the page can carry breaking news
at any time. What's more, subscribers need not download an entire
edition. From an image of each page, readers may preview stories
before choosing what to download. The software recognizes jumps -
- readers simply click on the "continued" notes. Readers also may
switch to "fast-view," in which the entire story appears in
straight text.
Men said ActivePaper Daily is devised to value editors' work.
Even when users personalize papers by setting up subject-based
profiles (company, country, team, etc.) that highlight relevant
editorial and advertising content, they may always read an entire
edition. Clicking on a keyword within a highlighted item moves to
the next highlighted item. A "see also" feature, said Men, "links
the article to other articles or to external sites" selected by
an editor. The original story can be "frozen" while checking
links. Because a word search (which may be restricted to stories
or ads) occurs on the original page image, it relies on bitmap
indexing of word-pattern images.
Copyright management built into ActivePaper Daily and a second
product, ActivePaper Archive, for material (from wire services,
free-lancers) permitted on paper but not on screen automatically
blurs unlicensed material on a page view, locks it from access or
searches, and notifies users that it is protected. The rules-
based process relies on intelligent recognition of bylines and
photo credits. Longer term, according to Men, because ActivePaper
can measure access by article or picture, should publishers and
contributors create clearinghouse arrangements, his product could
be used to pay copyright holders by frequency of access to,
revenues generated from, and/or size and position of articles.
Used by the British Library and in tests at newspapers,
ActivePaper Archive brings together in electronic form microfilm,
clippings, entire saved paper editions, and recent digital page
files. By affording control of components of digital page files -
- from live editions or from scans of archived paper or film --
the segmentation software takes over the job of a librarian with
a pair of scissors. Unlike a clip file, however, users may see
any story from a search as it appeared on a page, and may explore
other material in the same edition.
When a search returns a list of items (stories, columns,
captions, ads, etc.), clicking on an item brings up a page image,
with the item outlined and search terms highlighted. Clicking on
the outlined item brings it to the screen in a readable size (as
it originally appeared, unless "fast-view" is selected). Other
items from the same page or other pages in the same edition can
be similarly examined. Only when a component on a page is clicked
is a readable file downloaded. The software contains links
between pages of the same edition and between a page's image and
its content files.
Although verbal content of page images scanned from newsprint or
microfilm is extracted by optical-character-recognition software,
OCR is not 100% accurate. It gets worse, said Men, when applied
to newspapers' multiple typefaces and type sizes for everything
from headlines to agate statistics to body copy set in small type
with tight leading. Olive's software makes these distinctions in
segmenting components, and it can distinguish scratches on
microfilm. OCR handles searches with help from Olive's Adaptive
Probability software, said Men, to "apply fuzzy logic only ...
where there's a high probability of a mistake."
Using OCR only for searching limits the errors it can cause. A
user reads words as they appear on original images, so the mind,
said Men, does a better job of compensating for print defects and
aging paper.
Olive, too, digitizes paper or film. It runs sample customer
microfilm reels to calibrate its scanners, enabling them to run
as fast as possible without sacrificing accuracy. Scanning
original paper copies, said Men, takes longer and costs more.
Olive also will accept in-house and third-party volume scans.
The company has licensed its technology to the Online Computer
Library Center, Dublin, Ohio, which will market its 50,000
libraries' newspaper collections. Recognition and control of
individual stories, he said, is what will allow the sale of
collections by subject matter.
Because each job presents its own challenges, Lambrakis also
first processes sample customer data or scans, which it returns
for approval. Gouraros said it refined its techniques by trying
them out and making mistakes on its own publications first.
Lambrakis' software also uses segmentation of page components,
and it, too, limits display according to copyright constraint.
But one week before the U.S. Supreme Court ruled on the Tasini
case, Gouraros said, "After five or 10 years, I don't think we'll
have a copyright problem."
E-preserve automatically fills a column of metadata fields, which
can be changed and added to manually. Its archiving, said
Gouraros, provides "quality-control tools for all procedures."
Page scans saved in tagged image file format (TIFF) are converted
to portable document format (PDF), in which outlined stories
called up from search-result lists can be viewed on the browser
screen using the magnifier tool.
"The next step is ... automatic extraction of drawings, tables,
and adverts," said Gouraros, adding that work also is under way
on Web page archiving and on authentication that aims to ensure
there is no alteration from the first scanned image, giving
researchers the same confidence they have using bound volumes of
original pages.
Jim Rosenberg (tech@editorandpublisher.com) is a senior editor covering newspaper technology for E&P.
Copyright 2001, Editor & Publisher.
Comments
No comments on this item Please log in to comment by clicking here