Sunday, December 30, 2012
PDF/A format
Diary of a DAS Student
Thursday, April 26, 2012
Metadata Deluxe / Top 10 List of Embedded Metadata Properties
Metadata Deluxe / Top 10 List of Embedded Metadata Properties
Friday, October 21, 2011
NARAtions » National Archives Digitization Tools Now on GitHub
Tuesday, March 15, 2011
A New Image Site - Ookaboo
His group Ontology2, has created the Ookaboo website, which contains digital images that they claim are either in the public domain and or under Creative Commons licensing terms. It appears to be a collection of harvested images from the web, in particular Wikimedia, where people place images that they wish to share with the world, so the free access is probably correct.
The interesting part to me, and I think to many of you, is how they are indexing the images , which is by means of the semantic web.
Here is their description of what that means.
Images on Ookaboo are indexed by terms from the semantic web, the web of linked data. Although you're free to find images through the human interface, automated systems can quickly find and use images through the semantic API.
Ookaboo has two goals: (i) to dramatically improve the state of the art in image search for both humans and machines, and (ii) to construct a knowledge base about the world that people live in that can be used to help information systems better understand us.
Semantic Web, Linked Data
In the semantic web, we replace the imprecise words that we use everyday with precise terms defined by URLs. This is linked data because it creates a universal shared vocabulary.
For an example, in conventional image search, a person might use the word "jaguar" to search for
• the animal
• the automobile brand
• the Jacksonville Jaguars (NFL team)
• the game console from Atari
• ... and nearly 30 other things that are listed in Wikipedia.
Note in the cases above, there are pages in Wikipedia about each of the topics above: it's reasonable, therefore, that we could use these URLs as a shared vocabulary for referring to these things. However, we get some benefits when we use URLs that are linked to machine-readable pages, such as http://dbpedia.org/resource/Jaguar, or http://rdf.freebase.com/rdf/biology.itis.180593
Pages on Ookaboo are marked up with RDFa, a standard that lets semantic web tools extract machine readable information from the same pages that people view.
Named entitiesThe above information is from their About us page, which I highly recommend you check out.
Ookaboo is oriented around named entities, particularly 'concrete' things such as places, people and creative works. With current technology, it's more practical to create a taxonomy of things like "Manhattan", "Isaac Asimov" and "The Catcher In the Rye" than it is to tackle topics like "eating", "digestion" and "love". We believe that a comprehensive exploration of named entities will open pathways to an understanding of other terms, and hope to extend Ookaboo's capabilities as technology advances.
Oh and yes their images are pretty good too, especially for those interested in buildings. and other "concrete" things.
Monday, January 3, 2011
Embedded Metadata News
Embedded Metadata News
Friday, November 12, 2010
University of Chicago to use Finding Aid for large scale Digitization Initiative
But Kathleen describes the whole project so much better, Read it for yourself.
11/11/10 to Digipres from Katheen Arthur:
The University of Chicago Library’s Special Collections Research Center has a launched an initiative for the digitization of archives and manuscript collections. The digital images are being made available via the online finding aid for each collection. This will recreate for the online user the experience of a researcher encountering the original materials in the SCRC Reading Room, with documents displayed as they are housed in each folder, and with description of the contents in the form of folder headings.
Individual, high-resolution images of each page will be permanently preserved in the Library's digital repository, and can be made available for publication or other research needs. Due to provisions of copyright laws, digitization efforts are currently focused on materials in the public domain, or those for which the University of Chicago holds copyright.
Collections with digitized content now available online include:
- The Ida B. Wells Papers contain diaries, correspondence, manuscripts and photographs documenting the life of the teacher, journalist, and anti-lynching activist. http://pi.lib.uchicago.edu/1001/scrc/ead/ICU.SPCL.IBWELLS
- The Dr. Harry Bakwin and Dr. Ruth Morris Bakwin Soviet Posters Collection contains nineteen Soviet political posters produced in the early 1930s, collected by the American physicians Dr. Harry Bakwin and Dr. Ruth Morris Bakwin during two trips to the Soviet Union. http://pi.lib.uchicago.edu/1001/scrc/ead/ICU.SPCL.BAKWINPOSTERS
- The Fielding Lewis Papers contain business, personal and legal records documenting life on a plantation on the James River in Virginia, both before and after the Civil War. http://pi.lib.uchicago.edu/1001/scrc/ead/ICU.SPCL.LEWISF
- The University of Chicago Laboratory School Work Reports are made up of reports about the Elementary and Secondary division of the Laboratory School, and document classroom activities in the School's first decade. http://pi.lib.uchicago.edu/1001/scrc/ead/ICU.SPCL.LABSCHOOLREPORT
- The Jefferson Davis Trial Papers document the legal entanglements, ambiguous delays, political floundering, and shifting of responsibilities surrounding Jefferson Davis' first indictment for treason. http://pi.lib.uchicago.edu/1001/scrc/ead/ICU.SPCL.MS979JDavis
- The Thomas Winston Papers relate primarily to Winston's activities as a surgeon with Illinois troops during the Civil War. http://pi.lib.uchicago.edu/1001/scrc/ead/ICU.SPCL.WINSTONT
- The Middle Eastern Poster Collection produced by government offices and private organizations, primarily in Iran and Afghanistan. http://pi.lib.uchicago.edu/1001/scrc/ead/ICU.SPCL.MEPOSTERS
For those of you who are interested in the production details, the Large Scale Digitization Initiative is a collaborative effort involving the Special Collections Research, the Preservation Department, and the Digital Library Development Center. The initiative has been guided by definitions of, and requirements for, mass digitization provided by funding agencies such as the National Historical Publications and Records Commission. These guidelines stress expedited scanning workflows, without sacrifice of image quality, and with close attention to preservation concerns, and the use of existing descriptive metadata, such as that provided by a finding aid.
The collections are scanned by Preservation staff. The documents are scanned in color, in the order in which they are filed in each folder, and a .TIFF file is created for each page image. A naming scheme is used for the files which can be extended to other collections scanned as part of the initiative. The TIFF images from each folder in the physical collection are combined into PDFs for delivery. PDF was chosen as a delivery format because of its simplicity, stability and ubiquity. It is expected that the vast majority of users will have PDF viewers on their computers, and will be able to use them to enlarge, decrease, rotate, print, and otherwise easily view the images. Although the images are delivered as pdfs, the tifs of each page will be stored in the digital repository, and will be available if needed for other purposes.
Links to the digital files are added to the online finding aid by SCRC staff. The Encoded Archival Description (EAD) tags chosen allow links to be created at any level of description in the finding aid, from series, to folder, to item, and for multiple links to be attached to a particular description. DLDC staff updated the style sheets to allow display of the links in the finding aids database. DLDC is also hosting the digital files, which will be retained in and delivered from the digital repository.
The procedures developed for large scale digitization of archives and manuscript collections are simple and extensible. Future plans call for the delivery of digital audio files and, eventually, of born-digital content, and for full-text searching of digitized typescript documents, which can be made keyword searchable through optical character recognition (OCR).
Large Scale Digitization Team include: Eileen A. Ielmini, Kathleen Feeney, Daniel Meyer, Kathy Arthur, Karen Dirr, Charles Blair, and the student scanners in Preservation.
Friday, September 3, 2010
Metadata for Digital Content (MDC), Developing institution-wide policies and standards at the Library of Congress
From the site
Metadata for Digital Content (MDC)
Developing institution-wide policies and standards at the Library of Congress
Over the years the Library of Congress' digital projects have generated many digital objects and these objects have been given various levels and types of descriptive metadata. The Library has assembled several use cases that require a more coordinated and standardized approach to the creation and management of this descriptive metadata. A few examples of use cases are:
* Geographic navigation of Library of Congress digital content
* Temporal navigation of Library of Congress digital content
* Exchange video and audio data with external services
Metadata of varying degrees of richness is necessary to support the use cases.
As part of this effort an institution-wide working group was established and is making the following available for use by any interested institutions:
* A master metadata element list with recommendations on best practices for populating the elements to provide more consistency of new metadata creation throughout the institution, support the Library of Congress metadata use cases, and point to areas where metadata remediation of current metadata might be beneficial.
Check it out
Wednesday, May 12, 2010
How to catalog Apples and Oranges?
The key is flexible but clear rules.
First, define a core group of fields that are to be filled all assets. You will want that information which will aid in the management of the assets, such as:
- Location of digital file
- Title
- Usage rights
- Owner/creator of digital file
Secondly, define data standards before you start - i.e. decide what field will be used for which particular information. You can start with data standards developed for each specialty, but chances are you will also need to create a crosswalk for differing data standards. The important thing is to have each group use the assigned field for all the descriptive metadata, so that your crosswalks are accurate.
Thursday, April 8, 2010
Ah Metadata...
Thursday, April 1, 2010
How Do I Find the Right Words?
While I did not hear Dr. Judy Weidman explain how cataloging is an art and not a science, I believe that she is on the right track. As an image cataloger, I can relate to her assertions about the cataloging process particularly as it pertains to architecture: uncertainty is always present, design is non-linear and one is imposing an order. In many cases, I simply do not know enough.
Those of us who experienced the development, release, promulgation, and implementation of controlled vocabularies which are now used to tag works or documents that we are cataloging in order to facilitate straightforward retrieval, appreciate the consistency that authorized terminology brings to our cataloging. We avoid many of the pitfalls of databases replete with homographs, synonyms, and worse. Our choices when applying authorized terminology are supported by discipline specific user warrant thus providing welcome consistency both within our local databases and in the aggregate as they are shared within and beyond institutional borders. Our work became much easier when these authoritative vocabularies became widely available.
And, yet, there is so much that I as a visual resources’ cataloger fail to notice or simply don’t or can’t know. While my lack of knowledge might be discipline based, this is not necessarily the case. Images are used so ubiquitously by all sorts of people for many different reasons that it is impossible for a single cataloger facing a data input screen to capture it all. This is where folksonomy--also known as social tagging, social indexing or social classification--becomes useful. This collaborative creation and management of descriptive tags contributed by end users adds an important level of descriptive information that is both democratically based and current. It can provide information that I simply do not and cannot know. This is why the Library of Congress Flickr Project was so successful; this is why museums are building on the experience of Steve.Museum.
Yes, there is still a huge role for our favorite authorized vocabularies to play as we describe images. The Getty Art History Information Program vocabularies (AAT, ULAN, TGN, and soon CONA) along with the Thesaurus of Graphic Materials (TGM-I, TGM II), and the Library of Congress Authorities (NAF, SAF) will continue to provide a base level of consistency, timelessness, and stability to our descriptive practices. We need both approaches.