Showing posts with label descriptive metadata. Show all posts
Showing posts with label descriptive metadata. Show all posts

Sunday, December 30, 2012

PDF/A format

This is one person's blog about her experience as a DAS student in the Digital Archives Specialist certificate program that Society of American Archivists is offering.  I found this one post particularly helpful and wanted to share it. Also think the certificate program  looks quite solid.  Check it out
Diary of a DAS Student

Thursday, April 26, 2012

Metadata Deluxe / Top 10 List of Embedded Metadata Properties

Here is a great resource from the Embedded Metadata work group of the Visual Resources Association.  Its crosswalks all the common fields in most web and photo software tools: Check it out!

Metadata Deluxe / Top 10 List of Embedded Metadata Properties

Friday, October 21, 2011

NARAtions » National Archives Digitization Tools Now on GitHub

Just want to alert you to a great blog from the National Archives.  This post in particular discusses some digitisation tool that they are developing and making available. Check it out: NARAtions » National Archives Digitization Tools Now on GitHub

Tuesday, March 15, 2011

A New Image Site - Ookaboo

Recently I was contacted  by a Paul Houle about some broken links on the site Dig-mar.com which I maintain. I guess I might as well announce here that that site is no longer being maintained and will be closed completely this summer. I am focusing my energies on this blog and the Imageminders.net site of the ICCoop.. However, I am grateful to Paul for his reminder and would like to pass on the resource which he was attempting to post on that site. Ookaboo

His group  Ontology2, has created the Ookaboo website, which contains digital images that they claim  are either in the public domain and or under Creative Commons licensing terms. It appears to be a collection of harvested images from the web, in particular Wikimedia, where people place images that they wish to share with the world, so the free access is probably correct. 
The interesting part to me, and I think to many of you,  is how they are indexing the images , which is by means of the semantic web. 
Here is their description of what that means.
Images on Ookaboo are indexed by terms from the semantic web, the web of linked data. Although you're free to find images through the human interface, automated systems can quickly find and use images through the semantic API.
Ookaboo has two goals: (i) to dramatically improve the state of the art in image search for both humans and machines, and (ii) to construct a knowledge base about the world that people live in that can be used to help information systems better understand us.
Semantic Web, Linked Data
In the semantic web, we replace the imprecise words that we use everyday with precise terms defined by URLs. This is linked data because it creates a universal shared vocabulary.
For an example, in conventional image search, a person might use the word "jaguar" to search for
    •    the animal
    •    the automobile brand
    •    the Jacksonville Jaguars (NFL team)
    •    the game console from Atari
    •    ... and nearly 30 other things that are listed in Wikipedia.
Note in the cases above, there are pages in Wikipedia about each of the topics above: it's reasonable, therefore, that we could use these URLs as a shared vocabulary for referring to these things. However, we get some benefits when we use URLs that are linked to machine-readable pages, such as http://dbpedia.org/resource/Jaguar, or 
http://rdf.freebase.com/rdf/biology.itis.180593
Pages on Ookaboo are marked up with RDFa, a standard that lets semantic web tools extract machine readable information from the same pages that people view.
Named entities
Ookaboo is oriented around named entities, particularly 'concrete' things such as places, people and creative works. With current technology, it's more practical to create a taxonomy of things like "Manhattan", "Isaac Asimov" and "The Catcher In the Rye" than it is to tackle topics like "eating", "digestion" and "love". We believe that a comprehensive exploration of named entities will open pathways to an understanding of other terms, and hope to extend Ookaboo's capabilities as technology advances.
The above information is from their About us page, which I highly recommend you check out.

Oh and yes their images are pretty good too, especially for those interested in buildings. and other "concrete" things.

Monday, January 3, 2011

Embedded Metadata News

Fresh from the Visual Resource Association's (VRA) email list is this post from Greg Reser, which he has given me permission to repost here. Anyone with a Flickr account should check it out.

Embedded Metadata News

Like all of you, I spent my Holiday break reading about metadata, especially embedded metadata.  One blog post I found was from a Flickr developer in response to a complaint the social media companies "own" and "lock-in" their user's content.  The developer responded by saying that Flickr allows full access to all of the content users upload and create, including Descriptions and Tags, http://laughingmeme.org/2010/05/18/minimal-competence-data-access-data-ownership-and-sharecropping/ .  This is good news because it means that all that work you did tagging your photos is not limited to your Flickr page, you can download it to your computer or share it with other social media sites.  You enter the data once and then reuse it over and over.

Unfortunately, this data access comes via the Flickr API, meaning you have to be a programmer to get at it.  The good news is that some developers have created free or low cost applications to backup your Flickr content.  Best of all, they all have the cute missing "e" in their name.  For those of you have been looking for a way to get your metadata out of Flickr, you might try one of these apps.  The only limit I have found is that these apps do not export any custom metadata that was originally uploaded to Flickr.  So, if used a custom tool like the VRA Photoshop panel to add VRA metadata, it won't be downloaded.  It is possible to do this, just not with these apps.

Flickr Edit (Windows) - FREE (upload/download photos and EXIF and IPTC metadata metadata) http://www.flickr.com/services/apps/72157602367520002/

Downloadr (Windows) - FREE (download photos and metadata EXIF and IPTC metadata) http://www.flickr.com/services/apps/12400/

Bulkr (Windows, Mac) - FREE for basic (download photos), $24.95 for pro (download photos and metadata EXIF and IPTC metadata.  Also exports .txt file of metadata!) http://www.flickr.com/services/apps/72157622874451890/

You can find lots of other useful Flickr apps on the Flickr App Garden - http://www.flickr.com/services/

Besides retrieving metadata for your own photos, this ability could also be put to use for collecting images and metadata from groups.  For instance, you could have faculty and students upload and tag their own photos in Flickr, say of field work they did documenting architecture, and then you could download them and parse out the metadata into your institutional database.


Greg Reser

Arts Library
University of California, San Diego


Friday, November 12, 2010

University of Chicago to use Finding Aid for large scale Digitization Initiative

Kathleen E. Arthur, Head of Digitization at the Preservation Department of the University of Chicago Library recently announced their launch of a digitization initiative on the Digipres listserv. What I found particularly interesting is their use of finding aids, but these are not your mother's finding aids. As you can see by reading her post below, their use allows them to catalog at various levels of detail, through the use of EAD tags links placed at various places in the finding aid.

But Kathleen describes the whole project so much better, Read it for yourself.

11/11/10 to Digipres from Katheen Arthur:

The University of Chicago Library’s Special Collections Research Center has a launched an initiative for the digitization of archives and manuscript collections. The digital images are being made available via the online finding aid for each collection. This will recreate for the online user the experience of a researcher encountering the original materials in the SCRC Reading Room, with documents displayed as they are housed in each folder, and with description of the contents in the form of folder headings.

Individual, high-resolution images of each page will be permanently preserved in the Library's digital repository, and can be made available for publication or other research needs. Due to provisions of copyright laws, digitization efforts are currently focused on materials in the public domain, or those for which the University of Chicago holds copyright.

Collections with digitized content now available online include:

For those of you who are interested in the production details, the Large Scale Digitization Initiative is a collaborative effort involving the Special Collections Research, the Preservation Department, and the Digital Library Development Center. The initiative has been guided by definitions of, and requirements for, mass digitization provided by funding agencies such as the National Historical Publications and Records Commission. These guidelines stress expedited scanning workflows, without sacrifice of image quality, and with close attention to preservation concerns, and the use of existing descriptive metadata, such as that provided by a finding aid.

The collections are scanned by Preservation staff. The documents are scanned in color, in the order in which they are filed in each folder, and a .TIFF file is created for each page image. A naming scheme is used for the files which can be extended to other collections scanned as part of the initiative. The TIFF images from each folder in the physical collection are combined into PDFs for delivery. PDF was chosen as a delivery format because of its simplicity, stability and ubiquity. It is expected that the vast majority of users will have PDF viewers on their computers, and will be able to use them to enlarge, decrease, rotate, print, and otherwise easily view the images. Although the images are delivered as pdfs, the tifs of each page will be stored in the digital repository, and will be available if needed for other purposes.

Links to the digital files are added to the online finding aid by SCRC staff. The Encoded Archival Description (EAD) tags chosen allow links to be created at any level of description in the finding aid, from series, to folder, to item, and for multiple links to be attached to a particular description. DLDC staff updated the style sheets to allow display of the links in the finding aids database. DLDC is also hosting the digital files, which will be retained in and delivered from the digital repository.

The procedures developed for large scale digitization of archives and manuscript collections are simple and extensible. Future plans call for the delivery of digital audio files and, eventually, of born-digital content, and for full-text searching of digitized typescript documents, which can be made keyword searchable through optical character recognition (OCR).

Large Scale Digitization Team include: Eileen A. Ielmini, Kathleen Feeney, Daniel Meyer, Kathy Arthur, Karen Dirr, Charles Blair, and the student scanners in Preservation.

Friday, September 3, 2010

Metadata for Digital Content (MDC), Developing institution-wide policies and standards at the Library of Congress

Metadata for Digital Content (MDC), Developing institution-wide policies and standards at the Library of Congress
From the site
Metadata for Digital Content (MDC)
Developing institution-wide policies and standards at the Library of Congress

Over the years the Library of Congress' digital projects have generated many digital objects and these objects have been given various levels and types of descriptive metadata. The Library has assembled several use cases that require a more coordinated and standardized approach to the creation and management of this descriptive metadata. A few examples of use cases are:

* Geographic navigation of Library of Congress digital content
* Temporal navigation of Library of Congress digital content
* Exchange video and audio data with external services

Metadata of varying degrees of richness is necessary to support the use cases.

As part of this effort an institution-wide working group was established and is making the following available for use by any interested institutions:

* A master metadata element list with recommendations on best practices for populating the elements to provide more consistency of new metadata creation throughout the institution, support the Library of Congress metadata use cases, and point to areas where metadata remediation of current metadata might be beneficial.

Check it out

Wednesday, May 12, 2010

How to catalog Apples and Oranges?

So you are creating a digital asset system, which will serve a diverse user base. Your institution does not want to build a different management system for each user even if the marketing group may have different needs for digital assets than say the development or curator group. In an educational institution, there may be different subject areas or even in a design studio each individual may have a unique perspective. So, how to build an integrated system that supports all needs?

The key is flexible but clear rules
First, define a core group of fields that are to be filled all assets.  You will want that information which will aid in the management of the assets, such as:
  • Location of digital file
  • Title
  • Usage rights
  • Owner/creator of digital file
You will also want the fields that will aid users in cross collection searching, which must be derived from a study of your own users' needs when searching digital collection. For example in an art museum, all users would probably be interested in creator of original object, location of original object and possibly its provenance.  Thus a search on a particular object might not only turn up an image of that piece, but if doing a cross collection search, also promotional material about it and possibly informal images of it within an exhibit.  These fields would then be part of the core fields, which all groups would complete for their digital assets.

Secondly, define data standards before you start - i.e. decide what field will be used for which particular information.  You can start with data standards developed for each specialty, but chances are you will also need to create a crosswalk for differing data standards.  The important thing is to have each group use the assigned field for all the descriptive metadata, so that your crosswalks are accurate.

Thursday, April 8, 2010

Ah Metadata...


Jason Roy, at the Digital Collections Unit/Digital Library Development Lab, University of Minnesota had a word of encouragement for all those archivists out there while speaking at the VRA conference this year.  "You don't have to catalog on the item level." 
One of his archives received funding to scan an enormous collection for which they just had finding aids. For those not up on archival cataloging, it is customary to describe a collection by elements within a box, possibly a folder. So, you might have a folder of letters to Mr. Deere over a period of time. Alternatively, you might have a box of the Deere family memorabilia, estimated to cover 20 years of their lives.  The archivist will browse through said folder or box, looking for the item he desires.  The finding aid will give him a general idea where to best look.
Most digital collections have insisted that archivist must now throw some metadata at each item as it is converted to a digital format, so that an item is retrievable. 
Well, Jason, upon being confronted with the daunting task of converting this large collection to a digital format in the conventional manner, said no. Rather, they would place digital files in the appropriately labeled folder according to the box / folder information that existed already and would published a finding aid to direct the searcher to the appropriate "holder,"  which the searcher could then browse looking for the desired item.  No promises that it was there.  This is not ideal, but it is as good as the existing condition.  When you also throw in the technical ability to tag digital files with further information as researchers retrieve material, it means it will be a growing evolving catalog.  In addition, if you then embed the descriptive metadata that you have in each file, the digital file will always be identifiable without being bloated.
This approach would free up so many of our archives, if more special collections could accept it.  What do you think?

Thursday, April 1, 2010

How Do I Find the Right Words?

While I did not hear Dr. Judy Weidman explain how cataloging is an art and not a science, I believe that she is on the right track. As an image cataloger, I can relate to her assertions about the cataloging process particularly as it pertains to architecture: uncertainty is always present, design is non-linear and one is imposing an order. In many cases, I simply do not know enough.

Those of us who experienced the development, release, promulgation, and implementation of controlled vocabularies which are now used to tag works or documents that we are cataloging in order to facilitate straightforward retrieval, appreciate the consistency that authorized terminology brings to our cataloging. We avoid many of the pitfalls of databases replete with homographs, synonyms, and worse. Our choices when applying authorized terminology are supported by discipline specific user warrant thus providing welcome consistency both within our local databases and in the aggregate as they are shared within and beyond institutional borders. Our work became much easier when these authoritative vocabularies became widely available.

And, yet, there is so much that I as a visual resources’ cataloger fail to notice or simply don’t or can’t know. While my lack of knowledge might be discipline based, this is not necessarily the case. Images are used so ubiquitously by all sorts of people for many different reasons that it is impossible for a single cataloger facing a data input screen to capture it all. This is where folksonomy--also known as social tagging, social indexing or social classification--becomes useful. This collaborative creation and management of descriptive tags contributed by end users adds an important level of descriptive information that is both democratically based and current. It can provide information that I simply do not and cannot know. This is why the Library of Congress Flickr Project was so successful; this is why museums are building on the experience of Steve.Museum.

Yes, there is still a huge role for our favorite authorized vocabularies to play as we describe images. The Getty Art History Information Program vocabularies (AAT, ULAN, TGN, and soon CONA) along with the Thesaurus of Graphic Materials (TGM-I, TGM II), and the Library of Congress Authorities (NAF, SAF) will continue to provide a base level of consistency, timelessness, and stability to our descriptive practices. We need both approaches.

Monday, March 8, 2010

The Art of Cataloging

We all have heard the cataloging is an art not a science and last month at the Northern California VRA meeting, I heard Dr. Judy Weidman of San Jose State's SLIPS explain exactly how.  She is in the process of developing a thesis on "Design Theory: Creating local vocabularies for images."  She is basing much of her design theory on architectural writings, because as she said, they seem the most self-analytical, which is one way to put it.   I found her approach very refreshing. It may just be that, as a former architect, her approach is familiar to me, but also I was amazed at how it did work to guide my current efforts.  She had several premises: uncertainty is always present, Design is non linear and you are imposing an order.  The uncertainty comes from the fact that there is no absolutely right way to describe anything. It is non-linear because one choice influences another. Nor is it organic or meant to be, the cataloger is creating the structure, not revealing it.

The real crux of her argument for me was that one must define the problem before solving it, Obvious I know, but not always done. Many strive to fully and accurately describe an object, when maybe we should be striving to meaningfully describe it. To do that, we must define the audience that is looking for the meaning. As many of us know that can sometimes be the hardest description to accurately create.