Showing posts with label transcript. Show all posts
Showing posts with label transcript. Show all posts

Sunday, 22 February 2015

A Transcription Toolbox

In the early 1980s my grandmother taught me to crochet.  We worked together on a set of lacey mats for my dressing table.  Gran did the larger central mat while I worked on a smaller mat with matching pattern.  To facilitate the creative endeavor she wrote out the pattern for me to follow.
Crochet pattern transcribed by my gran (Mabel Adams, nee Coulson 1910-1991), the crocheted mat I made from the pattern, and the crochet hook used

For clarity, legibility and accuracy, Gran wrote in upper case, but the abbreviations and terminology would not be meaningful to someone who does not know crochet.  Gran's normal cursive (joined up) handwriting was a different style to mine and I remember thinking that some of her letters looked slightly odd.

Learning to read old handwriting, understanding why a document was created and knowing about its subject are key to genealogical source analysis. 

Why transcribe?

Handwriting recognition is a rapidly developing technology, but it may tempt us to be too lazy to learn about old handwriting.  Optical character recognition (OCR) was hailed as the solution for printed text, but the results are often not very accurate, particularly with poor quality print and older fonts.  These technologies are useful, but I think the "no more indexing" claim by the runners-up in the RootsTech 2015 Innovator Showdown , ArgusSearch, is misguided.

Apart from making a digital copy, transcribing documents fully makes you examine the details closely, which helps you better understand the contents.

Palaeography Lessons

Palaeography (also spelled paleography) is the study of old handwriting.  Fortunately there is a range of online tutorials that can help us learn to read scripts used in previous centuries.

The UK National Archives tutorials are an excellent practical introduction, with exercises for you to try.  Have some fun with the ducking stool gameScottish Handwriting.com from the National Records of Scotland offers practical exercises and a weekly poser.  English Handwriting 1500-1700 provides a good background information and transcription conventions.

These resources cover the period from 1500 to 1800.  For genealogists, learning to read secretary hand, a script widely used for administrative and legal purposes is important.

As handwriting developed from earlier scripts, an understanding of those scripts is helpful.  Nottingham University's Medieval Handwriting provides interactive exercises and information on understanding medieval documents.  Further early modern and medieval material is available at Palindex. DigiPal is a fabulous resource for very early scripts from Anglo-Saxon and Norman times (1000-1100).

Latin was both an international and legal language.  In England it was used in legal documents up to 1733.  Legal Latin is often heavily abbreviated which makes deciphering it a challenge.  A knowledge of the law at the time the document was written is greatly helpful.  Fortunately, legal documents tend to follow a pattern, which helps us predict what it might contain.  Once again the UK National Archives has good introduction to Latin, and Latin PalaeographyWhittaker's Words is an online Latin dictionary.

Decoding the script and abbreviations, and knowing what to expect are all pre-requisites for a quality transcript, just like the crochet pattern needed knowledge and skill to produce a finished article.

Many of the resources listed here include interactive practice.  Practice really does make perfect transcriptions.

Transcription Tools

Armed with knowledge gained from the resources above, you are ready to start transcribing.  For access reasons, it is common for genealogists to work from a digital image or a photocopy of the original manuscript.  Enlargement and zooming in on detail is an advantage of a good quality digital image.  A large screen or two screens can help make the image and transcript easier to see and read.  Before you rush to buy a second screen, check the ports on your computer and TV.  I connect to my TV using a HDMI cable.  You can use image program like Irfanview or Photoshop to manipulate images and a word processor, such as Word.  Some people find manipulating two programs in different windows clumsy.

Transcript 2.5 puts the image and transcript in a single window and provides basic image manipulation and text formatting tools.  Some people find it suits them very well, but I did not find it much easier to use.  Try both approaches and find what suits you.  For me, the deal breaker is that the formatting is not completely compatible with Word.

A plain or formatted text transcript is a good start, but I want a more sophisticated transcription tool:
  • Annotations note layout, corrections, obscured text and non-text parts of the manuscript.
  • Semantic mark-up, tags that tell the computer what the text means, enable further analysis.
  • Support for non-standard characters, such as the thorn and yogh that appear in 1700s manuscripts.

A more detailed explanation of why I want these features is explained in the worked example of a property document.

T-PEN transcription interface


T-PEN, Transcription for Palaeographical and Editorial Notation is an academic project intended for digital humanities scholars.  The web based tool includes image analysis functions that highlight lines of text.  The user-friendly interface presents each line of text for transcription.  Annotations, special characters and XML encoding are supported.  XML encoding is an exciting feature because it offers the potential of semantic mark-up.  This is the kind of transcription tool I would like to see developed and adapted for genealogy.

Tuesday, 22 April 2014

Criteria for Assessing the Quality of Genealogy Websites and Online Data

Academic researchers, commercial vendors and volunteer interest groups have produced a vast array of online resources useful to family historians and genealogists.  The quality of the websites and data contained in them varies hugely.  For this discussion, I will focus on websites that offer access to digital copies of original records.

Quality has nothing to do with the total number of records, or the number of collections or data sets.  Quality is unrelated to the cost of a website subscription or motivation of the provider.

Without documentation of all the processing of records and information in them a researcher cannot asses the reliability of records. We can’t change the imperfect state that the original records come to us in. Archivists work hard to preserve both the records themselves and the context of their creation and use, but online presentation is often performed by other parties. Digitisation, indexing, and search are just a few of the processes that happen before an online version of the record is presented. Presentation can have profound influence on how records are perceived and the conclusions drawn from them. Consequently, transparency is an ethical obligation.

What are the most important website features? How can they be assessed?

The following, in order of importance, are essential:
  1. Catalogue
  2. Transcript quality
  3. Search facilities
  4. Browsing facilities
  5. Record quality
The quality of other features are also important, and a bonus if included.  Examples include analytical tools, user data (e.g. family trees, imported sources, research notes etc.), collaborative tools and social networks. But for now, I will discuss the basic five points above.

Collections with different histories or characteristics should be assessed separately. Only the catalogue can be assessed across a whole website.

Catalogue First

Yes, I really do mean that the catalogue is more important than anything else.

Genealogists use archival material, whether in the form of original records or some kind of derivative.  Genealogy websites are really a digital archive of such materials, so a genealogy website’s catalogue should share many of the features of an archival catalogue.

In her blog post The Value of Archival Description, Considered, archivist Maureen Callaghan recognizes researcher’s needs:
“getting to understand who created records, why they were created, and what they provide evidence of – really gets to the nature of research. These are the questions that historians and journalists and lawyers and all of the communities that use our collections ask – they don’t just see artifacts, they see evidence that can help them make a principled argument about what happened in the past. They want to know about reliability, authenticity, chain of custody, gaps, absences and silences.”

So, a catalogue is more than just a list of collections. Such a list might be the starting point for creating a catalogue, but falls well short of the sophisticated database that comprises an archival catalogue. It contains information about the collections, so serves a quite different purpose to search and browsing facilities.

A good catalogue answers questions about the website’s collections with no fuss:
  1. Is the catalogue complete, including collections not yet digitised and indexed with a timescale of expected online availability? This information allows the researcher to make informed decisions using the database or seeking the records elsewhere.
  2. What record collections does it contain? You want to know that relevant records are included before paying a subscription or spending precious time searching for records, don’t you?
  3. Where did the collections come from? Typically records come from originals in an archive, or a publication. The barest minimum information for archival material is the archive and the archive reference, and for published information, the bibliographic reference. That allows the researcher to check the archive’s or bibliographic catalogues.
  4. How do the collections relate to one another? Logical groupings of record sets by record type reflect original function of the records, whilst groupings by creator reflect the history or provenance of the records. Both are important for understanding how the records can be used. Were several types of record created by a particular process e.g. collection of taxes involved assessment of liability, record of payments and penalties for late or non-payment.
  5. What is the structure of the data set?  How is the record set arranged? Is it by date, person or something else?
  6. Is each collection or record set complete?
  7. What is the extent of each collection and record set?  How many sub-sets, how many records in each?
  8. Does the catalogue entry describe the records? Is a brief history of the original records creation and provenance included? Were the digital records an image of the original, or derived from a microfilm or transcript?
  9. What information do the records typically contain?
  10. Is a scholarly work on the record type referenced, or a critique on the strengths and weaknesses of the records included?

Transcript Quality

Transcription transforms manuscript and typescript documents into computer readable text, essential for creating searchable records. In evaluating the quality of computerized records consider if the website documents the following:
  1. The completeness of the transcript.  A complete transcript captures the most information so is far more useful than an abstract or an index. 
  2. Accuracy of transcription is influenced by how it was produced.  Optical character recognition (OCR) is commonly used for typescript.  Human data entry of manuscript or handwritten material depends on palaeographic and keyboard skills.  Typically, OCR and unskilled data entry yield less accurate transcripts.
  3. Checking procedures should detect obvious gobbledy-gook, and common OCR and data entry errors.  Double data entry produces a con-census interpretation, but may not avoid common reading errors.
  4. Have error rates been assessed?

Search Facilities

Good search rests on an accurate transcript, not algorithms or user added ‘corrections’.  Repeatable search, essential for confidence in the validity of results, requires a complete data set and stable search methods.  Search is not a simple operation, so inexperienced users need coaching and encouragement, not dumbed-down, limited functionality.  Consider the website’s documentation and functionality of the following:
  1. Targeted search on individual data sets and collections as default. Choosing which collection or collections to search first is much more efficient than filtering out irrelevant collections.
  2. Search on all data items in the record.  This requires a complete transcript.
  3. Full text search, the ability to search everywhere in the record, also requires a complete transcript.
  4. Complex search, the ability to specify ‘AND’, ‘OR’ and other operators.
  5. Name matching algorithm choice.  Examples include soundex and metaphone, which perform phonetic matching for English-language names, and Daitch-Mokotoff, which is adapted for Slavic and German spellings of Jewish names.
  6. Date ranges. Can start and end dates be specified, or a central dates with accuracy?
  7. Are place searches restricted to place names? Is a proximity search based on distance included?
  8. Wildcards, replacement characters in the search term that stand in for unknown possibilities e.g. Sm*th returns Smith, Smyth.
  9. Separation of transcribed values and interpreted values with options to search on either or a combination. For example, the abbreviation ‘Wm’ or Latinised ‘Gulielmus’ can be interpreted as William. Standardised interpretations are known to librarians and archivists as authority control http://en.wikipedia.org/wiki/Authority_control . User added ‘corrections’ are another kind of interpreted value.
  10. Filtering of search results.
  11. Optional ‘sticky’ settings and well-chosen default settings.
  12. Result presentation.  Is it simple, clear and contain the information important to you?  Does it include the search terms used?
  13. Result sorting on data fields chosen by the user. What does ‘relevance’ mean?
  14. Result export in a variety of formats, ready for use by with software tools of your choice.
  15. Logged searches that document research activity.

Browsing Facilities

Browsing is a tool for examining records in the context of the record set.  It should replicate the experience of turning the pages of the original.  The order of digital images should exactly follow the order of the original pages.  The structure of the record set and relationships between individual records contain subtle information about the creation and use of the original.

Browsing does not replace search. When used as a last resort when search fails, it is an indicator of poor search or transcription.

Record Quality

Original records are typically presented as digital images. Genealogists use the most original source available so they can be confident that the information is as reliable as possible. A digital image is not the same as the original, but can come acceptably close, provided that:
  1. Good image quality that is legible.  Sharp focus, resolution, colour accuracy, and contrast all contribute to legibility. Digital image file types vary in the degree of data compression, which influences image quality.
  2. Information that identifies the record portrayed included in the image file. That means all the information that you want in your citation, such as the archive and archive reference of the original, page number, record identifier, person of interest etc. A meaningful file name is helpful, but enough detail makes an unreasonably long name, as is human readable text added to the image. Potentially most useful is embedded citation information in the image file metadata, which is computer readable.
  3. Technical camera or scanner metadata provides provenance of the image, including whether it has been modified.

Your Challenge – Review one data set

There is a lot to consider in assessing the quality of genealogy websites and the data they contain. Of course, we want all the features mentioned above in a user-friendly package, but I think there is quite enough to start with above. Have I omitted anything vital? Do you agree with these criteria?

Before we can hold genealogy data suppliers accountable, we need to fairly assess whether what they offer is of sufficiently good quality for our purposes.  What constitutes ‘fit for purpose’ is open to debate.  I think genealogy data consumers would benefit from setting expectations and demanding quality, and that suppliers would benefit greatly from carefully considered feedback.

In the interest of collaboration between suppliers and consumers, I challenge you to review one data set using these criteria.