Showing posts with label catalogue. Show all posts
Showing posts with label catalogue. Show all posts

Sunday, 22 November 2015

From the Top - Explore Archives


This week the annual Explore Your Archives campaign launched here in the UK. Remember only a tiny fraction of historical materials have been digitised and made available online, so archives still hold many treasures genealogists need.  Learning to use an archive is really important.

I am prepared to bet that most of us were taught how to use a library sometime in childhood, but we learn to use archives as adults by trial and error. Libraries organise their books and other materials and have catalogues to help you find what you want. Archives also organise and catalogue their holdings, but do so differently.

Archives are organised using a hierarchy with several levels. Cataloguing archives, in four very easy steps illustrates the four main levels with rather natty photos. The levels are:
  • Fonds or Collection
  • Series
  • File
  • Item
Genealogists often focus on the single item, and forget the context in which it was created, used and preserved. Much valuable information about an item is in the higher catalogue levels, so explore them all.

Archivists start at the top level and don't always have the resources to fully describe lower levels.  The best way of discovering what an archive holds is to look at the fonds level. The National Archives (TNA, the UK one) has over 400 fonds, which is somewhat overwhelming. According to the published guide 'Tracing your Ancestors in the National Archives' these are the fonds most commonly used by family historians:


I would like to see a summary like this one posted on every archive's website and displayed prominently in every search room.  I am sure UK researchers are familiar with some of the contents of the fonds HO, RG and WO. Have you looked at other fonds, and other series within fonds?  Take a close look at RG 101.

Sadly, online archive catalogues do not make examining fonds level entries as easy as it should be. To display all fonds in Discovery, TNA's catalogue, use the Advanced Search. A search term must be entered so I used the wildcard *, meaning everything.  Then I scrolled down the page and selected 'The National Archives' under Held by, which brought up further options. I scrolled a long way further and selected 'Department' under Catalogue Levels.  TNA confusingly uses the terms Department for Fonds and Piece for File. Once you have identified a Fond of interest, you can search using the reference.


Now it is your turn. Have fun exploring an archive!

Reference
Bevan, Amanda. 2006. Tracing your Ancestors in the National Archives. The National Archives: Richmond. pp 4-7.

Thursday, 22 May 2014

There Be Dragons – Finding Tithe Maps for England and Wales

Tithe maps are often the earliest detailed maps of parishes in England and Wales. They are important for genealogists, as well as historians of all kinds, cartographers, and other disciplines.

Tithes were a tax levied on land owners by the established churches, the Church of England and the Church of Wales in their respective countries. The property right to collect tithes, traditionally one tenth of agricultural produce, was a claimed by the church for the support of the clergy. In 1836, tithes were commuted (changed) to money payments rather than payments in kind. The Tithe Act 1836 set out the transition process, and established the Tithe Commission to oversee it. For each parish, the total revenue due was apportioned between the landowners, taking into account the area and productivity of the land. As the area of land was a factor in the calculation of the monetary amount due by each landowner, a survey was required. The tithe commission used the maps to check the accuracy of the survey, hence the fairness of the apportionment.

The Tithe Apportionments that accompany each map, record the owner, occupier, area, land use, tithable value and plot number for each land parcel. The plot number corresponds directly with the map, so people’s land ownership and occupation is precisely located. This is a genealogical treasure chest!

Archive Catalogues locate original Tithe Maps

Three copies of each map were produced for each of the parties involved: the Tithe Commission, the diocese (the bishop oversaw the clergy), and the parish (the rector was usually the main tithe owner).

The Tithe Commission copies at The National Archives (TNA), a collection of 11,785 tithe maps, were described in Kain and Oliver’s seminal work, The Tithe Maps of England and Wales.1 Many of the descriptions have been incorporated into the TNA catalogue.

TNA catalogue entry for tithe commissioner's copy of Llanychan tithe map

The diocese copies for Wales are at The National Library of Wales (NLW), and the catalogue descriptions benefit from Davies’ work, The tithe maps of Wales.2

NLW catalogue entry for the diocese copy of Llanychan tithe map & apportionment
 The diocese copies for England and parish copies are likely to be found in the relevant county archives, which vary in the availability of online catalogues and detail of description. Denbighshire Archives has an index that includes the third parish copy of tithe map for Llanychan, the example shown in the TNA and NLW catalogues above. An example from England, the parish of Claverley, Shropshire, is in the diocese of Hereford. The diocese copy is at the Herefordshire Archives, but is not included in an online catalogue, so I had to enquire to locate it.

Shropshire Archives’ catalogue includes copies of the tithe apportionment for Claverley, but no map. Enquiries have revealed that when the church deposited its records in 1955, no map was found. In Claverley, there were several tithe owners other than the parish incumbent. Consequently, the ‘parish’ copy of the map may have been kept by one of them and never deposited in the archive.

Which of the three tithe maps for any particular parish was the original?

In the case of Llanychan, it is hard to tell, as both the commissioner’s and diocese copies are manuscripts, or hand drawn. The diocese copy was apparently drawn from a new survey, so was current in 1838.

In the case of Claverley, the commissioner’s copy is a lithograph, a kind of printed copy, dated 1840, and the diocese copy bears a date of 1842. So, neither is the original.

Archives often make photocopies of tithe maps to prevent damage from handling the huge originals, which may be measured meters. Such copies may be reduced in size so do not faithfully reproduce scale and detail.

Scanning and photographing maps large in both scale and physical size, in a way that preserves both the detail and cartographic accuracy, is challenging. The resolution of this problem underpins the quality of digital copies, especially interactive online versions.

Online Editions and Peeling the Onion with GIS

Traditional maps pack lots of information into a flat piece of paper. There are multiple types of features like the roads, rivers and water bodies, buildings, and land parcels depicted on tithe maps. Think of these types as the layers of an onion.


The land ownership and occupation information could have been written on the map, but that would have made it difficult to use for tax purposes. Instead the plot number provides the key for the apportionment data, which is presented as a separate document. When the tithe maps were produced, this was the only practical option, but modern map making tools can handle this with ease.

Geographical Information Systems (GIS) have revolutionized map production from the 1960s right up to today. This is also the technology behind familiar online mapping like Google Maps and Bing. GIS editions of tithe maps online include:

Cheshire
Norfolk
East Sussex
Worcestershire
Leeds area, West Yorkshire
University of Portsmouth, a selection of parishes

Try them out. Do you prefer outlines of land parcels overlaid on a current map, or viewing old and new side-by-side? Which maps are easiest to use? Did you experience performance issues like slow responses? Can you pinpoint great-grand-daddy’s tiny hovel?

Each map has some advantages and drawbacks. These websites are limited implementations of the technology. If you know of website that does really impressive mapping, please share in the comments. I am impressed with the United States website Atlas of Historical County Boundaries which even offers the GIS data for download. I would like to see something similar for the tithe maps.

In February 2014, the UK-based genealogy subscription website The Genealogist announced that they will offer the TNA copies of tithe apportionments and maps online. Some apportionments are available now, and the maps are due to launch in 2015. It is not clear how the maps and apportionments will be linked. The time scale seems too short for the production of GIS editions.

According to the minutes (Item 6.1) of the TNA’s User Advisory Group meeting on 18 March 2014, in answer to questions in about the tithe map digitization, Commercial Director Mary Gledhill stated that
“the digitisation being carried out by S and N is targeted towards a genealogist market. They are paying for the work to be carried out, and make the decision on how the work is produced”. 
S and N are the company that runs The Genealogist. I am perturbed by the TNA’s apparent lack of engagement with the digitisation of the tithe maps. Done well, it could be an important research resource for all disciplines. Family and house historians need a high quality, and cartographically accurate rendition. Then you stand a real chance of locating the small dwelling of a peasant ancestor.

My comments on the production of GIS versions of tithe maps is based on personal experience of digitising tithe maps gained during research for my Master’s degree dissertation.3

References

1 Kain, Roger & Oliver, Richard (1995) The tithe maps of England and Wales : a cartographic analysis and county-by-county catalogue, Cambridge: Cambridge University Press.
2 Davies, Robert. (1999) The tithe maps of Wales :a guide to the tithe maps and apportionments of Wales in the National Library of Wales Aberystwyth: National Library of Wales.
3 Adams, Susan Yvonne (2012) To what extent can Cartographic, Land and Genealogical data be combined to establish Land Ownership in England and can a Geographical Information System (GIS) tie it all together? unpublished MSc thesis, University of Strathclyde, Glasgow.

Tuesday, 22 April 2014

Criteria for Assessing the Quality of Genealogy Websites and Online Data

Academic researchers, commercial vendors and volunteer interest groups have produced a vast array of online resources useful to family historians and genealogists.  The quality of the websites and data contained in them varies hugely.  For this discussion, I will focus on websites that offer access to digital copies of original records.

Quality has nothing to do with the total number of records, or the number of collections or data sets.  Quality is unrelated to the cost of a website subscription or motivation of the provider.

Without documentation of all the processing of records and information in them a researcher cannot asses the reliability of records. We can’t change the imperfect state that the original records come to us in. Archivists work hard to preserve both the records themselves and the context of their creation and use, but online presentation is often performed by other parties. Digitisation, indexing, and search are just a few of the processes that happen before an online version of the record is presented. Presentation can have profound influence on how records are perceived and the conclusions drawn from them. Consequently, transparency is an ethical obligation.

What are the most important website features? How can they be assessed?

The following, in order of importance, are essential:
  1. Catalogue
  2. Transcript quality
  3. Search facilities
  4. Browsing facilities
  5. Record quality
The quality of other features are also important, and a bonus if included.  Examples include analytical tools, user data (e.g. family trees, imported sources, research notes etc.), collaborative tools and social networks. But for now, I will discuss the basic five points above.

Collections with different histories or characteristics should be assessed separately. Only the catalogue can be assessed across a whole website.

Catalogue First

Yes, I really do mean that the catalogue is more important than anything else.

Genealogists use archival material, whether in the form of original records or some kind of derivative.  Genealogy websites are really a digital archive of such materials, so a genealogy website’s catalogue should share many of the features of an archival catalogue.

In her blog post The Value of Archival Description, Considered, archivist Maureen Callaghan recognizes researcher’s needs:
“getting to understand who created records, why they were created, and what they provide evidence of – really gets to the nature of research. These are the questions that historians and journalists and lawyers and all of the communities that use our collections ask – they don’t just see artifacts, they see evidence that can help them make a principled argument about what happened in the past. They want to know about reliability, authenticity, chain of custody, gaps, absences and silences.”

So, a catalogue is more than just a list of collections. Such a list might be the starting point for creating a catalogue, but falls well short of the sophisticated database that comprises an archival catalogue. It contains information about the collections, so serves a quite different purpose to search and browsing facilities.

A good catalogue answers questions about the website’s collections with no fuss:
  1. Is the catalogue complete, including collections not yet digitised and indexed with a timescale of expected online availability? This information allows the researcher to make informed decisions using the database or seeking the records elsewhere.
  2. What record collections does it contain? You want to know that relevant records are included before paying a subscription or spending precious time searching for records, don’t you?
  3. Where did the collections come from? Typically records come from originals in an archive, or a publication. The barest minimum information for archival material is the archive and the archive reference, and for published information, the bibliographic reference. That allows the researcher to check the archive’s or bibliographic catalogues.
  4. How do the collections relate to one another? Logical groupings of record sets by record type reflect original function of the records, whilst groupings by creator reflect the history or provenance of the records. Both are important for understanding how the records can be used. Were several types of record created by a particular process e.g. collection of taxes involved assessment of liability, record of payments and penalties for late or non-payment.
  5. What is the structure of the data set?  How is the record set arranged? Is it by date, person or something else?
  6. Is each collection or record set complete?
  7. What is the extent of each collection and record set?  How many sub-sets, how many records in each?
  8. Does the catalogue entry describe the records? Is a brief history of the original records creation and provenance included? Were the digital records an image of the original, or derived from a microfilm or transcript?
  9. What information do the records typically contain?
  10. Is a scholarly work on the record type referenced, or a critique on the strengths and weaknesses of the records included?

Transcript Quality

Transcription transforms manuscript and typescript documents into computer readable text, essential for creating searchable records. In evaluating the quality of computerized records consider if the website documents the following:
  1. The completeness of the transcript.  A complete transcript captures the most information so is far more useful than an abstract or an index. 
  2. Accuracy of transcription is influenced by how it was produced.  Optical character recognition (OCR) is commonly used for typescript.  Human data entry of manuscript or handwritten material depends on palaeographic and keyboard skills.  Typically, OCR and unskilled data entry yield less accurate transcripts.
  3. Checking procedures should detect obvious gobbledy-gook, and common OCR and data entry errors.  Double data entry produces a con-census interpretation, but may not avoid common reading errors.
  4. Have error rates been assessed?

Search Facilities

Good search rests on an accurate transcript, not algorithms or user added ‘corrections’.  Repeatable search, essential for confidence in the validity of results, requires a complete data set and stable search methods.  Search is not a simple operation, so inexperienced users need coaching and encouragement, not dumbed-down, limited functionality.  Consider the website’s documentation and functionality of the following:
  1. Targeted search on individual data sets and collections as default. Choosing which collection or collections to search first is much more efficient than filtering out irrelevant collections.
  2. Search on all data items in the record.  This requires a complete transcript.
  3. Full text search, the ability to search everywhere in the record, also requires a complete transcript.
  4. Complex search, the ability to specify ‘AND’, ‘OR’ and other operators.
  5. Name matching algorithm choice.  Examples include soundex and metaphone, which perform phonetic matching for English-language names, and Daitch-Mokotoff, which is adapted for Slavic and German spellings of Jewish names.
  6. Date ranges. Can start and end dates be specified, or a central dates with accuracy?
  7. Are place searches restricted to place names? Is a proximity search based on distance included?
  8. Wildcards, replacement characters in the search term that stand in for unknown possibilities e.g. Sm*th returns Smith, Smyth.
  9. Separation of transcribed values and interpreted values with options to search on either or a combination. For example, the abbreviation ‘Wm’ or Latinised ‘Gulielmus’ can be interpreted as William. Standardised interpretations are known to librarians and archivists as authority control http://en.wikipedia.org/wiki/Authority_control . User added ‘corrections’ are another kind of interpreted value.
  10. Filtering of search results.
  11. Optional ‘sticky’ settings and well-chosen default settings.
  12. Result presentation.  Is it simple, clear and contain the information important to you?  Does it include the search terms used?
  13. Result sorting on data fields chosen by the user. What does ‘relevance’ mean?
  14. Result export in a variety of formats, ready for use by with software tools of your choice.
  15. Logged searches that document research activity.

Browsing Facilities

Browsing is a tool for examining records in the context of the record set.  It should replicate the experience of turning the pages of the original.  The order of digital images should exactly follow the order of the original pages.  The structure of the record set and relationships between individual records contain subtle information about the creation and use of the original.

Browsing does not replace search. When used as a last resort when search fails, it is an indicator of poor search or transcription.

Record Quality

Original records are typically presented as digital images. Genealogists use the most original source available so they can be confident that the information is as reliable as possible. A digital image is not the same as the original, but can come acceptably close, provided that:
  1. Good image quality that is legible.  Sharp focus, resolution, colour accuracy, and contrast all contribute to legibility. Digital image file types vary in the degree of data compression, which influences image quality.
  2. Information that identifies the record portrayed included in the image file. That means all the information that you want in your citation, such as the archive and archive reference of the original, page number, record identifier, person of interest etc. A meaningful file name is helpful, but enough detail makes an unreasonably long name, as is human readable text added to the image. Potentially most useful is embedded citation information in the image file metadata, which is computer readable.
  3. Technical camera or scanner metadata provides provenance of the image, including whether it has been modified.

Your Challenge – Review one data set

There is a lot to consider in assessing the quality of genealogy websites and the data they contain. Of course, we want all the features mentioned above in a user-friendly package, but I think there is quite enough to start with above. Have I omitted anything vital? Do you agree with these criteria?

Before we can hold genealogy data suppliers accountable, we need to fairly assess whether what they offer is of sufficiently good quality for our purposes.  What constitutes ‘fit for purpose’ is open to debate.  I think genealogy data consumers would benefit from setting expectations and demanding quality, and that suppliers would benefit greatly from carefully considered feedback.

In the interest of collaboration between suppliers and consumers, I challenge you to review one data set using these criteria.