Showing posts with label Challenge. Show all posts
Showing posts with label Challenge. Show all posts

Monday, 2 May 2016

Worldwide Genimates

I'm a survivor!
Last month I blogged about my rash decision to join the 2016 Blogging from A to Z Challenge. I am pleased to report that I crossed the finish line on time last Saturday with around 1300 other bloggers.

It was a hard slog that ate into my precious time but it was worth the effort. One of the unexpected benefits was that it introduced me to a number of genealogy and history bloggers from around the world. I even discovered that one of my Surname Study contacts was a prolific blogger.

The result is that the number of feeds in my RSS reader has grown significantly.

Reading others' blog posts confirmed my bias towards reminiscing and stories hence the two standout blogs for me were from two Australians, a new contact and a social media mate. If you fancy trips down memory lane do visit Linda and Maureen.

While I saw many old friends in the challenge list some of the new to me blogs I discovered were:

GenWestUK - Ros - England
History Roundabout - ? - England
Molly's Canopy - Molly - USA
My Genealogy Challenges - Dianne - Canada
Old Scottish Genealogy and Blog - Penny and Fergus - Scotland
The Past Whispers - USA
Roots and Stuff - Mary - USA
Southern Graves - Stephanie - USA
Treetrack'n - Mary
The Writing Desk - Ros - England

But wait, there's more - Pauleen has compiled a list of geneabloggers who participated in the challenge.

Tuesday, 22 April 2014

Criteria for Assessing the Quality of Genealogy Websites and Online Data

Academic researchers, commercial vendors and volunteer interest groups have produced a vast array of online resources useful to family historians and genealogists.  The quality of the websites and data contained in them varies hugely.  For this discussion, I will focus on websites that offer access to digital copies of original records.

Quality has nothing to do with the total number of records, or the number of collections or data sets.  Quality is unrelated to the cost of a website subscription or motivation of the provider.

Without documentation of all the processing of records and information in them a researcher cannot asses the reliability of records. We can’t change the imperfect state that the original records come to us in. Archivists work hard to preserve both the records themselves and the context of their creation and use, but online presentation is often performed by other parties. Digitisation, indexing, and search are just a few of the processes that happen before an online version of the record is presented. Presentation can have profound influence on how records are perceived and the conclusions drawn from them. Consequently, transparency is an ethical obligation.

What are the most important website features? How can they be assessed?

The following, in order of importance, are essential:
  1. Catalogue
  2. Transcript quality
  3. Search facilities
  4. Browsing facilities
  5. Record quality
The quality of other features are also important, and a bonus if included.  Examples include analytical tools, user data (e.g. family trees, imported sources, research notes etc.), collaborative tools and social networks. But for now, I will discuss the basic five points above.

Collections with different histories or characteristics should be assessed separately. Only the catalogue can be assessed across a whole website.

Catalogue First

Yes, I really do mean that the catalogue is more important than anything else.

Genealogists use archival material, whether in the form of original records or some kind of derivative.  Genealogy websites are really a digital archive of such materials, so a genealogy website’s catalogue should share many of the features of an archival catalogue.

In her blog post The Value of Archival Description, Considered, archivist Maureen Callaghan recognizes researcher’s needs:
“getting to understand who created records, why they were created, and what they provide evidence of – really gets to the nature of research. These are the questions that historians and journalists and lawyers and all of the communities that use our collections ask – they don’t just see artifacts, they see evidence that can help them make a principled argument about what happened in the past. They want to know about reliability, authenticity, chain of custody, gaps, absences and silences.”

So, a catalogue is more than just a list of collections. Such a list might be the starting point for creating a catalogue, but falls well short of the sophisticated database that comprises an archival catalogue. It contains information about the collections, so serves a quite different purpose to search and browsing facilities.

A good catalogue answers questions about the website’s collections with no fuss:
  1. Is the catalogue complete, including collections not yet digitised and indexed with a timescale of expected online availability? This information allows the researcher to make informed decisions using the database or seeking the records elsewhere.
  2. What record collections does it contain? You want to know that relevant records are included before paying a subscription or spending precious time searching for records, don’t you?
  3. Where did the collections come from? Typically records come from originals in an archive, or a publication. The barest minimum information for archival material is the archive and the archive reference, and for published information, the bibliographic reference. That allows the researcher to check the archive’s or bibliographic catalogues.
  4. How do the collections relate to one another? Logical groupings of record sets by record type reflect original function of the records, whilst groupings by creator reflect the history or provenance of the records. Both are important for understanding how the records can be used. Were several types of record created by a particular process e.g. collection of taxes involved assessment of liability, record of payments and penalties for late or non-payment.
  5. What is the structure of the data set?  How is the record set arranged? Is it by date, person or something else?
  6. Is each collection or record set complete?
  7. What is the extent of each collection and record set?  How many sub-sets, how many records in each?
  8. Does the catalogue entry describe the records? Is a brief history of the original records creation and provenance included? Were the digital records an image of the original, or derived from a microfilm or transcript?
  9. What information do the records typically contain?
  10. Is a scholarly work on the record type referenced, or a critique on the strengths and weaknesses of the records included?

Transcript Quality

Transcription transforms manuscript and typescript documents into computer readable text, essential for creating searchable records. In evaluating the quality of computerized records consider if the website documents the following:
  1. The completeness of the transcript.  A complete transcript captures the most information so is far more useful than an abstract or an index. 
  2. Accuracy of transcription is influenced by how it was produced.  Optical character recognition (OCR) is commonly used for typescript.  Human data entry of manuscript or handwritten material depends on palaeographic and keyboard skills.  Typically, OCR and unskilled data entry yield less accurate transcripts.
  3. Checking procedures should detect obvious gobbledy-gook, and common OCR and data entry errors.  Double data entry produces a con-census interpretation, but may not avoid common reading errors.
  4. Have error rates been assessed?

Search Facilities

Good search rests on an accurate transcript, not algorithms or user added ‘corrections’.  Repeatable search, essential for confidence in the validity of results, requires a complete data set and stable search methods.  Search is not a simple operation, so inexperienced users need coaching and encouragement, not dumbed-down, limited functionality.  Consider the website’s documentation and functionality of the following:
  1. Targeted search on individual data sets and collections as default. Choosing which collection or collections to search first is much more efficient than filtering out irrelevant collections.
  2. Search on all data items in the record.  This requires a complete transcript.
  3. Full text search, the ability to search everywhere in the record, also requires a complete transcript.
  4. Complex search, the ability to specify ‘AND’, ‘OR’ and other operators.
  5. Name matching algorithm choice.  Examples include soundex and metaphone, which perform phonetic matching for English-language names, and Daitch-Mokotoff, which is adapted for Slavic and German spellings of Jewish names.
  6. Date ranges. Can start and end dates be specified, or a central dates with accuracy?
  7. Are place searches restricted to place names? Is a proximity search based on distance included?
  8. Wildcards, replacement characters in the search term that stand in for unknown possibilities e.g. Sm*th returns Smith, Smyth.
  9. Separation of transcribed values and interpreted values with options to search on either or a combination. For example, the abbreviation ‘Wm’ or Latinised ‘Gulielmus’ can be interpreted as William. Standardised interpretations are known to librarians and archivists as authority control http://en.wikipedia.org/wiki/Authority_control . User added ‘corrections’ are another kind of interpreted value.
  10. Filtering of search results.
  11. Optional ‘sticky’ settings and well-chosen default settings.
  12. Result presentation.  Is it simple, clear and contain the information important to you?  Does it include the search terms used?
  13. Result sorting on data fields chosen by the user. What does ‘relevance’ mean?
  14. Result export in a variety of formats, ready for use by with software tools of your choice.
  15. Logged searches that document research activity.

Browsing Facilities

Browsing is a tool for examining records in the context of the record set.  It should replicate the experience of turning the pages of the original.  The order of digital images should exactly follow the order of the original pages.  The structure of the record set and relationships between individual records contain subtle information about the creation and use of the original.

Browsing does not replace search. When used as a last resort when search fails, it is an indicator of poor search or transcription.

Record Quality

Original records are typically presented as digital images. Genealogists use the most original source available so they can be confident that the information is as reliable as possible. A digital image is not the same as the original, but can come acceptably close, provided that:
  1. Good image quality that is legible.  Sharp focus, resolution, colour accuracy, and contrast all contribute to legibility. Digital image file types vary in the degree of data compression, which influences image quality.
  2. Information that identifies the record portrayed included in the image file. That means all the information that you want in your citation, such as the archive and archive reference of the original, page number, record identifier, person of interest etc. A meaningful file name is helpful, but enough detail makes an unreasonably long name, as is human readable text added to the image. Potentially most useful is embedded citation information in the image file metadata, which is computer readable.
  3. Technical camera or scanner metadata provides provenance of the image, including whether it has been modified.

Your Challenge – Review one data set

There is a lot to consider in assessing the quality of genealogy websites and the data they contain. Of course, we want all the features mentioned above in a user-friendly package, but I think there is quite enough to start with above. Have I omitted anything vital? Do you agree with these criteria?

Before we can hold genealogy data suppliers accountable, we need to fairly assess whether what they offer is of sufficiently good quality for our purposes.  What constitutes ‘fit for purpose’ is open to debate.  I think genealogy data consumers would benefit from setting expectations and demanding quality, and that suppliers would benefit greatly from carefully considered feedback.

In the interest of collaboration between suppliers and consumers, I challenge you to review one data set using these criteria.

Saturday, 12 April 2014

Can We Step up to the Challenge

How Are Your Analysis Skills?

 Background

Pat Richley-Erickson Aka DearMYRTLE posted a challenge to any of her followers/readers on her blog on 2nd April 2014 (1) .
In this post I want to discuss how I approached this, and how on reviewing, I discovered I had issues when analysing my sources.
If you are not aware I am one of the panellists discussing Mastering Genealogical Proof (2) in Study Group 2 a hangout on air recording being held on Sundays. (Schedule available on DearMYRTLE blog (3)) So this challenge is putting any knowledge I have retained from these discussions to the test.

Preparation

Deciding what to write about to create an interesting but not too lengthy post was my first challenge. I have plenty of sources but which ones do I use and should I approach this as a new researcher should and decide on what question I wanted to answer. One might think this was an easy thing to do but like many new researchers I found these sources with an almost "scattergun" approach and it can take something to look at them outside of what is recorded in the standard genealogy software programs.
After at least one failed attempt this was the question I came up with " Who were the siblings and parents of my grandmother who was orphaned? ".
The 3 documents were a birth certificate, census and orphanage document.

Reasons for not publishing

Before I finished writing my post whilst in the process of writing my citations I realised that 2 of the sources I was about to use had the same author.
Both the Birth Certificate and the letter for the orphanage were created by the Superintendent Registrar for Warminster and he had signed both documents.

My original hypothesis to the question about what siblings my grandmother had would have been incorrect. I only ever knew she had one brother, if you look at the census record she is the youngest of 5 children (4). The next question was why had I only ever known the 1 brother.
I should have laid out what I knew about my grandmother and possibly cited this as personal knowledge. I could have also cited her son and daughters for the additional information they had provided.

Also, in retrospect, I had set a question whose parameters were much too broad.
My research question could have been as simple as When was my grandmother born? or Who were her parents?.

In my attempt to show what I had found I had not considered how I had got there.
The records I have contain a large amount of useful information, much of it can be classed as primary. But the number of documents I will need to cite, to fully build my conclusions for the question I originally posed, will most certainly exceed the three I was initially going to use.

I hope that this exercise has shown that we not only need to analyse the documents that we use but we need to think about how we obtained them and have they answered the question that we originally wanted to answer. Why the records were originally created, is a question we need to ask if we do not wanted to duplicate sources from the same informant.

So are YOU and your ANALYSIS SKILLS up to the CHALLENGE if so then please join in.

We can all benefit from peer review and even if you don't plan to publish your research in a recognised journal it is well worth it for the thought process it takes to put your case.

When I get more time in the coming week I hope to step up to this challenge and I will post a link in the comments. 


References

  1. http://blog.dearmyrtle.com/2014/04/the-ragu-challenge-3-2-1-cite.html
  2. Thomas W. Jones, Mastering Genealogical Proof (Arlington, Virginia: National Genealogical Society, 2013), 6. [Book available from the publisher at http://www.ngsgenealogy.org/cs/mastering_genealogical_proof ]
  3. 1901 census of England, Wiltshire, Warminster, Christchurch parish, folio 37 recto, p.9, household 53, Edmond Compton; digital image, Find My Past (http://new.findmypast.co.uk/ : accessed 5th April 2014) ; citing PRO RG13 /1943