Showing posts with label Data-analysis. Show all posts
Showing posts with label Data-analysis. Show all posts

Saturday, 1 June 2013

The ancient of day's marking



During my career as an academic member of staff at the University of Birmingham, I probably marked thousands of essays, from both undergraduates and postgraduates. The most miserable experience was marking hundreds of medical student examination papers at a time. As with most markers, I became a sort of machine. I would devise a checklist of suitable responses to each question and tick the number of times each student used one. I could thereby mark each mini-essay in two or three minutes. Much more enjoyable were the essays written by undergraduate medical students for their elective project on learning disability and health, which I still teach and mark. Students choose their own topic within this field and always produce long essays that are stimulating and, in some cases, worthy of publication.

A further problem with marking is not just the number of scripts to mark, but also the requirement to assign a numerical score to a long essay. It was different when I worked with my colleague Dr Beryl Smith when we set up our postgraduate masters course in intellectual (‘learning’) disabilities in 1992. Students completed eight assignments (each of which would be answered by long essays of between 1500 and 3000 words) and a dissertation of 15,000 words. Students were encouraged to write about their area of special clinical interest (such as challenging behaviour, epilepsy and so on). We decided that only four grades were required to mark each assignment to an acceptable standard of reliability. We gave a B if the student met the specified requirement for the assignment, a C if they answered the question but did not argue their case well or failed to draw on sufficient evidence. We gave an A where the answer was a high standard and would be publishable in a professional journal. Finally, the failed D grade was given where the student did not meet the requirements of the essay assignment. To make this scheme work, we needed to make sure the requirements of the assignment were stated clearly, and we included details of our marking-scheme. This system was reliable because there were only three grade boundaries (A/B, B/C and C/D) to decide. We double-marked all assignments and usually came to rapid agreement on the grading for each essay.

All this of course would look like mollycoddling to the sort of academic who believes that the only aims of examinations are to catch students out and identify a would-be elite. Beryl and I took a different approach: we believed that the purpose of our course and hence the marking of assignments was to help students learn to become more reflective and effective practitioners. This would, indirectly, be our contribution to improving the lives of people with intellectual disabilities. The purpose of the assignments was not just to decide whether each student’s work was of an acceptable standard, but also to help us measure their progress, and see what areas required some individual attention. As important as the grade, therefore, were the detailed comments we completed on each essay, identifying how the student could improve what they had written, areas of strength they could develop, and areas of weakness they could concentrate on improving.

Beryl eventually retired and the University introduced new regulations that stipulated that all marking should be numerical. Instead of our four grades with clear descriptions and marked to a high degree of reliability, we had to use percentages. What nobody could tell me was what they were a percentage of. If students are set a large number of questions to answer (for instance in a maths exam), then it is possible to calculate the percentage they answered correctly. This only has any meaning of course if each question is deemed to be of equal difficulty and all questions can be marked as either ‘correct’ or ‘incorrect’. But to give a ‘percentage’ for an essay suggests that this is the degree to which it approximates to some perfect essay. No such perfect essay exists. Indeed, it seems rare in universities for any essay, however good, to be given more than 85 ‘percent’. At the lowest end of the scale, I have once or twice given a mark of 35 ‘percent’, but I think that almost all marks for essays fall somewhere between these two extremes.

Since it is not possible to define what marks are a percentage of, the marks are not actually percentages at all: they are ‘pseudo-percentages’. They are an example of the belief that scores and numbers are preferable to description, even when they are used to supposedly measure things that are inherently non-numerical. Academic psychologists are probably the most prone to this disorder, ‘measuring’ such diverse concepts as intelligence, affection, extroversion and so on with numerical scales calculated by summing answers to sets of questions or assigning weights to responses to various ingenious types of numeric scales. Sooner or later, people come to believe that because it is possible to assign a numerical score, then there must be a thing corresponding to it. So  many believe there is an entity called ‘intelligence’ which you can either have a lot of or a little of. This is despite the common-sense observation that people often have a very uneven pattern of mental skills, being, for instance, brilliant at thinking through maths problems but incompetent at remembering times and dates.

The urge to assign scores to people also says something about the people assigning the scores. No competent clinical psychologist would believe that an individual patient could be summarised by a few numerical test scores. Instead, each patient is seen as unique, perhaps having some familiar categories of problems, but still assessed and treated as an individual. A clinical psychologist and academics like myself and Beryl can afford to treat people as individuals because we are assessing so few of them. Once organisations become involved in processing large numbers of people, they see people as numbers. This has happened to many universities. There may be hundreds of students in a single year of an undergraduate course, knowing little of the academics who teach it and having few chances to exchange ideas with them. Indeed, some universities try hard to prevent any such exchange, placing their academic staff in research centres and laboratory blocks behind locked doors, inaccessible to mere undergraduate students. They have become people-processing institutions which, like business corporations, are judged not by how they improve the lives of ordinary citizens but by how much income they generate. Cash has thus become the supreme number which measures all. It is the Modern of Days, replacing the Ancient of Days in William Blake’s painting.

See also

Wednesday, 10 April 2013

Laws of Information 4 and 5

A long time ago, I proposed some ‘laws of information’, looking particularly at the kind of information available to manage large public organisation. These are as below:
1. Information is costly
2. Data is always less reliable than you think.
3. data that is collected to measure performance loses reliability
You can click these weblinks to access the relevant text for each law.

Here is another law:
4. There is always more information available than you first thought.
Natural science has advanced through ever-improved measurement techniques, of which the Large Hadron Collider is the most recent and by far the most expensive. Each area of science has its own preferred measurement technique, and great effort is expended in improving its accuracy and reliability. Social science works on very different principles: no single measure remotely approaches the levels of accuracy taken for granted in the natural sciences, and so social scientists draw on multiple sources of data to base their conclusions. This principle is also a good one for managing large public organisations like universities and hospitals, where there is also a mass of different information, sometimes of dubious reliability (for reasons why it is dubious, see Laws of Information 1-3).

Sadly, this principle is not always applied. Managers and politicians often focus on one (unreliable) measurement and ignore the others. In part, this is a consequence of a commitment to the written word. Nothing is truly believed to exist unless it has been written down, preferably on a form. Once written down, it is believed superior to all other forms of information such as observation, informal discussions with staff, or patients’ letters of complaint. The most holy of all written data is quantitative data, especially that emanating from a computer. This tendency is reinforced by the use by governments of simple quantitative targets to measure the complex activities of complex institutions.

As an example, look at the Stafford Hospital case. Analyses of routine data showed that the hospital was an outlier in mortality statistics for some surgical procedures (ie people were much more likely to die). This all explained in an excellent article in the London Review of Books, available here:
Rigging the death rate
If the hospital management had followed the Fourth Law of Information, they would have seen this analysis of mortality statistics as a sign that they should gather data from other sources. They could then have visited the wards and observed daily care, spoken to patients and staff, reviewed casenotes, checked how their staffing levels compared with those of other similar hospitals, or brought in some outside experts to do these things and advise. They don’t seem to have done any of this: instead, they decided to discredit the mortality statistics. Management consultants were brought in to change the diagnostic codes of patients who died in hospital. Researchers at Birmingham University were funded to discredit the use of statistics to assess hospital performance. Their report made the correct conclusion that statistics can be misleading and that one set of them should not be used exclusively to assess performance. But that of course misses the point. Being an outlier should be regarded as a warning sign rather than definitive proof. It should have indicated a need to collect other data. In other words, the truth is found not in one set of data, however tidily it is presented and however quantitative, but in a wide range of information, from which an informed person can make a judgement. This leads to the fifth law of information:
5. Interpreting information requires judgement.

The word ‘judgement’ of course will sound a warning bell to some. How much better to pretend that decisions follow automatically from the data without human intervention or the exercise of personal responsibility. Then all that is needed when things go wrong is for the relevant procedures to be blamed and amended. This defence (“I was only carrying out procedures”) is an effective life strategy in any large organisation, and may be a more reliable path to promotion than anticipating problems and taking the initiative in solving them. Look around, and you will see the consequences.

Friday, 31 August 2012

Rank bad

The London Olympics were a festival for statisticians. As an example, let’s look at the men’s 100 metres. Usain Bolt won this race with a time of 9.63 seconds (a new Olympic record). But all the entrants had very similar times apart from Asafa Powell, who pulled up with a groin injury and completed the race in 11.99 seconds. Leaving him out of the calculations, the difference between Usain Bolt and the last-but-one runner was only 0.35 seconds. This tiny sliver of time is only 3.6% longer than Bolt’s winning time. So a statistical summary of the race (using nonparametric statistics because of the skewed distribution) would say that the median race time was 9.84 seconds, with an interquartile range of  0.76-0.97 seconds. Or, in everyday language, they all (but one) ran very fast and there was little variation in the times they took.

Of course, this misses the point. The 100 metres final is a race, and what matters in a race is the ranking of the runners - who comes first (and to a less degree second and third). But just as it makes no sense to ignore ranking in races, it is equally pointless to analyse all areas of human activity as if they are races. This does not stop people doing it, and rankings are now common not just for things that are simple to measure (such as the time taken to run a race), but also for institutions and the complex range of activities they perform. This means that such rankings have to be derived from multiples of ‘scores’ based on a range of unreliable data about something that can never be reliably measured.

As an example, let’s take university research and teaching. Recent years have seen a growing collection of international ‘league tables’ of universities. These use a wide range of data, including statistics about numbers of staff and students, numbers of publications in academic journals, income from research grants, and rankings by panels of academics. This data is then scored, weighted and combined to produce a combined score which can be ranked. Individual universities can then congratulate themselves that they have ascended from the 79th best university in the world to the 78th best, while those that have descended a place or two can fret, threaten their academic staff and sack their vice chancellor.

Yet higher education systems differ greatly between countries, and universities themselves are usually a diverse ragbag of big and small research groups, teaching teams, departments, and schools. This makes a single league table a dubious affair, even if it is based on reliable data. But that of course is not the case. The university in which I worked (by most standards a well-managed institution) struggled to find out what its academic staff were doing with their time, or the quality of their achievements. In other universities, the data on which international league tables are based may be little more than a work of mystery and imagination.

But the problem lies not so much in the dubious quality of the data, but the very act of ranking. Even when the results are derived from a single survey in one country, the results are often analysed in a misleading way. As an example, look at the National Student Survey in the UK, which is taken very seriously by the UK higher education sector. In its most recent form, this comprises 22 questions about different aspects of the student’s university and course. All ranked on a five-point Likert scale and given a score from ‘definitely agree’ (scored 5) to ‘definitely disagree’ (scored 1). A conventional way of comparing universities and courses would therefore be to take the mean score on each scale, and this is how they are analysed in papers like the Guardian. So, by institution, the results for the overall satisfaction question vary from the maximum of 4.5 (the Open University) to 3.5 (the University of the Arts, London).

This is all seen as being a bit too technical for prospective students, and so the Unistats website reports only the responses to a single statement in the Survey: “Overall, I am satisfied with the quality of the course”. It then adds together the number of students with scores of 5 (‘definitely agree’) and 4 (‘mostly agree’) to produce the percentage of ‘satisfied’ students. This is quite a common survey procedure (I admit to having done it myself), but it is flawed. Someone who ‘mostly agrees’ may still have important reservations about their course.

If you actually analyse satisfaction scores, you find that the great majority of universities fall within a narrow range, with some outliers. As an example, ‘satisfaction’ among students with degrees in medicine in English universities range from 99% in Oxford to 69% in Manchester. The median satisfaction is 89%, with the interquartile range between 84% and 94%. So half of all medical schools fall within only ten percentage points around the middle of the range. This makes ranking pointless, because a small change in percentage satisfaction from one year to the next could send an individual medical schools several places up the rankings, but would amount to little more than the usual fluctuations common to surveys of this kind.

A far more useful step is to look at the outliers. What is so special about Oxford (99%) and Leeds (97%)? Alternatively, are there problems at Manchester (69%) and Liverpool (70%)? Before we get too excited about the latter two universities, note that they have levels of satisfactions that most politicians and people in the media would only dream about. However, to see if there are particular problems, we need to look in more depth at the full range of results. We could also see if there is anything distinctive in the way they teach. Actually, we do know that both universities have been very committed to problem-based learning (PBL). This is a way of teaching medicine that involves replacing conventional lecture-based teaching by a system whereby small groups of students are set a series of written case descriptions. Students then work as a group to investigate the scientific basis for the presenting problem and the evidence for the most effective treatment.

Research on PBL in medicine is (in common with a lot of research in education) inconclusive. But medical students are very bright and highly-motivated, and would probably triumph if their education amounted to little more than setting them a weekly question to answer and presenting them with a pile of textbooks to read. Come to think of it, this more or less describes how PBL operates.

Tuesday, 9 March 2010

A survey has shown that...

If you want to get publicity for some idea, promote a product, or just get in the news, then you should report the results of a meaningless survey. Search on Google using the phrase "a survey has shown that...", and you will see what I mean. You will learn that one in six therapists have tried to cure homosexuals, that more than 70% of people would exchange their computer password for a bar of chocolate, that Americans who attend church are more likely to favour torture than those that do not, and so on. You don't need to bother with getting a good response rate, a representative sample, or even a valid and reliable questionnaire. Just circulate some questions to a few people, and send the most eye-catching result to the press.

There are also plenty of meaningless surveys which never get to the press, but are circulated within companies, government departments and universities. These are often promoted as 'quality assurance', and are even taken seriously by some people. Management boards ponder reasons for a fall in satisfaction ratings by 5% on a survey with a response rate of 20%, without admitting that the whole exercise does not mean very much. Truth to tell, survey results might not mean much even if the response rate was 100%. Many meaningless surveys use ambiguous questions coupled with dubious Likert scales (the kind which assign numerical scores to a range of five or so questions from 'very satisfied' to 'very dissatisfied'). These have the apparent advantage of producing a numerical score and hence allowing statistical analysis. Usually however, people only look at mean scores, and these can be misleading. A survey in which 50% of respondents were 'very satisfied' and 50% 'very dissatisfied' would produce the same mean score as one in which 100% said they were 'neither satisfied or dissatisfied'.

What's the alternative? It is essential for organisations to assess the quality of what they do, and their customers/citizens/students are in a good position to assess this. Rather than assessing mean scores on Likert scales, organisations should concentrate their attention on the causes of satisfaction and dissatisfaction, and ideas for improverment. The best way of doing this is probably to use open-ended interviews or focus groups. Of course, this would require quality assurance staff to be skillful in survey techniques, to be creative, and to be prepare to co-operate with front-line staff rather than stand in judgement over them.

Monday, 10 August 2009

The Laws of Information No. 3

The third law of information is:

3. Data that is collected to measure performance loses validity.

First a confession. In the late 1980s, I worked as Director of Planning and Information in a mental health service in the NHS. One of my tasks was to organise the statistical returns on clinical activity for despatch to the Department of Health. At that time, these were based on a set of standard definitions called the ‘Korner system’ (after a woman who chaired a committee which recommended them). Our service included a brilliant and very hard-working consultant psychiatrist for the elderly. She believed that assessments of new patients should initially be in the patient’s own home (a ‘domiciliary visit’ or ‘DV’). Since she was an orthodox Jew, this meant a lot of walking when her duty days coincided with the Sabbath. Unfortunately, the Korner system required information about scheduled outpatient clinics but not domiciliary visits. Following the Korner rules would have meant that our most active consultant would appear as our least active. This was obviously unjust, so I modified the returns for her clinical activity to record each DV as an attendance at a (non-existent) outpatient clinic.

Paradoxically, my data-adjustment produced statistical returns which were a more accurate reflection of clinical activity than would have been the case without such adjustment. Nonetheless, they became an inaccurate record of inpatient clinics in the service in which I worked. I suspect that data-adjustment in the desired direction was and is rife in the NHS. Although this is dishonest, it can cause far less damage than changing reality to generate honest statistics. A well-known example of changing reality in the NHS is to make patients wait in ambulances outside A&E departments. This reduces the time the patient spends in A&E for the purposes of official statistics, and hence enables the hospital to meet a government target. There are many, many more examples in the NHS of how meeting centrally-imposed targets can damage patient care.

This is not a recent phenomenon. The whole technology of corporate strategic planning and management by targets owes its origins to Gosplan, the state planning commission in the USSR. Studies of the Soviet economy in the 1960s and onwards were full of examples of how rational responses by individual enterprises to centrally-determined targets could produce absurd results. These included the shoe factory that met its target number of shoes by producing shoes all in one size, and the steel factory that met it target for weight of steel by producing a few huge ingots.

Monday, 27 July 2009

The Laws of Information No. 2

Staff in offices, universities, schools and almost everywhere else are communication victims. The management in my own university is excellent at communicating to its staff. There are attractive magazines full of good news, regular staff meetings in which college heads present their challenges and achievements, all backed up by daily emails from an array of administrators to guide staff about their business. Yet a recent survey of staff has found dissatisfaction with ‘communication’. What could be the solution? More attractive magazines? More meetings? One answer that has not been considered is less (but more useful) information. As I noted in my posting on the First Law of Information, information is costly. Staff believe that all information emerging from senior management must be important, and therefore it must be read and understood. They do not have the time to do this in addition to all the other emails they receive daily, so messages accumulate in inboxes unread.

The cost of information is felt most acutely by staff when it is required from them. There are routine statistics to be completed, forms to be filled on staff and student progress (including one for every single meeting with a research student!), surveys of staff satisfaction, and one-off requests for information which have descended the management line (usually with shorter and shorter deadlines at each stage of transmission). Staff usually see these requests as a chore to be completed quickly, and do not therefore strive tirelessly for accuracy in collecting and recording the required data. This leads to the second law of information:

2. Data is always less reliable than you think.
Scientific texts emphasise the potential pitfalls in gathering data, and careful scientists have standard routines for checking its validity and reliability. Gathering research data from people is particularly troublesome because of their capacity to fabricate, to rationalise, to forget, and even to avoid telling the truth as an act of politeness. Even in a world where people did none of these things, there would still be lag between events occurring and data being collected, inconsistent application of rules for categorising data, and missing data. Yet these limitations are usually ignored when organisations collect and process information from their staff or from the public. Instead of using wide confidence intervals when reporting the information they have collected, organisations glibly report data to an exact percentage point. There are earnest debates about small changes in statistics from one reporting period to another, even though these are probably within confidence intervals.

Interpreting data would be difficult enough if it was simply a matter of general unreliability, but there is the far bigger problem of biassed unreliability. This is the third law of information:

3. Data that is collected to measure performance loses validity.
I will deal with this in the next posting.

Tuesday, 14 July 2009

The Laws of Information No. 1

After finishing my first degree, I worked for a summer in a typewriter factory. Typewriters are now so obsolete that it is usually necessary to explain to younger people what they were for. But this experience taught me a lot about information and how it is used in organisations. This was because I worked on what was then called ‘O&M’, reporting to a rather odd but very clever Welshman. The first law of information that I learnt was:

1. Information is costly. Back in 1968, there were no photocopiers, office computers, or emails. If you wanted a copy of a letter, the typist had to insert carbons and additional sheets of paper when she typed. There was a limit of about three or four copies that could be made this way. If you needed more, then a different process was required. The typist would type the report on specially-waxed ‘skins’, which were attached to the drum of a machine we called a ‘Gestetner’. The drum would contain thick black ink, which you always got on your hands. Both methods of copying were costly and time-consuming, and a major O&M task was therefore to reduce the amount of unnecessary information circulating round the factory. We did this by creating a flow diagram for all the routine reports generated by staff, and asking their recipients whether they found them useful. We found that many reports had begun as one-off requests by management to meet a specific need, but had then become routinised. Some reports went straight from the envelope to the waste paper bin.

This seems a lost world now because photocopiers, word-processing and emailing have successively made the production of multiple copies much easier. But this has had the effect of shifting the cost of information to the reader. People in offices now spend hours a week sifting through emails, most of which come from their seniors but are irrelevant to their work. Emails accumulate in inboxes, and the ones which require rapid attention are missed. Because the idea has taken root that information is cheap to reproduce, staff are required, often at short notice, to produce data and statistics for senior management. As in the past, these requests can become routinised even when the original need for the information has passed.

This indicates that organisations should revert to the O&M principle of reducing the flow of unnecessary information, to release staff time and speed up their response to the information that really matters. Without this, problems develop with the data we do have, which I will look at in a later posting.