Pages

Showing posts with label validity. Show all posts
Showing posts with label validity. Show all posts

Monday, 26 January 2026

Are LLM apps the 'man behind the curtain'?

I have noticed more and more that people are using AI apps to put together text for novel situations. For example, a colleague of mine recently used the Open AI app, ChatGPT, to write a response to a student complaint, and the resulting answer was a superb piece of sympathetic, supportive writing intended to smooth those ruffled feathers. The response was read for sense, edited to suit our context, and emailed out. It seemed to work in practice, as well.

Personally, the idea of pasting another's email into such an app feels like stealing: I am giving away another's words under my own aegis. I am uncomfortable with that. If it is my own writing that needs editing down (e.g. sweating a 500 word draft into two paragraphs), I semi-OK with it, but even then I am struggling. Perhaps 'hesitant' might better describe how I feel. It feels ...dishonest.

Perhaps it feels like cheating because writing is a "time skill", like keyboarding (Canning, 1975, p. 277). It is “where the ticking away of the unforgiving seconds plays a dominant part in both learning and application of the skill” (p. 277), now known as sequential learning (what comes next) and map learning (where we find that next thing). Our skills become automatic though practice, helping us to embed both types of learning (Canning, 1975; Rao et al, 2000). Like learning to drive, like cooking, like carpentry, like becoming a doctor, or like glassblowing; there are no shortcuts. We learn by doing year on year, gathering mastery. Once we gain mastery, we can take some shortcuts: but too many will erode our hard-won skills. I worry that LLMs writing on our behalf will erode our writing mastery. Everyone may be doing it, but I am not sure that makes it any better. 

Adding more to my unease, if we are using the large language model AI apps, how do we know the apps we use are valid, and don't suffer from hallucinations (Lingard, 2023)? Trying to compare AI LLM apps is difficult for the average user: there are so many options, and how do we actually weight them against each other? 

Well, that is where LM Arena https://beta.lmarena.ai/ comes in (Bansal, 2025). We can ask a question in this online app, and we get pointed to a pair of LLMs, and asked to chose our preferred answer from the two results from two anonymised LLMs. We can chose to ask again; to clarify further; or accept the result (which will then unmask which LLM we prefer). Having tried a few times, I am apparently drawn to Claude.ai and DeepSeek. No surprises there: these are the two models I have found in my own limited usage to provide the most accurate responses to my personal requests.

However, I am still uneasy, and the Bansal article further explores yet another reason for my unease: the inequity of human AI 'raters': those underpaid workers who 'massage' those responses we receive to our AI app requests (2025). This is truly smoke and mirrors: AIs need a human baby-sitter. 

Human AI raters apparently don't "trust [...] the products they are helping build and train. Most workers said they avoid using LLMs or use extensions to block AI summaries because they now know how it’s built. Many also discourage their family and friends from using it, for the same reason" (Bansal, 2025). It is reminiscent of the Wizard of Oz, perhaps the most famous "man behind the curtain"; the charlatan of our childhood imaginings (Baum, 1900). It looks and sounds great, but it is all a trick.

If those who work on the apps avoid them, perhaps some unease is warranted.


Sam

References:

Bansal, V. (2025, September 11). How thousands of ‘overworked, underpaid’ humans train Google’s AI to seem smart. The Guardian. https://www.theguardian.com/technology/2025/sep/11/google-gemini-ai-training-humans

Baum, L. F. (1900). The Wonderful Wizard of Oz (illustrations W. W. Denslow). George M. Hill Company.

Canning, B. W. (1975). Keyboard skill-a useful business accompaniment. Education + Training, 17(10), 277-278. https://doi.org/10.1108/eb016409

Lingard, L. (2023). Writing with ChatGPT: An illustration of its capacity, limitations & implications for academic writers. Perspectives on Medical Education, 12(1), 261-270. https://doi.org/10.5334/pme.1072

Rao, S. M., Harrington, D., & Parsons, M. W. (2000). Acquisition of keyboarding skills: An event-related fMRI study. NeuroImage, 11(5), S361. https://doi.org/10.1016/S1053-8119(00)91292-8

read more "Are LLM apps the 'man behind the curtain'?"

Friday, 9 May 2025

Making sense of testing

We use career assessments in order to help our clients in identifying their unique characteristics. Each assessment is designed to measure different components, thus - with appropriate interpretation - assisting our clients to find career options which match their particular attributes, values, and skills (Osborn & Zunker, 2016).

While tests can assist client's decision making processes (Whitfield et al., 2009), to be effective, those tests need to be reliable and valid (Walsh & Betz, 2000). If a test is valid, it means that it actually measures what it says it measures: it does what it says on the tin (Heale & Twycross, 2015). There are three key types of validity: content validity (test accuracy); construct validity (does what it says on the tin - e.g. testing for job search skills might inadvertently be evaluating problem-solving skills); and criterion-related validity (where the same factor - or variable - is measured each time, through 'convergent' validity which is strongly correlated with similar tests; 'divergent' validity with poor correlation to different tests; and 'predictive' validity where the test is highly correlated to related factors - e.g. being task-oriented should lead to being a completer/finisher) (Heale & Twycross, 2015). 

Tests also need to have been normalised for the population group our client affiliates (awhis) to. That means that, when assessments are created, researchers have run a number of sample tests (usually around 300; Steve Evans, personal communication, 13 September 2021) on each population group, seeking normal distribution in the test results via cultural, ethnic, gender, political and socio-economic group factors (Hansen, 2003; Osborn & Zunker, 2016). We can see that normalising tests is going to be an expensive business, in giving 300 tests to measure each norm group.

We also need to have consistent test-retest rates: the same result needs to be achieved each time the test is run (Heale & Twycross, 2015). If our client does a test in March, we don't want to see that they obtain a completely different result when they repeat the test in July (one of the main bug-bears of MBTI; Mastrangelo, 2001). While it’s not possible to perfectly assess each career instrument, we can estimate their replicability (Heale & Twycross, 2015) through “internal [...] and test-retest reliability” (Osborn & Zunker, 2016, p. 37).

And, while we might have all reliability, validity and representative norm groups, we might still find that our client does not suit the test we propose. The client may complete the test and end up with results which make no sense. For example, each time I complete a RIASEC test, I get a different score. Over the years, I think I have seen a pattern: that in those of us with very generalist skills, the RIASEC test may lose it's test-retest reliability. I offer RIASEC here as one example: it is not the only one I have noticed. I have had clients who achieve poor results from HBDI, from MBTI, and from DiSC. All tests do not necessarily suit all people.

We must take all quantitative tests with a pinch of salt :-)


Sam

References:

Hansen, S. S. (2003). Career counselors as advocates and change agents for equality. The Career Development Quarterly, 52(1), 43-53. https://doi.org/10.1002/j.2161-0045.2003.tb00626.x

Heale, R., & Twycross, A. (2015). Validity and reliability in quantitative studies. Evidence Based Nursing, 18(3), 66-67. https://doi.org/10.1136/eb-2015-102129

Herr, E. A. (2001). Chapter 2: Career Assessment: Perspectives on trends and issues. In J. T. Kapes, E. A. Whitfield (Eds.), A counselor's guide to career assessment instruments (4th ed., pp. 15-26). National Career Development Association.

Mastrangelo, P. M. (2001). [251] Myers-Briggs Type Indicator [Form M]. In B. S. Plake & J. C. Impara (Eds.), The fourteenth mental measurements yearbook (816-820). Buros Center for Testing.

Osborn, D. S., & Zunker, V. G. (2016). Using Assessment Results for Career Development (9th ed.). Cengage Learning.

Walsh, W. B., & Betz, N. E. (2000). Tests and Assessment (4th ed.). Prentice Hall.

Whitfield, E. A., Feller, R. W., & Wood, C. (Eds.). (2009). A counselor’s guide to career assessment instruments (5th ed., pp. 13–25). National Career Development Association.

read more "Making sense of testing"

Friday, 25 April 2025

Checking email addresses

As emails is still the most common way we communicate with our network, it is great having somewhere that an address can be checked before sending something out to a new contact. And there is a website - Clean Talk - where we can load an email address, and the site will check whether the email is viable. 

Sometimes websites make their email contacts hard to find. If faced with a 'contact us' form, most of us won't bother trying. Contact forms are usually set up to be assigned to one person, but due to staff turnover lose the allocation, so end up going to a long-lost and never explored folder. We can avoid the contact form with Clean Talk, instead trying a few email formats to check to see if we can work out a viable email address, and email the organisation directly.

Why else would we want to check an email? Apparently around "30% of email addresses [...] used to spam websites are fake" (Clean Talk, 2025). So the site also checks to see if an email is 'real', or will re-route the email to the actual client email; and will check to see if it has been blacklisted (i.e. it has been reported as a spam email or site). We don't want to email a blacklisted email as that can splash back on us.

I have used this intermittently, without an account, but I would imagine if we were wanting to verify a lot of emails, then having a paid account would be required. 

This is quite a handy site for intermittent use, however!


Sam

References:

Clean Talk. (2025). Email Checker. https://cleantalk.org/email-checker/

read more "Checking email addresses"

Wednesday, 2 April 2025

Using career assessments from other countries

It is quite a process to create, test and normalise career assessment instruments (Stuart, 2004), but living in the Antipodes, where we have such a small population - only 5m - it would also be a costly procedure. Pretty much the only quantitative tools we have in Aotearoa are tests which have been internationally-developed. So, if we career practitioners in New Zealand want to give our clients evidence-based assessments, we have to rely on those which have been developed elsewhere. But are those international assessments worth using, from a cultural appropriateness point of view, or should we avoid quantitative testing altogether?

Due to our geographic isolation, rural roots, and confluence of Pākehā, Māori and Pasifika ethnicities, New Zealand's multicultural society is unique. Māori and Pasifika cultures have tended to focus more on collective well-being, interdependence, and respect for the environment; as opposed to the Western individualism arising from the Pākehā settlers (Harmsworth, 2005; while noting that all three culture are moving closer together). Our social norms, leadership styles, and personal interactions of Aotearoa mean that we prize modesty, practicality, and resilience (Harmsworth, 2005). Due to Māori and Pasifika cultural influence, all New Zealanders may have more community-oriented career goals, on average, than other nations. In fact, the John Hopkins Institute collected and cross-tabulated UN volunteer data, which showed that New Zealand has the most volunteers by a third, even though our not-for-profit sector is smaller than some other nations (Belgium, Australia and Israel; GMVP, 2013). Volunteering in New Zealand appears more culturally endemic than in Australia; apparently 50% of Kiwis volunteer versus 5% of Aussies volunteer (SNZ, 2006; VNZ, 2024). 

Our differing values may mean that international test validity may not translate to test validity here in Aotearoa. But why should we use quantitative assessments anyway? Well, there are good reasons. It seems that clients who complete assessment instruments have a deeper understanding of their own interests, values and strengths (Heppner et al., 1994). In addition, clients tend to make more informed career decisions, and seem to experience less career indecision as a result of testing (Heppner et al., 1994). Even better, clients who took assessments as part of seeing a career practitioner experienced more positive career outcomes, including better career goal alignment, increased job satisfaction, and improved career advancement (Heppner et al., 1994).

It appears that knowing ourselves may assist our career decision making, how we further our careers, and make us happier in our work. So, as long as we don't put too much emphasis on the tests (don't treat them as gospel), then the tests give our clients some clarity.

Bonus.



Sam

References:

GVMP. (2011). The Global Volunteer Measurement Project. http://volunteermeasurement.org/

Heppner, M. J., O'Brien, K. M., Hinkelman, J. M., & Humphrey, C. F. (1994). Shifting the paradigm: The use of creativity in career counseling. Journal of Career Development, 21(2), 77-86. https://doi.org/10.1177/089484539402100202

SNZ. (2006). Finding and Keeping Volunteers [report]. Sport New Zealand [formerly SPARC]. http://www.sparc.org.nz/filedownload?id=850d18af-002f-40b7-b989-5a99e5b40f82

Stuart, B. (2004). Twelve Practical Suggestions for Achieving Multicultural Competence. Professional Psychology: Research and Practice 35(1) 3–9. https://doi.org/10.1037/0735-7028.35.1.3

VNZ. (2024). State of Volunteering Report 2024 [report]. Tuao Aotearoa | Volunteering New Zealand. https://www.volunteeringnz.org.nz/wp-content/uploads/f_SOV-report_2024_web.pdf

read more "Using career assessments from other countries"

Monday, 24 February 2025

The Emperor has no clothes & AI

While I am sure that AI platforms will improve, I was struck by a Guardian long read article last year where a journalist reported that, "when I asked ChatGPT to write a bio for me, it told me I was born in India, went to Carleton University and had a degree in journalism – about which it was wrong on all three counts (it was the UK, York University and English). To ChatGPT, it was the shape of the answer, expressed confidently, that was more important than the content, the right pattern mattering more than the right response" (Alang, 2024).

I think that is the core of the AI problem. The confidence of the delivery from the AIs we consult (Alang, 2024). The large language models which AI is trained upon is logically North American. That is where the tech companies are. The USA has driven much of the research and IT work for the past half century. The US is probably the most WEIRD society (here; Henrich et al., 2010): a Western, educated, industrialised, rich and democratic society, which collectively make up 12% of the global population. Researchers have considered "how WEIRD [society populations] measure up relative to the available reference populations" (p. 62), finding that in most behavioural research studies, a full "68% of [research participants] came from the United States, and a full 96% of subjects were from Western industrialized countries, specifically those in North America and Europe, as well as Australia and Israel" (p. 63); and even more narrow, that "67% of the American [participants] (and 80% of the [participants] from other countries) were composed solely of undergraduates in psychology courses" (p. 63).

So not very representative then. And if we think of the 12% of global population in WEIRD societies, 50% will be female. Around 40% of Americans go to college (National Center for Education Statistics, 2020). So lets assume of the 6% of WEIRD societies which are male, that 40% have gone to college. While this is very rough maths, at the most, AI is based on 2.4% of the global population (and it will be a fraction of that number, because few will have completed an IT degree, let alone a behavioural science degree, as per Henrich et al., 2010). Yes, I know I am comparing apples with oranges, but I don’t think we can safely assume that the data being used to 'train' the AI models is unbiased. I think it is pretty clear that the training data is based on a tiny non-representative percentage of the global population. 

The software and hardware engineers working on AI are also likely to be male, with a good chunk from North America (Alang, 2024). While, 23% of workers in IT are women (Deloitte, 2021), it was noted at one large US company that there were "641 people working on 'machine intelligence,' of whom only 10 percent were women" (Simonite, 2018). So yes, while nearly a quarter of the IT sector has women in it, the gender distribution is uneven. And if we come back to Von Bertalanffy's system theory (1968), this shows that the input is definitely biased. Thus the transformation - no matter what we do elsewhere - will also be biased. This means that the output too will be biased. 

We are used to consulting the internet for factual answers. Yet there is a growing trend that what is on the internet is a mashup of fact and fiction. Since the early 1990s, we 'little people' have been able to create our voices without the peer review of publishers and others to filter what we say. And now, perhaps throwing a massive spanner in the works, generative AI creates blends of fiction and fact... and - unless we know our field - we have little idea which elements are fiction, and which are factual (Alang, 2024; Lingard, 2023). We consult the oracle and lack the understanding to be able to point out that the emperor has no clothes.

But the more I read, the more I think that the emperor is indeed naked. So far, anyway.


Sam

References:

Alang, N. (2024, August 8). No god in the machine: the pitfalls of AI worship. The Guardian. https://www.theguardian.com/news/article/2024/aug/08/no-god-in-the-machine-the-pitfalls-of-ai-worship

Deloitte. (2021, December 1). Women in the tech industry: Gaining ground, but facing new headwind. https://www2.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2022/statistics-show-women-in-technology-are-facing-new-headwinds.html

Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world?. Behavioral and Brain Sciences, 33(2-3), 61-83. https://doi.org/10.1017/S0140525X0999152X

Lingard, L. (2023). Writing with ChatGPT: An illustration of its capacity, limitations & implications for academic writers. Perspectives on Medical Education, 12(1), 261-270. https://doi.org/10.5334/pme.1072

National Center for Education Statistics. (2020). Chapter 2: College Enrollment Rates. In The Condition of Education. https://nces.ed.gov/programs/coe/pdf/coe_cpb.pdf

Simonite, T. (2018, August 17). AI Is the Future—But Where Are the Women?. WIRED. https://www.wired.com/story/artificial-intelligence-researchers-gender-imbalance/

Von Bertalanffy, L. (1968). General System Theory:  Foundations, Development, Applications. George Braziller.

read more "The Emperor has no clothes & AI"

Friday, 25 October 2024

AI Hallucination

I have a burning question around the validity of AI. I have run my own tests (here) in ChatGPT, where I felt that the hype around AI was just that: hype. My brief and rather unscientific experiments found that the AI I had used (ChatGPT) effectively made up the answers I obtained, which is called "AI hallucination" (Lingard, 2023). I know the answers I was given by the AI were largely nonsense because I carefully validated what the AI had delivered to me.

The trouble is, when we are writing academically, we cannot afford to stand our arguments on dodgy evidence. So if we lack a reasonable expectation that ChatGPT will supply us with sound evidence, its 'use' becomes useless. If it becomes “crucial for students to factcheck all ChatGPT output during interaction with the system to identify potential biases or inaccuracies to construct an accurate understanding of the topic” (Rasul et al., 2023, p. 8), how many students are going to do that? And if students DON'T fact-check, what does that do for their quality of their work? Or the overall quality of academic writing? 

We will not only mark students down for insufficient understanding, we will also ping them for using AI in their written work. The institutions I teach at require students to declare where and how they have used AI in their work. Lingard notes that academic publications are stating "that ChatGPT cannot be a co-author because it cannot take responsibility for the work, and they require that researchers document any use of ChatGPT in their Methods or Acknowledgements sections" (2023, p. 261). 

Just as I and others have noticed, Rasul et al. (2023, p. 3) point out:

That, yes, "ChatGPT can act as a research assistant, answering users’ questions based on the related literature [...], analysing data [, ...] serve as a writing assistant [,...] and provide writing support". However, "users should exercise caution as ChatGPT may be prone to hallucinations (Alkaissi & McFarlane, 2023) and fabricate references and quotes (Sallam, 2023; Shen et al., 2023)".

I continue to be concerned about AI. It needs to get much, much better before it can become a useful, reliable and valid tool.


Sam

References:

Lingard, L. (2023). Writing with ChatGPT: An illustration of its capacity, limitations & implications for academic writers. Perspectives on Medical Education, 12(1), 261-270.  https://doi.org/10.5334/pme.1072

Rasul, T., Nair, S., Kalendra, D., Robin, M., de Oliveira Santini, F., Ladeira, W. J., ... & Heathcote, L. (2023). The role of ChatGPT in higher education: Benefits, challenges, and future research directions. Journal of Applied Learning and Teaching, 6(1), 1-16. https://doi.org/10.37074/jalt.2023.6.1.29

read more "AI Hallucination"

Monday, 4 September 2023

The trustworthy qualitative data tangle

When evaluating our qualitative research data - our evidence - we need to consider what our standards for quality will be... but the terms 'validity' and 'reliability' are not necessarily terms which are used with qualitative approaches.

That is not to say that qualitative researchers should not seek validity and reliability in our research work: it is to say that qualitative terms, meanings and methods seem differ from those used in quantitative work. Validity is a term which Bazeley uses (2013, p. 58, citing Maxwell, 2013), in outlining five key elements required for a sound qualitative research design as "purpose, conceptual framework, research questions, methods, and approach to validity", later implying that 'internal validity' is synonymous with trustworthiness. The biggest proponents of the shift in terminology are Guba & Lincoln (1985, 1989, 2017), who get quite excited about trustworthiness. Further, Bazeley implies that trustworthiness may have a different meaning to validity as she also mentions that qualitative research requires "rigour, reliability, credibility, trustworthiness, and validity" (Bazeley, 2021, p. 492).

So what do these terms mean in qualitative research? Thus far it seems that qualitative researchers are yet to come to a cohesive agreement on what those terms are, and how to define them in the field (Morse, 2015b). In fact, it becomes a tangled web, as illustrated here:

  • Rigour. Used synonymously with trustworthiness "to evaluate the credibility, transferability, dependability" of research outcomes (Morse, 2015b, p. 1212, citing Guba & Lincoln, 1989). First tangle.
  • Reliability. Or 'dependability': "Attainable through credibility, the use of “overlapping methods” (triangulation), “stepwise replication” (splitting data and duplicating the analysis), and use of an “inquiry audit” (p. 317) or audit trail" (Morse, 2015b, p. 1212, citing Guba & Lincoln, 1985, 1989). Second tangle with credibility.
  • Credibility. Third tangle. We can see from above that credibility is related to rigour. This is also used synonymously with quality (Corbin & Strauss, 2014). Fourth tangle: credibility implies 'trustworthiness' in findings where they accurately reflect participant and researcher experiences (Corbin & Strauss, 2014). This can be termed "internal validity [...]: [p]rolonged engagement, persistent observation, triangulation, peer debriefing, negative case analysis, referential adequacy, and member checks" (Morse, 2015b, p. 1212, citing Guba & Lincoln, 1989). Fifth tangle (with validity).
  • Trustworthiness. Apparently synonymous with rigour. OK... sixth tangle.
  • Validity. Aka 'transferability' or 'generalisability'; and really related to external validity. "Thick description is essential for 'someone interested' (p. 316) to transfer the original findings to another context, or individuals" (Morse, 2015b, p. 1212-13, citing Guba & Lincoln, 1989)

Morse (2015b) gets stroppy and points out the idiocy of all this, suggesting that we just step back into the current 'quantitative' terminology and stop all this faffing about:

  1. Validity. "(or internal validity) is usually defined as, the 'degree to which inferences made in a study are accurate and well-founded' (Polit & Beck, 2012, p. 745). Miller (2008b) further defined it as 'the ‘goodness’ or ‘soundness’ of a study' (p. 909). In qualitative inquiry, this is usually 'operationalized' by how well the research represents the actual phenomenon" (Morse, 2015b, p. 1213)
  2. Reliability. "'broadly described as the dependability, consistency, and/or repeatability of a project’s data collection, interpretation, and/or analysis' (Miller, 2008a, p. 745). Basically, it is the ability to obtain the same results if the study were to be repeated" (Morse, 2015b, p. 1213)
  3. Generalisability. "(or external validity) is 'extending the research results, conclusions, or other accounts that are based on the study of particular individuals, setting, times or institutions, to other individuals, setting, times or institutions than those directly studied' (Maxwell & Chmiel, 2014; Polit & Beck, 2012). In qualitative inquiry, the application of the findings to another situation or population is achieved through decontextualization and abstraction of emerging concepts and theory" (Morse, 2015b, p. 1213)
  4. Rigour. If the three elements above are present, then the study should be rigorous. "Both criteria of reliability and validity are intended to make qualitative research rigorous (formerly referred to as trustworthy). Through particular representation, abstraction, and theory development, validity enables qualitative theories to be generalizable and useful when recontextualized and applied to other settings. Reliability makes replication possible, although qualitative researchers themselves recognize induction is difficult" (Morse, 2015b, p. 1213).

It is all research, after all, whether it be qualitative or quantitative. Why not expand the definitions so that they cover both fields?


Sam

References:

Bazeley, P. (2013). Qualitative Data Analysis: Practical strategies (1st ed.). SAGE Publications Ltd.

Bazeley, P. (2021). Qualitative Data Analysis: Practical strategies (2nd ed.). SAGE Publications Ltd.

Corbin, J., & Strauss, A. (2014). Basics of qualitative research techniques and procedures for developing grounded theory (4th ed.). SAGE Publications Ltd.

Denzin, N. K., & Lincoln, Y. S. (Eds.) (2017). The SAGE Handbook of Qualitative Research (5th ed.). SAGE Publications, Inc.

Guba, E. G., & Lincoln, Y. S. (1985). Naturalistic Inquiry. Sage Publications, Inc.

Guba, E. G., & Lincoln, Y. S. (1989). Fourth Generation Evaluation. Sage Publications, Inc.

Morse, J. M. (2015a). Analytic Strategies and Sample Size. Qualitative Health Research, 25(10), 1317-1318. https://doi.org/10.1177/1049732315602867

Morse, J. M. (2015b). Critical Analysis of Strategies for Determining Rigor in Qualitative Inquiry. Qualitative Health Research, 25(9), 1212–1222. https://doi.org/10.1177/104973231558850

read more "The trustworthy qualitative data tangle"

Monday, 18 April 2022

Evidence and MBTI

Last year I read Merv Emre's book on the Myers-Briggs Type Indicator (2018), which was a fascinating read. I had expected a clear-eyed exploration of the type indicator, and I think the resulting book was fair in its approach to the story of the development of the MBTI tool. Ms Emre's book delivered more than a hint of admiration for the doggedness and drive of the tool's founders: Katharine Briggs, and her daughter, Isabel Briggs-Myers (2018).

This book enlightened me in a number of ways. For example, I had not realised the considerable amount of work undertaken between 1957 to 1975 from the Educational Testing Service (ETS) in attempting to validate the MBTI tool, which was pretty much an epic fail. It was only after ETS had given up on validation and returned the rights to Katharine Briggs - who by that time had dementia - and her daughter Isobel Myers-Briggs that the tool actually took off. The rise and rise of MBTI was via a deal with Consulting Psychologists Press (CPP) in California, which still markets the tool (Emre, 2018).

I found the chapters on the issues with validation to be very enlightening. There were clearly issues with replicability, construct and content validity, and generalisability. Isabel Myers-Briggs was reported as having closely held the process of scoring the returned, completed test scripts, appearing reluctant to allow computer scoring of data ...which was likely to have been more independent and objective (Emre, 2018). Other sources also mention issues of validity: MBTI's "ipsative construction of the questions, lack of criterion-related validity, and its tendency to be a 'feel good' instrument" are common complaints (Wood & Hay, 2013, p. 54). Test re-test replicability is as low as 50%: half those who have taken the test will fail to get the same result when retaken after three months (Emre, 2018; McCaully & Moody, 2007; Menand, 2018).

Despite my management and career background - or perhaps because of it! - I am a personality instrument skeptic. I am happy to use a personality instrument as a starting point for a career conversation: but not as an end point to categorise people; to put them in a box; to confine them.

Further, I like to have sound evidence of reliability, validity and generalisability, which many tests cannot provide; including MBTI (Mastrangelo, 2001). For me, my MBTI instrument type test re-test has always been consistent - unlike RIASEC, where my test-re-test is consistently inconsistent (I suspect that my interests are too broad to return a meaningful RIASEC score). What is interesting though is that - in light of Ms Emre's book (2018) - when I reread my MBTI ENTJ type profile, the actual type wording now seems to possess the vagueness of a horoscope:

"The ENTJ personality type is a competitive, highly motivated and focused person who sees just about everything by focusing on the bigger picture. ENTJs thrive by setting long-term goals and making highly analytical decisions, and they often do well in high-stress leadership roles. ENTJ types tend to see things in black and white, or by the numbers. In personal relationships they are fair, measured, and supportive" (MBTI Online, 2021).

So... hang on a minute. Competitive AND highly analytical AND fair AND supportive AND black and white AND big picture... right. Some of these things are not like the others.

What is particularly troubling is that MBTI has grown so popular that it is now blindly being used for making decisions in organisations; to determine who manages; who is in the team; who is hired; and who is let go (Macabasco, 2021). Additionally, MBTI also appears to be being used the classroom like the debunked learning styles (here) to categorise how different students learn (Emre, 2018, 2020; Harel, 2021; Macabasco, 2021).

We should not be making such decisions without good quality evidence. And - in my view - MBTI does not yet provide good enough quality evidence.

An HBO documentary documentary, "Persona: The Dark Truth Behind Personality Tests", was released on HBO last year (Macabasco, 2021). While I have not yet seen this film, it apparently explores the pervasiveness of personality testing use in organisational HR decisions. I can't wait to see it, but I am sure it will be somewhat depressing viewing. Ah well.


Sam

References:

Emre, M. (2018). The Personality Brokers: The Strange History of Myers-Briggs and the Birth of Personality Testing. Penguin Random House.

Emre, M. (23 December 2020). TED-Ed: Do personality tests work? [video]. https://youtu.be/lN7Fmt1i5TI

Harel, A. (29 April 2021). The Problems With Using Personality Tests For Hiring. Vervoe. https://vervoe.com/personality-tests-hiring/

HBO Max (23 February 2021). Persona | Official Trailer [video]. https://youtu.be/XWBXniurrA0

Macabasco, L. W. (4 March 2021). 'They become dangerous tools': the dark side of personality tests. The Guardian. https://www.theguardian.com/tv-and-radio/2021/mar/03/they-become-dangerous-tools-the-dark-side-of-personality-tests

Mastrangelo, P. M. (2001). Myers-Briggs Type Indicator [Form M]. In B. S. Plake & J. C. Impara (Eds.) The Fourteenth Mental Measurements Yearbook (pp. 818-820). Buros Center for Testing.

MBTI Online (2021). ENTJ. https://www.mbtionline.com/en-US/MBTI-Types/ENTJ

McCaully, M. H. & Moody, R. A. (2007) Multicultural applications of the Myers-Briggs Type Indicator, in L. A. Suzuki:, J. G. Ponterotto, & P. J. Meller (Eds.), Handbook of Multicultural Assessment: Clinical, psychological, and educational applications (3rd ed., pp. 402-424). Jossey-Bass.

Menand, L. (10 September 2018). What Personality Tests Really Deliver. The New Yorker. https://www.newyorker.com/magazine/2018/09/10/what-personality-tests-really-deliver

read more "Evidence and MBTI"

Friday, 2 April 2021

Methodolatry

I read a new term the other day: methodolatry (Bazeley, 2021). It was coined by Valerie Janesick (1994, 2000), a portmanteau word of methodology and idolatry, meaning:

"a preoccupation with selecting and defending methods to the exclusion of the actual substance of the story being told. Methodolatry is the slavish attachment and devotion to method that [... may overtake the research]. In my lifetime I have witnessed an almost constant obsession with the trinity of validity, reliability, and generalizability" (1994, p. 215)

"at its worst [, it] is found in the cases of survey researchers who throw out survey responses that don’t match the answers they are looking for in order that they might do the only statistical techniques they were taught (and this is aside from the ethical issues raised by such practices)" (2000, p. 390).

Ouch. Methodolatry is a great term: instead of letting form follow function, wedging the function into the form. It can be tempting to do that. We need to understand our ontology and epistemology. We need to have performed a broad literature review to understand what other researchers have done in the field. We need to be guided by the voices of experts, but not confined by them. We need to let our data, our research question, our 'natural' inclinations, and our skills lead our methods.

I was also interested in was Janesick's mention of validity, reliability and generalisability: three quality measures which don't really apply to qualitative data. However, Janesick further explores these, saying:

  1. Validity: "a set of technical microdefinitions of which the reader is most likely well aware. Validity in qualitative research has to do with [...] is the explanation credible?" (1994, p. 2016). Janesick draws on Donmoyer (1990) in suggesting that the likelihood of only one "correct" meaning is pretty thin.
  2. Reliability: Janesick implies that it is reliability which is verified by either participant or independent data cross-checks. While she thinks this is useful in publicly funded research, she doesn't think that these checks add any value to research (1994).
  3. Generalisability: Again, drawing on Donmoyer, Janesick (1994) says that generalisability is a flawed principle in qualitative research, but does not explain further. Donmoyer (1990, p. 178) suggests that "human action is constructed, not caused, and that to expect Newton-like generalizations describing human action [...] is to engage in a process akin to 'waiting for Godot'." Generalisability turns on focus: social science qualitative research focuses on a few carefully interrogated cases; physical sciences focus on many. "Determining where a particular leaf would land when it falls off a tree would be a task no less complex" (p. 178) for physical scientists than for social scientists... except the physical scientist is not interested in the leaf, but in the forest.

This last point may be why some researchers avoid the word generalisability, instead using 'trustworthiness' as a marker within qualitative research (Denzin & Lincoln, 2018). Some also vehemently oppose the change (Bazeley, 2021; Morse, 2015).

A topic for another day!


Sam

References:

  • Bazeley, P. (2021). Qualitative Data Analysis: Practical strategies (2nd ed.). SAGE Publication Ltd.
  • Denzin, N. K. & Lincoln, Y. S. (Eds.) (2018). The SAGE Handbook of Qualitative Research (5th ed.). SAGE Publications, Inc.
  • Donmoyer, R. (1990). Generalizability and the single case study. In E. W. Eisner & A. Peshkin (Eds.), Qualitative inquiry in education: The continuing debate (pp. 175-200). Teachers College Press.
  • Janesick, V. J. (1994). Chapter 12: The Dance of Qualitative Research Design: Metaphor, Methodolatry, and Meaning. In N. K. Denzin & Y. S. Lincoln (Eds.) The SAGE handbook of qualitative research (1st ed., pp. 209–219). SAGE Publications, Inc.
  • Janesick, V. J. (2000). Chapter 13: The choreography of qualitative research design: Minuets, improvisations, and crystallization. In N. K. Denzin & Y. S. Lincoln (Eds.) The SAGE handbook of qualitative research (2nd ed., pp. 397–400). SAGE Publications, Inc.
  • Morse, J. M. (2015). Analytic strategies and sample size. Qualitative Health Research, 25(10), 1317-1318. https://doi.org/10.1177/1049732315602867

read more "Methodolatry"

Sunday, 28 July 2013

Accurate Performance Predictors?

As a member of the HRINZ LinkedIn group, I was reading a member post the other day that was really fascinating. Posted by Anna Sage of Sage Advice in Wellington, it detailed the predictive validities of a variety of hiring tools:
  1. Assessment centres - potential (0.53)
  2. Ability tests - job performance and training (0.50)
  3. Structured interviews (0.44)
  4. Bio-data (0.37)
  5. Assessment centres - performance (0.36)
  6. Personality tests (0.33)
  7. Unstructured interviews (0.33)
  8. References (0.17)
  9. Self-assessments (0.15)

All pretty poor, really, at predicting success - performance - on the job! It amazes me that we still use references, if they are less useful than 1 in 5 of being accurate. In fact, why on earth we use anything from Bio-data on down is pretty moot. CVs don't even get a rating.

But what really surprised me was the follow up list that Anna posted; her "what is most popular" hiring assessment tools with employers (in decreasing order of popularity, based on some research Anna did between 1991 and 2006):
  1. References - 93% (predictive validity 0.17)
  2. Structured panel interviews - 88% (predictive validity 0.44)
  3. Structured one-to-one interviews - 85% (predictive validity 0.44)
  4. Competency-based interviews - 85%
  5. Ability tests - 75% (predictive validity 0.50)
  6. CVs - 74%
  7. Personality questionnaires - 60% (predictive validity 0.33)
  8. Assessment centres - 48% (predictive validity 0.53 or 0.36)
  9. Online selection tests - 25% (predictive validity 0.15)
  10. Bio-data - 7% (predictive validity 0.37)
If Anna's data is accurate, then why do employers and recruitment agencies still request the same old materials and hire as they do? It beggars belief.



Sam
read more "Accurate Performance Predictors?"