Overview of the TREC 2022 Fair Ranking Track
Michael D. Ekstrand, Graham McDonald, Amifa Raj, Isaac Johnson
Introduction
The TREC Fair Ranking Track aims to provide a platform for participants to develop and evaluate novel retrieval algorithms that can provide a fair exposure to a mixture of demographics or attributes, such as ethnicity, that are represented by relevant documents in response to a search query. For example, particular demographics or attributes can be represented by the documents’ topical content or authors.
The 2022 Fair Ranking Track adopted a resource allocation task. The task focused on supporting Wikipedia editors who are looking to improve the encyclopedia’s coverage of topics under the purview of a WikiProject.https://en.wikipedia.org/wiki/WikiProject WikiProject coordinators and/or Wikipedia editors search for Wikipedia documents that are in need of editing to improve the quality of the article. The 2022 Fair Ranking track aimed to ensure that documents that are about, or somehow represent, certain protected characteristics receive a fair exposure to the Wikipedia editors, so that the documents have an fair opportunity of being improved and, therefore, be well-represented in Wikipedia. The under-representation of particular protected characteristics in Wikipedia can result in systematic biases that can have a negative human, social, and economic impact, particularly for disadvantaged or protected societal groups .
Task Definition
The 2022 Fair Ranking Track used an ad hoc retrieval protocol. Participants were provided with a corpus of documents (a subset of the English language Wikipedia) and a set of queries. A query was of the form of a short list of search terms that represent a WikiProject. Each document in the corpus was relevant to zero to many WikiProjects and associated with zero to many fairness categories.
There were two tasks in the 2022 Fair Ranking Track. In each of the tasks, for a given query, participants were to produce document rankings that are:
Provide a fair exposure to articles that are associated to particular protected attributes.
The tasks shared a topic set, the corpus, the basic problem structure and the fairness objective. However, they differed in their target user persona, system output (static ranking vs. sequences of rankings) and evaluation metrics. The common problem setup was as follows:
Queries were provided by the organizers and derived from the topics of existing or hypothetical WikiProjects.
Documents were Wikipedia articles that may or may not be relevant to any particular WikiProject that is represented by a query.
Rankings were ranked lists of articles for editors to consider working on.
Fairness of exposure should be achieved with respect to the protected attributes associated with the documents. Documents can be associated to many different fairness attributes. The official track evaluation focused on intersectional fairness and, as such, evaluated how fairly systems rank documents with respect to all of the fairness categories. However, individual teams could choose whether to optimise their systems with respect to all, a subset of, or individual fairness categories.
The first task focused on WikiProject coordinators as users of the search system; their goal is to search for relevant articles and produce a ranked list of articles needing work that other editors can then consult when looking for work to do.
Output: The output for this task was a single ranking per query, consisting of 500 articles.
Evaluation was a multi-objective assessment of rankings by the following two criteria:
Relevance to a WikiProject topic. We will provide relevance assessments for the articles derived from existing Wikipedia data; Ranking relevance will be computed with nDCG, using binary relevance and logarithmic decay.
Fairness with respect to the exposure of different fairness categories associated to the articles returned in response to a query.
Section 4.2 contains details on the evaluation metrics.
2 Task 2: Wikipedia Editors
The second task focused on individual Wikipedia editors looking for work associated with a project. The conceptual model is that rather than maintaining a fixed work list as in Task 1, a WikiProject coordinator would create a saved search, and when an editor looks for work they re-run the search. This means that different editors may receive different rankings for the same query, and differences in these rankings may be leveraged for providing fairness.
Output: The output of this task is 100 rankings per query, each consisting of 20 articles.
Evaluation was a multi-objective assessment of rankings by the following three criteria:
Relevance to a WikiProject topic. We will provide relevance assessments for articles derived from existing Wikipedia data. Ranking relevance will be computed with nDCG.
Work needed on the article (articles needing more work preferred). We provide the output of an article quality assessment tool for each article in the corpus; for the purposes of this track, we assume lower-quality articles need more work.
Fairness with respect to the exposure of different fairness categories associated to the articles returned in response to a query.
The goal of this task was not to be fair to work-needed levels; rather, we consider work-needed and topical relevance to be two components of a multi-objective notion of relevance, so that between two documents with the same topical relevance, the one with more work needed is more relevant to the query in the context of looking for articles to improve.
This task used expected exposure to compare the exposure article subjects receive in result rankings to the ideal (or target) exposure they would receive based on their relevance and work-needed . This addresses fundamental limits in the ability to provide fair exposure in a single ranking by examining the exposure over multiple rankings.
For each query, participants provided 100 rankings, which we considered to be samples from the distribution realized by a stochastic ranking policy (given a query , a distribution over truncated permutations of the documents). Note that this is how we interpret the queries, but it did not mean that a stochastic policy is how the system should have been implemented — other implementation designs were certainly possible. The objective was to provide equitable exposure to documents of comparable relevance and work-needed, aggregated by protected attribute. Section 4.3 has details on the evaluation metrics.
Data
This section provides details of the format of the test collection, topics and ground truth. Further details about data generation and limitations can be found in Section 6.
The corpus and query data set is distributed via Globus, and can be obtained in two ways. First, it can be obtained via Globus, from our repository at https://boi.st/TREC2022Globus. From this site, you can log in using your institution’s Globus account or your own Google account, and synchronize it to your local Globus install or download it with Globus Connect Personal.https://www.globus.org/globus-connect-personal This method has robust support for restarting downloads and dealing with intermittent connections. Second, it can be downloaded directly via HTTP from: https://data.boisestate.edu/library/Ekstrand-2021/TRECFairRanking2022/.
The runs and evaluation qrels will be made available in the ordinary TREC archives.
2 Corpus
The corpus consisted of articles from English Wikipedia. We removed all redirect articles, but left the wikitext (markup Wikipedia uses to describe formatting) intact. This was provided as a JSON file, with one record per line, and compressed with gzip (trec_corpus.json.gz). Each record contains the following fields:
The unique numeric Wikipedia article identifier.
The article URL, to comply with Wikipedia licensing attribution requirements.
The three available formats of the corpus are as follows:
The full article text, with Wiki markup (text file only)
The full article text, without Wiki markup (plain file only)
The full article text, rendered into HTML (html file only)
The contents of this corpus were prepared in accordance with, and licensed under, the CC BY-SA 3.0 license.https://creativecommons.org/licenses/by-sa/3.0/ The raw Wikipedia dump files used to produce this corpus are available in the source directory; this is primarily for archival purposes, because Wikipedia does not publish dumps indefinitely.
3 Queries
The queries are in the 2022 directory, in the file train_topics_meta.jsonl. Each of the queries map to a single Wikiproject. The queries are constructed from extracted keywords from articles that are relevant to a Wikiproject. The following fields are provided:
A collection of search keywords forming the query text (list of str). We cleaned and parsed the Wiki articles and then used KeyBert to extract the most representative words of those articles. For each Wikiproject, we aggregated the extracted keywords from relevant articles and, after some manual filtering, used those as query texts for that particular Wikiproject.
The URL for the Wikiproject. This is provided for attribution and not expected to be used by your system as it will not be present in the evaluation data (string)
A list of the page IDs of relevant pages (list of int)
The keywords are the primary query text. The scope is there to provide some additional context and potentially support techniques for refining system queries.
In addition to query relevance, for Task 2: Wikipedia Editors (Section 2.2), participants will also be expected to return relevant documents that need more editing work done more highly than relevant documents that need less work done.
4 Fairness Categories
Fairness ground truth labels for the following fairness categories are also in the 2022 directory, in the trec_2022_articles_discrete.json.gz file. While we provide the raw values for each fairness category with the data, for most categories we also map the raw values to a reduced, fixed set of categories that will be used to judging systems.For more information, see: https://public.paws.wmcloud.org/User:Isaac_(WMF)/TREC/TREC_2022_Data.ipynb
The geographical location associated with the article topic. Both the associated countries—e.g., United Kingdom—and sub-continental regions—e.g., Northern Europe—are provided but systems will be evaluated using sub-continental regions (and not countries). An article can have 0 to many regions associated with it.
The geographic location associated with the article based on the article’s sources. Same categories as article geographic location above.
The gender of the individual about which the biography pertains. Gender has been reduced to four distinct categories: Man, Woman, Non-binary, and Unknown (missing data or not a biography).
How old the subject of the article is. For example, the birth date of a person in a biographical article, the date that an event occurred for articles that are about an event, or the creation date of a piece of art or music when the article is about the piece of art or music. The raw years are mapped to four distinct categories: Unknown, Pre-1900s, 20th century, and 21st century.
The occupation of the subject of an article. An article have 0 (unknown) to many occupations associated with it. There are 32 distinct occupation categories included in the data.
Editors often work through articles in alphabetical order and this can result in articles about subjects / topics that start with letters that appear earlier in the alphabet getting more exposure to the editors. Therefore, it is important that articles from later in the alphabet also get a fair exposure to the editors. The first letter is mapped to four discrete categories: a-d, e-k, l-r, and s-.
The length of time the article has existed. The date is mapped to one of four discrete categories: 2001-2006, 2007-2011, 2012-2016, and 2017-2022.
Number of times the page was viewed in February 2022. The number of pageviews are normalized and mapped to four discrete categories: Low, Medium-Low, Medium-High, and High.
The number of other language Wikipedias that the article is replicated in. This can range from English-only to all 300+ languages of Wikipedia but is mapped to three discrete categories: English only, 2-4 languages, and 5+ languages.
For the purposes of multidimensional fairness, we treated the dimensions as independent, and took the outer product of the fairness categories as a combined fairness space. The resulting space had over 11M dimensions, requiring care in implementing target alignments and metrics.
5 Metadata
We provide a simple Wikimedia quality score (a float between 0 and 1 where 0 is no content on the page and 1 is high quality) for optimizing for work-needed in Task 2. Work-needed can be operationalized as the reverse—i.e. 1 minus this quality score. The discretized quality scores will be used as work-needed for final system evaluation.
This data is provided together in a metadata file (trec_metadata.json.gz), in which each line is the metadata for one article represented as a JSON record with the following keys:
Continuous measure of article quality with 0 representing low quality and 1 representing high quality (float in range $$)
Discrete quality score in which the quality score is mapped to six ordinal categories from low to high: Stub, Start, C, B, GA, FA (string)
The group alignments associated to an article as described in Section 3.4.
6 Output
For Task 1, participants outputted results in rank order in a tab-separated file with two columns:
For Task 2, this file had 3 columns, to account for repeated rankings per query:
Evaluation Metrics
Each task was evaluated with its own metric designed for that task setting. The goal of these metrics was to measure the extent to which a system (1) exposed relevant documents, and (2) exposed those documents in a way that is fair to article topic groups, defined by the mentioned fairness constraints of the article’s subject.
This faces a problem in that Wikipedia itself has well-documented biases: if we target the current group distribution within Wikipedia, we will reward systems that simply reproduce Wikipedia’s existing biases instead of promoting social equity. However, if we simply target equal exposure for groups, we would ignore potential real disparities in topical relevance. Due to the biases in Wikipedia’s coverage, and the inability to retrieve documents that don’t exist to fill in coverage gaps, there is not good empirical data on what the distribution for any particular topic should be if systemic biases did not exist in either Wikipedia or society (the “world as it could and should be” ). Therefore, in this track we adopted a compromise: we averaged the empirical distribution of groups among relevant documents with the world population (for location) or equality (for gender) to derive the target group distribution.
Code to implement the metrics is found at https://github.com/fair-trec/trec2022-fair-public.
The tasks were to retrieve documents from a corpus that are relevant to a query . is a vector of relevance judgements for query . We denote a ranked list by ; is the document at position (starting from 1), and is the rank of document . For Task 1, each system returned a single ranked list; for Task 2, it returned a sequence of rankings .
In all metrics, we use log discounting to compute attention weights:
Task 2 also considered the work each document needs, represented by .
2 Task 1: WikiProject Coordinators (Single Rankings)
For the single-ranking Task 1, we adopted attention-weighted rank fairness (AWRF), first described by Sapiezynski et al. and named by Raj et al. . AWRF computes a vector of the cumulated exposure a list gives to each group, and a target vector ; we then compared these with the Jenson-Shannon divergence:
For Task 1, we ignored documents that are fully unknown for the purposes of computing and ; they do not contribute exposure to any group.
The resulting metric is in the range $$, with 1 representing a maximally-fair ranking (the distance from the target distribution is minimized). We combined it with an ordinary nDCG metric for utility:
To score well on the final metric , a run must be both accurate and fair.
3 Task 2: Wikipedia Editors (Multiple Rankings)
For Task 2, we used Expected Exposure to compare the exposure each group receives in the sequence of rankings to the exposure it would receive in a sequence of rankings drawn from an ideal policy with the following properties:
Relevant documents come before irrelevant documents
Relevant documents are sorted in nonincreasing order of work needed
Within each work-needed bin of relevant documents, group exposure is fairly distributed according to the average of the distribution of relevant documents and the distribution of global population (the same average target as before).
We have encountered some confusion about whether this task is requiring fairness towards work-needed; as we have designed the metric, work-needed is considered to be a part of (graded) relevance: a document is more relevant if it is relevant to the topic and needs significant work. In the Expected Exposure framework, this combined relevance is used to derive the target policies.
To apply expected exposure, we first define the exposure a document receives in sequence :
Our implementation rearranges the mean and aggregate operations, but the result is mathematically equivalent.
We then compare these system exposures with the target exposures for each query. This starts with the per-document ideal exposure; if is the number of relevant documents with work-needed level , then according to Diaz et al. the ideal exposure for document is computed as:
Since we include “unknown” as a group, we have a challenge with computing the target distribution by averaging the empirical distribution of relevant documents and the global population — global population does not provide any information on the proportion of relevant articles for which the fairness attributes are relevant. Our solution, therefore, is to average the distribution of known-group documents with the world population, and re-normalize so the final distribution is a probability distribution, but derive the proportion of known- to unknown-group documents entirely from the empirical distribution of relevant documents. Extended to handle partially-unknown documents, this procedure proceeds as follows:
Average the distribution of fully-known documents (both gender and location are known) with the global intersectional population (global population by location and equality by gender).
Average the distribution of documents with unknown location but known gender with the equality gender distribution.
Average the distribution of documents with unknown gender but known location with the world population.
The result is the target group exposure . We use this to measure the expected exposure loss:
Lower is better. It decomposes into two submetrics, the expected exposure disparity (EE-D) that measures overall inequality in exposure independent of relevance, for which lower is better; and the expected exposure relevance (EE-L) that measures exposure/relevance alignment, for which higher is better .
Results
This year 5 different teams submitted a total of 24 runs. All 5 teams participated in Task 1: Single Rankings (27 runs total), while 2 groups participated in Task 2: Multiple Rankings (11 runs total).
Relevance ranking by ColBERT-E2E and a heuristic approach that re-ranks to match target exposure using diversification.
Query rewriting strategy to expand query.
BM25 ranking from pyserini and pre-traned BERT for semantic score and re-ranked to fit target distribution at each ranking position by using greedy diversification, providing higher attention to protected group, and ensuring fairness in each position.
Relevance ranking using BM25 from pyserini and LambdaMART lerning-to-rank model and a multi-layer-perception to re-rank.
BM25 ranking from pyserini and weighted Reciprocal Ranking Fusion to diversify the ranking.
Table 1 shows the submitted systems ranked by the official Task 1 metric and its component parts nDCG and AWRF. Figure 1 plots the runs with the component metrics on the and axes. Unlike last year, we see less clustering of approaches from individual teams: two teams approaches had similar performance, while others are more scattered throughout the space.
We also computed fairness on individual dimension (Table 2 and Figure 2), and on the three subsets identified in Section 3.4 (Table 3 and Figure 3). For Task 1 the best-performing system overall also performed best on each individual category and subset; however, ordering of other systems changed between subsets or categories.
2 Task 2: Wikipedia Editors (Multiple Rankings)
Multi armed bandit strategies to select rankings from a pool of rankings considering each fairness category and observing the exposure and fairness-relevance relationship.
Epsilon-greedy with weighted ranking, epsilon-decay strategy, and randomisation are used in ranking selection process.
Table 4 shows the submitted systems ranked by the official Task 2 metric EE-L and its component parts EE-D and EE-R. Figure 4 plots the runs with the component metrics on the and axes. Overall, the submitted systems generally performed better for one of the component metrics than they did for the other.
We also computed fairness on individual dimension (Table 5 and Figure 5), and on the three subsets identified in Section 3.4 (Table 6 and Figure 6). Unlike Task 1, we see more difference in fairness between different single attributes and subsets.
Limitations
The data and metrics in this task address a few specific types of unfairness, and do so partially. This is fundamentally true of any fairness intervention, and does not in any way diminish the value of the effort — it is impossible for any data set, task definition, or metric to fully capture fairness in a universal way, and all data and analyses have limitations.
Some of the limitations of the data and task include:
Gender: For each Wikipedia article, we ascertain whether it is a biography, and, if so, which gender identity can be associated with the person it is about.Code: https://github.com/geohci/miscellaneous-wikimedia/blob/master/wikidata-properties-spark/wikidata_gender_information.ipynb This data is directly determined via Wikidata based on the instance-of property indicating the article is about a human (P31:Q5 in Wikidata terms) and then collecting the value associated with the sex-or-gender property (P21). Coverage here is quite high at 99.98% of biographies on Wikipedia having associated gender data on Wikidata.
Assigning gender identities to people is not a process without errors, biases, and ethical concerns. Applying the taxonomy developed by Pinney et al. to this work yields the following summary: the primary referent of the gender data is the subject; we do use a gender variable and it’s binary+other; gender determination is done via annotators (see details below); the gender data is used to measure bias and the goal is to audit system behavior. The process for assigning gender (annotation) is subject to some community-defined technical limitationshttps://www.wikidata.org/wiki/Property_talk:P21#Documentation and the Wikidata policy on living peoplehttps://www.wikidata.org/wiki/Wikidata:Living_people. While a separate project, English Wikipedia’s policies on gender identityhttps://en.wikipedia.org/wiki/Wikipedia:Manual_of_Style/Gender_identity likely inform how many editors handle gender; in particular, this policy explicitly favors the most recent reliably-sourced self-identification for gender, so misgendering a biography subject is a violation of Wikipedia policy; there may be erroneous data, but such data seems to be a violation of policy instead of a policy decision. Wikidata:WikiProject LGBT has documented some clear limitations of gender data on Wikidata and a list of further discussions and considerations.https://www.wikidata.org/wiki/Wikidata:WikiProject_LGBT/gender Since we are using gender data to calculate aggregate statistics, we judged these limitations to be less problematic than it would be if we were making decisions about individuals.
In our analysis (see Appendix A), we handle nonbinary gender identities by using four gender categories: unknown, male, female, and third.
We advise great care when working with the gender data, particularly outside the immediate context of the TREC task (either its original instance or using the data to evaluate comparable systems).
Geography: For each Wikipedia article, we also ascertained which, if any, countries and continents are relevant to the content.Code: https://github.com/geohci/wiki-region-groundtruth/blob/main/wiki-region-data.ipynb This was determined by directly looking up several community-maintained Wikidata structured data statements about the article. These properties were checked for the presence of countries, which were then mapped to continents via the United Nation’s geoscheme.https://en.wikipedia.org/wiki/United_Nations_geoscheme While this data must meet Wikidata’s verifiability guidelines,https://www.wikidata.org/wiki/Wikidata:Verifiability it does suffer from varying levels of incompleteness. For example, only 74% of people on Wikidata have a country of citizenship property.https://humaniki.wmcloud.org/gender-by-country Furthermore, structured data is itself limited—e.g., country of citizenship does not appropriately capture people who are considered stateless though these people may have many strong ties to a country. It is not easy to evaluate whether this data is missing at random or biased against certain regions of the world. Care should be taken when interpreting the absence of associated continents in the data. Further details can be found in the code repository.https://github.com/geohci/wiki-region-groundtruth
We also identify the associated countries and continents with the sources in the article. Each source is mapped to a country based on the URL or publisher associated with it. These mappings are built via a mixture of Wikidata, country extraction from whois records, and heuristics related to the top-level domain of the URL.See code for more details: https://github.com/geohci/geo-provenance This is the only inferred attribute used that is not maintained by Wikimedians and thus is much more likely to contain errors. Because this data was also incomplete, we had the assessors annotate an additional 15,000 items to help add to the data and better understand its quality. The feedback from assessors was that it was a difficult task—i.e. inferring country information for a generic website or publisher is often not easy and can be quite ambiguous at times. While most items were only assessed once, 101 publishers received multiple assessments with 85 (84%) of these in agreement and 32 URLs received multiple assessments with 26 (81%) of these in agreement. We also had a few items for which we had already inferred regions that we had the assessors check: only 6 publishers were checked but 5 (83%) were in agreement and 90 URLs were checked with 82 (91%) in agreement. While this leaves uncertainty about the publisher data, it does suggest that the URL data is reasonable quality because a few of those in disagreement appear to be assessor errors.
Age: We calculate the associated age of the subject of the article in a similar manner to article topic geography (extracting dates from several pre-determined properties on Wikidata).See the code for more details: https://gitlab.wikimedia.org/isaacj/miscellaneous-wikimedia/-/blob/master/wikidata-properties-spark/article-age.ipynb While geography is a multi-label feature, age is mapped to a single value via the median of the associated years and that median value is bucketed as pre-1900s, 20th century, or 21st century (and beyond). Like geography, it is not clear how many articles should have associated categories—e.g., many articles like those for plant species, do not clearly map to any specific time period—but it is safe to assume all biographies should have associated year data and we see 91.5% coverage suggesting relatively complete data.
Occupation: For each Wikipedia biography, we also ascertained which occupations could be associated with the person it is about. This data is directly determined via Wikidata by collecting the values associated with the occupation property (P106). For each occupation value, we then mapped it to one of 32 higher-order occupations based on the occupation ontology (using P279, the sub-class of, property for each occupation)For more details, see the code: https://gitlab.wikimedia.org/isaacj/miscellaneous-wikimedia/-/blob/master/wikidata-properties-spark/wikidata_occupation_taxonomy.ipynb. The 32 higher-order occupations were hand-selected to give sufficient detail while remaining a manageable number of categories. On English Wikipedia, 92.1% of biographies have at least one associated occupation value that could be mapped to the 32 higher-order occupations.
Popularity: For each article, we calculated how many pageviews it received in February 2022. These pageview counts are based on webrequest logshttps://wikitech.wikimedia.org/wiki/Analytics/Data_Lake/Traffic/Webrequest and filter out views from user-agents that explicitly identify themselves as spidersSee: https://meta.wikimedia.org/wiki/Research:Page_view and actors (shared user-agent and IP address) that seem to be automated in that they view more than 800 pages per hourhttps://wikitech.wikimedia.org/wiki/Analytics/Data_Lake/Traffic/BotDetection. These heuristics are not perfect however and traffic can be easily miscategorized if, for example, automated requests come from many different IPs or devices or actual users share a proxy that gives them the same IP and user-agent. The raw counts of pageviews were then converted into relative values between 0 and 1 by square-root transforming the value and normalizing to the 99th percentile of pageviews. Finally, these values were bucketed as [0 - 0.125), [0.125 - 0.250), [0.25 - 0.5), [0.5 - 1].
Sitelinks: The Wikimedia editor community maintains article sitelinks, or interlanguage links—i.e. explicit connection of articles about the same subject across language editions—via Wikidata. Almost all Wikipedia articles (99.92% for English)https://wikidata-analytics.wmcloud.org/app/WD_percentUsageDashboard have a corresponding Wikidata item and editors work to merge Wikidata items that are about the same subject so the sitelinks are aligned. Though there is no empirical data, it is generally accepted that most articles are appropriately linked to their corresponding other-language equivalents, especially in languages with shared scripts where simple approaches such as searching for an article title is often sufficient to identify matches.
Other: several fairness criteria are relatively straightforward and thus do not have many attached limitations. Specifically, the first letter of the article title (Alphabetical) and age of the article.
WikiProject Relevance: For the training queries, relevance was obtained from page lists for existing WikiProjects. While WikiProjects have broad coverage of English Wikipedia and we selected for WikiProjects that had tagged new articles in the recent months in the training data as a proxy for activity, it is certain that almost all WikiProjects are incomplete in tagging relevant content (itself a strong motivation for this task). While it is not easy to measure just how incomplete they are, it should not be assumed that content that has not been tagged as relevant to a WikiProject in the training data is indeed irrelevant.Current Wikiproject tags were extracted from the database tables maintained by the PageAssessments extension: https://www.mediawiki.org/wiki/Extension:PageAssessments
Work-needed: Our proxy for work-needed is a coarse proxy. It is based on just a few simple features (page length, sections, images, categories, links, and references) and does not reflect the nuances of the work needed to craft a top-quality Wikipedia article.For further details, see: https://meta.wikimedia.org/wiki/Research:Prioritization_of_Wikipedia_Articles/Language-Agnostic_Quality#V2 A fully-fledged system for supporting Wikiprojects would also include a more nuanced approach to understanding the work needed for each article and how to appropriately allocate this work.
Existing Article Bias: The task is limited to topics for which English Wikipedia already has articles. These tasks are not able to counteract biases in the processes by which articles come to exist (or are deleted )—recommending articles that should exist but don’t is an interesting area for future study.
Fairness constructs: we focus on several fairness constructs in this challenge as metrics for which there is high data coverage and a clear mechanism for which ”unfair” coverage might arise. That does not mean these are the most important constructs, but others—e.g., religion, sexuality, culture, race—generally are either more challenging to model or map to fairness goals .
References
Appendix A Page Alignments
This notebook computes the page alignments from the Wikipedia metadata. These are then used by the task-specific alignment notebooks to compute target distributions and page alignment subsets for retrieved pages.
Warning: this notebook takes quite a bit of memory to run.
import sysfrom pathlib import Pathimport pandas as pdimport xarray as xrimport numpy as npimport matplotlib.pyplot as pltimport seaborn as snsimport gzipimport jsonfrom natural.size import binarysize
reSet up progress bar and logging support:
from tqdm.auto import tqdmtqdm.pandas(leave=False)
import sys, logginglogging.basicConfig(level=logging.INFO, stream=sys.stderr)log = logging.getLogger('PageAlignments')
from wptrec.save import OutRepooutput = OutRepo('data/metric-tables')
A.2 Loading Data
We need a set of subregions that are folded into Oceania:
oc_regions = [ 'Australia and New Zealand', 'Melanesia', 'Micronesia', 'Polynesia',]
A.2.2 Page Data
Finally, we load the page metadata. This is a little manual to manage memory usage. Two memory usage tricks:
Use sys.intern for strings representing categoricals to decrease memory use
Bonus is that, through careful logic, we get a progress bar.
# META_FILE_TAG = 'discrete'META_FILE_TAG = 'discrete_assessed'
page_path = Path(f'data/trec_2022_articles_{META_FILE_TAG}.json.gz')page_file_size = page_path.stat().st_sizebinarysize(page_file_size)
Let’s define the different attributes we need to extract:
SUB_GEO_ATTR = 'page_subcont_regions'SRC_GEO_ATTR = 'source_subcont_regions'GENDER_ATTR = 'gender'OCC_ATTR = 'occupations'BASIC_ATTRS = [ 'page_id', 'first_letter_category', 'creation_date_category', 'relative_pageviews_category', 'num_sitelinks_category',]
Now, we’re going to process by creating lists we can reassemble with pd.DataFrame.from_records. We’ll fill these with tuples and dictionaries as appropriate.
qual_recs = []sub_geo_recs = []src_geo_recs = []gender_recs = []occ_recs = []att_recs = []seen_pages = set()
with tqdm(total=page_file_size, desc='compressed input', unit='B', unit_scale=True) as fpb: with open(page_path, 'rb') as gzf, gzip.GzipFile(fileobj=gzf, mode='r') as decoded: for line in decoded: line = json.loads(line) page = line['page_id'] if page in seen_pages: continue else: seen_pages.add(page) # page quality qual_recs.append((page, line['qual_cat'])) # page geography for geo in line[SUB_GEO_ATTR]: sub_geo_recs.append((page, sys.intern(geo))) # src geography psg = {'page_id': page} for g, v in line[SRC_GEO_ATTR].items(): if g == 'UNK': g = UNKNOWN psg[sys.intern(g)] = v src_geo_recs.append(psg) # genders for g in line[GENDER_ATTR]: gender_recs.append((page, sys.intern(g))) # occupations for occ in line[OCC_ATTR]: occ_recs.append((page, sys.intern(occ))) # other attributes att_recs.append(tuple((sys.intern(line[a]) if isinstance(line[a], str) else line[a]) for a in BASIC_ATTRS)) fpb.update(gzf.tell() - fpb.n) # update the progress bar
{"model_id":"7a8ed81f35ca4fa0b50c58c43638f3e0","version_major":2,"version_minor":0}
Now we will assemble these records into data frames.
quality = pd.DataFrame.from_records(qual_recs, columns=['page_id', 'quality'])
sub_geo = pd.DataFrame.from_records(sub_geo_recs, columns=['page_id', 'sub_geo'])sub_geo.info()
src_geo = pd.DataFrame.from_records(src_geo_recs)src_geo.info()
gender = pd.DataFrame.from_records(gender_recs, columns=['page_id', 'gender'])gender.info()
occupations = pd.DataFrame.from_records(occ_recs, columns=['page_id', 'occ'])occupations.info()
cat_attrs = pd.DataFrame.from_records(att_recs, columns=BASIC_ATTRS)cat_attrs.info()
all_pages = np.array(list(seen_pages))all_pages = np.sort(all_pages)all_pages = pd.Series(all_pages)
del src_geo_recs, sub_geo_recsdel gender_recs, occ_recsdel seen_pages
A.3 Helper Functions
These functions will help with further computations.
We are going to compute a number of data frames that are alignment vectors, such that each row is to be a multinomial distribution. This function normalizes such a frame.
def norm_align_matrix(df): df = df.fillna(0) sums = df.sum(axis='columns') return df.div(sums, axis='rows')
A.4 Page Alignments
All of our metrics require page ”alignments”: the protected-group membership of each page.
Quality isn’t an alignment, but we’re going to save it here:
output.save_table(quality, 'page-quality', parquet=True)
A.4.2 Page Geography
Let’s start with the straight page geography alignment for the public evaluation of the training queries. We’ve already loaded it above.
We need to do a little cleanup on this data:
Align pages with no known geography with ’@UNKNOWN’ (to sort before known categories)
reLet’s start by turning this into a wide frame:
sub_geo_align = sub_geo.assign(x=1).pivot(index='page_id', columns='sub_geo', values='x')sub_geo_align.fillna(0, inplace=True)sub_geo_align.head()
reNow we need to collapse Oceania into one column.
ocean = sub_geo_align.loc[:, oc_regions].sum(axis='columns')sub_geo_align = sub_geo_align.drop(columns=oc_regions)sub_geo_align['Oceania'] = ocean
reNext we need to add the Unknown column and expand this.
pSum the items to find total amounts, and then create a series for unknown:
sub_geo_sums = sub_geo_align.sum(axis='columns')sub_geo_unknown = ~(sub_geo_sums > 0)sub_geo_unknown = sub_geo_unknown.astype('f8')sub_geo_unknown = sub_geo_unknown.reindex(all_pages, fill_value=1)
reNow let’s join this with the original frame:
sub_geo_align = sub_geo_unknown.to_frame(UNKNOWN).join(sub_geo_align, how='left')sub_geo_align = norm_align_matrix(sub_geo_align)sub_geo_align.head()
sub_geo_align.sort_index(axis='columns', inplace=True)sub_geo_align.info()
reAnd convert this to an xarray for multidimensional usage:
sub_geo_xr = xr.DataArray(sub_geo_align, dims=['page', 'sub_geo'])sub_geo_xr
output.save_table(sub_geo_align, 'page-sub-geo-align', parquet=True)
A.4.3 Page Source Geography
We now need to do a similar setup for page source geography, which comes to us as a multinomial distribution already.
src_geo.set_index('page_id', inplace=True)
reExpand, then put 1 in UNKNOWN for everything that’s missing:
src_geo_align = src_geo.reindex(all_pages, fill_value=0)src_geo_align.loc[src_geo_align.sum('columns') == 0, UNKNOWN] = 1src_geo_align
ocean = src_geo_align.loc[:, oc_regions].sum(axis='columns')src_geo_align = src_geo_align.drop(columns=oc_regions)src_geo_align['Oceania'] = ocean
src_geo_align = norm_align_matrix(src_geo_align)
src_geo_align.sort_index(axis='columns', inplace=True)src_geo_align.info()
src_geo_xr = xr.DataArray(src_geo_align, dims=['page', 'src_geo'])src_geo_xr
output.save_table(src_geo_align, 'page-src-geo-align', parquet=True)
A.4.4 Gender
Now let’s work on extracting gender - this is going work a lot like page geography.
reNow, we’re going to do a little more work to reduce the dimensionality of the space. Points:
Cis/trans status is an adjective that can be dropped for the present purposes
The result is that we will collapse ”transgender female” and ”cisgender female” into ”female”.
The downside to this is that trans men are probabily significantly under-represented, but are now being collapsed into the dominant group.
pgcol = gender['gender']pgcol = pgcol.str.replace(r'(?:tran|ci)sgender\s+((?:fe)?male)', r'\1', regex=True)pgcol.value_counts()
reNow, we’re going to group the remaining gender identities together under the label ’NB’. As noted above, this is a debatable exercise that collapses a lot of identity.
gender_labels = [UNKNOWN, 'female', 'male', 'NB']pgcol[~pgcol.isin(gender_labels)] = 'NB'pgcol.value_counts()
reNow put this column back in the frame and deduplicate.
page_gender = gender.assign(gender=pgcol)page_gender = page_gender.drop_duplicates()
kg_mask = all_pages.isin(page_gender['page_id'])unknown = all_pages[~kg_mask]page_gender = pd.concat([ page_gender, pd.DataFrame({'page_id': unknown, 'gender': UNKNOWN})], ignore_index=True)page_gender
gender_align = page_gender.reset_index().assign(x=1).pivot(index='page_id', columns='gender', values='x')gender_align.fillna(0, inplace=True)gender_align = gender_align.reindex(columns=gender_labels)gender_align.head()
reLet’s see how frequent each of the genders is:
gender_align.sum(axis=0).sort_values(ascending=False)
gender_xr = xr.DataArray(gender_align, dims=['page', 'gender'])gender_xr
output.save_table(gender_align, 'page-gender-align', parquet=True)
A.4.5 Occupation
Occupation works like gender, but without the need for processing.
occ_align = occupations.assign(x=1).pivot(index='page_id', columns='occ', values='x')occ_align.head()
occ_unk = pd.Series(1.0, index=all_pages)occ_unk.index.name = 'page_id'occ_kmask = all_pages.isin(occ_align.index)occ_kmask.index = all_pagesocc_unk[occ_kmask] = 0occ_align = occ_unk.to_frame(UNKNOWN).join(occ_align, how='left')occ_align = norm_align_matrix(occ_align)occ_align.head()
occ_xr = xr.DataArray(occ_align, dims=['page', 'occ'])occ_xr
output.save_table(occ_align, 'page-occ-align', parquet=True)
A.4.6 Other Attributes
The other attributes don’t require as much re-processing - they can be used as-is as categorical variables. Let’s save!
pages = cat_attrs.set_index('page_id')pages
reNow each of these needs to become another table. The get_dummies function is our friend.
alpha_align = pd.get_dummies(pages['first_letter_category'])
output.save_table(alpha_align, 'page-alpha-align', parquet=True)
alpha_xr = xr.DataArray(alpha_align, dims=['page', 'alpha'])
age_align = pd.get_dummies(pages['creation_date_category'])output.save_table(age_align, 'page-age-align', parquet=True)
age_xr = xr.DataArray(age_align, dims=['page', 'age'])
pop_align = pd.get_dummies(pages['relative_pageviews_category'])output.save_table(pop_align, 'page-pop-align', parquet=True)
pop_xr = xr.DataArray(pop_align, dims=['page', 'pop'])
langs_align = pd.get_dummies(pages['num_sitelinks_category'])output.save_table(langs_align, 'page-langs-align', parquet=True)
langs_xr = xr.DataArray(langs_align, dims=['page', 'langs'])
A.5 Working with Alignments
At this point, we have computed an alignment matrix for each of our attributes, and extracted the qrels.
We will use the data saved from this in separate notebooks to compute targets and alignments for tasks.
Appendix B Task 1 Alignment
This notebook computes the target distributions and retrieved page alignments for Task 1. It depends on the output of the PageAlignments notebook.
reThis notebook can be run in two modes: ’train’, to process the training topics, and ’eval’ for the eval topics.
import sysimport warningsfrom collections import namedtuplefrom functools import reducefrom itertools import productimport operatorfrom pathlib import Path
import pandas as pdimport xarray as xrimport numpy as npimport matplotlib.pyplot as pltimport seaborn as snsimport gzipimport jsonfrom natural.size import binarysizefrom natural.number import number
reSet up progress bar and logging support:
from tqdm.auto import tqdmtqdm.pandas(leave=False)
import sys, logginglogging.basicConfig(level=logging.INFO, stream=sys.stderr)log = logging.getLogger('Task1Alignment')
from wptrec.save import OutRepooutput = OutRepo('data/metric-tables')
B.2 Data and Helpers
Most data loading is outsourced to MetricInputs. First we save the data mode where metric inputs can find it:
import wptrecwptrec.DATA_MODE = DATA_MODE
We want a function to join alignments with qrels:
def qr_join(align): return qrels.join(align, on='page_id').set_index(['topic_id', 'page_id'])
B.2.2 norm_dist
And a function to normalize to a distribution:
def norm_dist_df(mat): sums = mat.sum('columns') return mat.divide(sums, 'rows')
B.3 Prep Overview
Now that we have our alignments and qrels, we are ready to prepare the Task 1 metrics.
We’re first going to prepare the target distributions; then we will compute the alignments for the retrieved pages.
B.4 Subject Geography
Subject geography targets the average of the relevant set alignments and the world population.
qr_sub_geo_align = qr_join(sub_geo_align)qr_sub_geo_align
reFor purely geographic fairness, we just need to average the unknowns with the world pop:
qr_sub_geo_tgt = qr_sub_geo_align.groupby('topic_id').mean()qr_sub_geo_fk = qr_sub_geo_tgt.iloc[:, 1:].sum('columns')qr_sub_geo_tgt.iloc[:, 1:] *= 0.5qr_sub_geo_tgt.iloc[:, 1:] += qr_sub_geo_fk.apply(lambda k: world_pop * k * 0.5)qr_sub_geo_tgt.head()
output.save_table(qr_sub_geo_tgt, f'task1-{DATA_MODE}-sub-geo-target', parquet=True)
B.5 Source Geography
qr_src_geo_align = qr_join(src_geo_align)qr_src_geo_align
qr_src_geo_tgt = qr_src_geo_align.groupby('topic_id').mean()qr_src_geo_fk = qr_src_geo_tgt.iloc[:, 1:].sum('columns')qr_src_geo_tgt.iloc[:, 1:] *= 0.5qr_src_geo_tgt.iloc[:, 1:] += qr_src_geo_fk.apply(lambda k: world_pop * k * 0.5)qr_src_geo_tgt.head()
output.save_table(qr_src_geo_tgt, f'task1-{DATA_MODE}-src-geo-target', parquet=True)
B.6 Gender
Now we’re going to grab the gender alignments. Again, we ignore UNKNOWN.
qr_gender_align = qr_join(gender_align)qr_gender_align.head()
qr_gender_tgt = qr_gender_align.groupby('topic_id').mean()qr_gender_fk = qr_gender_tgt.iloc[:, 1:].sum('columns')qr_gender_tgt.iloc[:, 1:] *= 0.5qr_gender_tgt.iloc[:, 1:] += qr_gender_fk.apply(lambda k: gender_tgt * k * 0.5)qr_gender_tgt.head()
output.save_table(qr_gender_tgt, f'task1-{DATA_MODE}-gender-target', parquet=True)
B.7 Remaining Attributes
The remaining attributes don’t need any further processing, as they aren’t averaged.
qr_occ_align = qr_join(occ_align)qr_occ_tgt = qr_occ_align.groupby('topic_id').sum()qr_occ_tgt = norm_dist_df(qr_occ_tgt)qr_occ_tgt.head()
output.save_table(qr_occ_tgt, f'task1-{DATA_MODE}-occ-target', parquet=True)
qr_age_align = qr_join(age_align)qr_age_tgt = norm_dist_df(qr_age_align.groupby('topic_id').sum())output.save_table(qr_age_tgt, f'task1-{DATA_MODE}-age-target', parquet=True)
qr_alpha_align = qr_join(alpha_align)qr_alpha_tgt = norm_dist_df(qr_alpha_align.groupby('topic_id').sum())output.save_table(qr_alpha_tgt, f'task1-{DATA_MODE}-alpha-target', parquet=True)
qr_langs_align = qr_join(langs_align)qr_langs_tgt = norm_dist_df(qr_langs_align.groupby('topic_id').sum())output.save_table(qr_langs_tgt, f'task1-{DATA_MODE}-langs-target', parquet=True)
qr_pop_align = qr_join(pop_align)qr_pop_tgt = norm_dist_df(qr_pop_align.groupby('topic_id').sum())output.save_table(qr_pop_tgt, f'task1-{DATA_MODE}-pop-target', parquet=True)
B.8 Multidimensional Alignment
Now, we need to set up the multidimensional alignment. The basic version is just to multiply the targets, but that doesn’t include the target averaging we want to do for geographic and gender targets.
Doing that averaging further requires us to very carefully handle the unknown cases.
Define the averaged dimensions (with their background targets) and the un-averaged dimensions
Demonstrate the logic by working through the alignment computations for a single topic
Let’s define background distributions for some of our dimensions:
dim_backgrounds = { 'sub-geo': world_pop, 'src-geo': world_pop, 'gender': gender_tgt,}
reNow we’ll make a list of dimensions to treat with averaging:
DR = namedtuple('DimRec', ['name', 'align', 'background'], defaults=[None])avg_dims = [ DR(d.name, d.page_align_xr, xr.DataArray(dim_backgrounds[d.name], dims=[d.name])) for d in dimensions if d.name in dim_backgrounds][d.name for d in avg_dims]
raw_dims = [ DR(d.name, d.page_align_xr) for d in dimensions if d.name not in dim_backgrounds][d.name for d in raw_dims]
reNow: these dimension are in the original order - dimensions has the averaged dimensions before the non-averaged ones. This is critical for the rest of the code to work.
B.8.2 Demo
To demonstrate how the logic works, let’s first work it out in cells for one query (1).
qno = qrels['topic_id'].ilocqdf = qrels[qrels['topic_id'] == qno]qdf.name = qnoqdf
reWe can use these page IDs to get its alignments.
reWe’re now going to grab the dimensions that have targets, and create a single xarray with all of them:
q_xta = reduce(operator.mul, [d.align.loc[q_pages] for d in avg_dims])q_xta
reWe can similarly do this for the dimensions without targets:
q_raw_xta = reduce(operator.mul, [d.align.loc[q_pages] for d in raw_dims])q_raw_xta
reNow, we need to combine this with the other matrix to produce a complete alignment matrix, which we then will collapse into a query target matrix. However, we don’t have memory to do the whole thing at one go. Therefore, we will do it page by page.
q_tam = mean_outer(q_xta, q_raw_xta)q_tam
reIn 2021, we ignored fully-unknown for Task 1. However, it isn’t clear hot to properly do that with some attributes that are never fully unknown - they still need to be counted. Therefore, we consistently treat fully-unknown as a distinct category for both Task 1 and Task 2 metrics.
Before we average, we need to be able to select data by its known/unknown status.
Let’s start by making a list of cases - the known/unknown status of each dimension.
avg_cases = list(product(*[[True, False] for d in avg_dims]))avg_cases
reThe last entry is the all-unknown case - remove it:
reWe now want the ability to create an indexer to look up the subset of the alignment frame corresponding to a case. Let’s write that function:
def case_selector(case): def mksel(known): if known: # select all but 1st column return slice(1, None, None) else: # select 1st column return 0 return tuple(mksel(k) for k in case)
reFantastic! Given a case (known and unknown statuses), we can select the subset of the target matrix with exactly those.
Ok, now we have to - very carefully - average with our target modifier. For each dimension that is not fully-unknown, we average with the intersectional target defined over the known dimensions.
At all times, we also need to respect the fraction of the total it represents.
We’ll use the selection capabilities above to handle this.
First, let’s make sure that our target matrix sums to 1 to start with:
reFantastic. This means that if we sum up a subset of the data, it will give us the fraction of the distribution that has that combination of known/unknown status.
pFor each condition, we are going to proceed as follows:
Compute an appropriate intersectional background distribution (based on the dimensions that are ”known”)
Select the subset of the target matrix with this known status
Compute a normalization table such that each coordinate in the distributions to correct sums to 1 (so multiplying this by the background distribution spreads the background across the other dimensions appropriately), and use this to spread the background distribution
Average with the spread background distribution
Re-normalize to preserve the original sum
Let’s define the whole process as a function:
def avg_with_bg(tm, verbose=False): tm = tm.copy() tail_names = [d.name for d in raw_dims] # compute the tail mass for each coordinate (can be done once) tail_mass = tm.sum(tail_names) # now some things don't have any mass, but we still need to distribute background distributions. # solution: we impute the marginal tail distribution # first compute it tail_marg = tm.sum([d.name for d in avg_dims]) # then impute that where we don't have mass tm_imputed = xr.where(tail_mass > 0, tm, tail_marg) # and re-compute the tail mass tail_mass = tm_imputed.sum(tail_names) # and finally we compute the rescaled matrix tail_scale = tm_imputed / tail_mass del tm_imputed for case in avg_cases: # for deugging: get names known_names = [d.name for (d, known) in zip(avg_dims, case) if known] if verbose: print('processing known:', known_names) # Step 1: background bg = reduce(operator.mul, [ d.background for (d, known) in zip(avg_dims, case) if known ]) if not np.allclose(bg.sum(), 1.0): warnings.warn('background distribution for {} sums to {}, expected 1'.format(known_names, bg.values.sum())) # Step 2: selector sel = case_selector(case) # Steps 3: sum in preparation for normalization c_sum = tm[sel].sum() # Step 5: spread the background bg_spread = bg * tail_scale[sel] * c_sum if not np.allclose(bg_spread.sum(), c_sum): warnings.warn('rescaled background sums to {}, expected c_sum'.format(bg_spread.values.sum())) # Step 4 & 6: average with the background tm[sel] *= 0.5 bg_spread *= 0.5 tm[sel] += bg_spread if not np.allclose(tm[sel].sum(), c_sum): warnings.warn('target distribution for {} sums to {}, expected {}'.format(known_names, tm[sel].values.sum(), c_sum)) return tm
q_target = avg_with_bg(q_tam, True)q_target.sum()
print(number(q_target.values.size), 'values taking', binarysize(q_target.nbytes))
reWe can unravel this value into a single-dimensional array representing the multidimensional target:
reNow we have all the pieces to compute this for each of our queries.
B.8.3 Implementing Function
To perform this combination for every query, we’ll use a function that takes a data frame for a query’s relevant docs and performs all of the above operations:
def query_xalign(pages): # compute targets to average avg_pages = reduce(operator.mul, [d.align.loc[pages] for d in avg_dims]) raw_pages = reduce(operator.mul, [d.align.loc[pages] for d in raw_dims]) # convert to query distribution tgt = mean_outer(avg_pages, raw_pages) # average with background distributions tgt = avg_with_bg(tgt) # and return the result return tgt
B.8.4 Computing Query Targets
reNow with that function, we can compute the alignment vector for each query. Extract queries into a dictionary:
queries = { t: df['page_id'].values for (t, df) in qrels.groupby('topic_id')}
reMake an index that we’ll need later for setting up the XArray dimension:
q_ids = pd.Index(queries.keys(), name='topic_id')q_ids
reNow let’s create targets for each of these:
q_tgts = [query_xalign(queries[q]) for q in tqdm(q_ids)]
{"model_id":"d7cf659921754083b3c99d2a487c1f52","version_major":2,"version_minor":0}
reSave this to NetCDF (xarray’s recommended format):
output.save_xarray(q_tgts, f'task1-{DATA_MODE}-int-targets')
Appendix C Task 2 Alignment
This notebook computes the target distributions and retrieved page alignments for Task 2. It depends on the output of the PageAlignments notebook, as imported by MetricInputs.
reThis notebook can be run in two modes: ’train’, to process the training topics, and ’eval’ for the eval topics.
import sysimport operatorfrom functools import reducefrom itertools import productfrom collections import namedtuplefrom pathlib import Pathimport pandas as pdimport xarray as xrimport numpy as npimport matplotlib.pyplot as pltimport seaborn as snsimport gzipimport jsonfrom natural.size import binarysize
reSet up progress bar and logging support:
from tqdm.auto import tqdmtqdm.pandas(leave=False)
import sys, logginglogging.basicConfig(level=logging.INFO, stream=sys.stderr)log = logging.getLogger('Task2Alignment')
from wptrec.save import OutRepooutput = OutRepo('data/metric-tables')
from wptrec import metricsfrom wptrec.dimension import sum_outer
C.2 Data and Helpers
Most data loading is outsourced to MetricInputs. First we save the data mode where metric inputs can find it:
import wptrecwptrec.DATA_MODE = DATA_MODE
We want a function to join alignments with qrels:
def qr_join(align): return qrels.join(align, on='page_id').set_index(['topic_id', 'page_id'])
C.2.2 norm_dist
And a function to normalize to a distribution:
def norm_dist_df(mat): sums = mat.sum('columns') return mat.divide(sums, 'rows')
C.3 Work and Target Exposure
The first thing we need to do to prepare the metric is to compute the work-needed for each topic’s pages, and use that to compute the target exposure for each (relevant) page in the topic.
This is because an ideal ranking orders relevant documents in decreasing order of work needed, followed by irrelevant documents. All relevant documents at a given work level should receive the same expected exposure.
First, look up the work for each query page (’query page work’, or qpw):
qpw = qrels.join(page_quality, on='page_id')qpw
reAnd now use that to compute the number of documents at each work level:
qwork = qpw.groupby(['topic_id', 'quality'])['page_id'].count()qwork
reNow we need to convert this into target exposure levels. This function will, given a series of counts for each work level, compute the expected exposure a page at that work level should receive.
def qw_tgt_exposure(qw_counts: pd.Series) -> pd.Series: if 'topic_id' == qw_counts.index.names: qw_counts = qw_counts.reset_index(level='topic_id', drop=True) qwc = qw_counts.reindex(work_order, fill_value=0).astype('i4') tot = int(qwc.sum()) da = metrics.discount(tot) qwp = qwc.shift(1, fill_value=0) qwc_s = qwc.cumsum() qwp_s = qwp.cumsum() res = pd.Series( [np.mean(da[s:e]) for (s, e) in zip(qwp_s, qwc_s)], index=qwc.index ) return res
reWe’ll then apply this to each topic, to determine the per-topic target exposures:
qw_pp_target = qwork.groupby('topic_id').apply(qw_tgt_exposure)qw_pp_target.name = 'tgt_exposure'qw_pp_target
reWe can now merge the relevant document work categories with this exposure, to compute the target exposure for each relevant document:
qp_exp = qpw.join(qw_pp_target, on=['topic_id', 'quality'])qp_exp = qp_exp.set_index(['topic_id', 'page_id'])['tgt_exposure']qp_exp
C.4 Subject Geography
Subject geography targets the average of the relevant set alignments and the world population.
qr_sub_geo_align = qr_join(sub_geo_align)qr_sub_geo_align
reCompute a raw target, factoring in weights:
qr_sub_geo_tgt = qr_sub_geo_align.multiply(qp_exp, axis='rows').groupby('topic_id').sum()
reAnd now we need to average the known-geo with the background.
qr_sub_geo_fk = qr_sub_geo_tgt.iloc[:, 1:].sum('columns')qr_sub_geo_tgt.iloc[:, 1:] *= 0.5qr_sub_geo_tgt.iloc[:, 1:] += qr_sub_geo_fk.apply(lambda k: world_pop * k * 0.5)qr_sub_geo_tgt.head()
reThese are not distributions, let’s fix that!
qr_sub_geo_tgt = norm_dist_df(qr_sub_geo_tgt)
output.save_table(qr_sub_geo_tgt, f'task2-{DATA_MODE}-sub-geo-target', parquet=True)
C.5 Source Geography
qr_src_geo_align = qr_join(src_geo_align)qr_src_geo_align
qr_src_geo_tgt = qr_src_geo_align.multiply(qp_exp, axis='rows').groupby('topic_id').sum()
qr_src_geo_fk = qr_src_geo_tgt.iloc[:, 1:].sum('columns')qr_src_geo_tgt.iloc[:, 1:] *= 0.5qr_src_geo_tgt.iloc[:, 1:] += qr_src_geo_fk.apply(lambda k: world_pop * k * 0.5)qr_src_geo_tgt.head()
qr_src_geo_tgt = norm_dist_df(qr_src_geo_tgt)
output.save_table(qr_src_geo_tgt, f'task2-{DATA_MODE}-src-geo-target', parquet=True)
C.6 Gender
Now we’re going to grab the gender alignments. Works the same way.
qr_gender_align = qr_join(gender_align)qr_gender_align.head()
qr_gender_tgt = qr_gender_align.multiply(qp_exp, axis='rows').groupby('topic_id').sum()
qr_gender_fk = qr_gender_tgt.iloc[:, 1:].sum('columns')qr_gender_tgt.iloc[:, 1:] *= 0.5qr_gender_tgt.iloc[:, 1:] += qr_gender_fk.apply(lambda k: gender_tgt * k * 0.5)qr_gender_tgt.head()
qr_gender_tgt = norm_dist_df(qr_gender_tgt)
output.save_table(qr_gender_tgt, f'task2-{DATA_MODE}-gender-target', parquet=True)
C.7 Occupation
Occupation is more straightforward, since we don’t have a global target to average with. We do need to drop unknown.
qr_occ_align = qr_join(occ_align).multiply(qp_exp, axis='rows')qr_occ_tgt = qr_occ_align.iloc[:, 1:].groupby('topic_id').sum()qr_occ_tgt = norm_dist_df(qr_occ_tgt)qr_occ_tgt.head()
output.save_table(qr_occ_tgt, f'task2-{DATA_MODE}-occ-target', parquet=True)
C.8 Remaining Attributes
The remaining attributes don’t need any further processing, as they are completely known.
qr_age_align = qr_join(age_align).multiply(qp_exp, axis='rows')qr_age_tgt = norm_dist_df(qr_age_align.groupby('topic_id').sum())output.save_table(qr_age_tgt, f'task2-{DATA_MODE}-age-target', parquet=True)
qr_alpha_align = qr_join(alpha_align).multiply(qp_exp, axis='rows')qr_alpha_tgt = norm_dist_df(qr_alpha_align.groupby('topic_id').sum())output.save_table(qr_alpha_tgt, f'task2-{DATA_MODE}-alpha-target', parquet=True)
qr_langs_align = qr_join(langs_align).multiply(qp_exp, axis='rows')qr_langs_tgt = norm_dist_df(qr_langs_align.groupby('topic_id').sum())output.save_table(qr_langs_tgt, f'task2-{DATA_MODE}-langs-target', parquet=True)
qr_pop_align = qr_join(pop_align).multiply(qp_exp, axis='rows')qr_pop_tgt = norm_dist_df(qr_pop_align.groupby('topic_id').sum())output.save_table(qr_pop_tgt, f'task2-{DATA_MODE}-pop-target', parquet=True)
C.9 Multidimensional Alignment
Now let’s dive into the multidmensional alignment. This is going to proceed a lot like the Task 1 alignment.
Let’s define background distributions for some of our dimensions:
dim_backgrounds = { 'sub-geo': world_pop, 'src-geo': world_pop, 'gender': gender_tgt,}
reNow we’ll make a list of dimensions to treat with averaging:
DR = namedtuple('DimRec', ['name', 'align', 'background'], defaults=[None])avg_dims = [ DR(d.name, d.page_align_xr, xr.DataArray(dim_backgrounds[d.name], dims=[d.name])) for d in dimensions if d.name in dim_backgrounds][d.name for d in avg_dims]
raw_dims = [ DR(d.name, d.page_align_xr) for d in dimensions if d.name not in dim_backgrounds][d.name for d in raw_dims]
reNow: these dimension are in the original order - dimensions has the averaged dimensions before the non-averaged ones. This is critical for the rest of the code to work.
C.9.2 Data Subsetting
avg_cases = list(product(*[[True, False] for d in avg_dims]))avg_cases.pop()avg_cases
def case_selector(case): def mksel(known): if known: # select all but 1st column return slice(1, None, None) else: # select 1st column return 0 return tuple(mksel(k) for k in case)
C.9.3 Background Averaging
We’re now going to define our background-averaging function; this is reused from the Task 1 alignment code.
For each condition, we are going to proceed as follows:
Compute an appropriate intersectional background distribution (based on the dimensions that are ”known”)
Select the subset of the target matrix with this known status
Compute a normalization table such that each coordinate in the distributions to correct sums to 1 (so multiplying this by the background distribution spreads the background across the other dimensions appropriately), and use this to spread the background distribution
Average with the spread background distribution
Re-normalize to preserve the original sum
Let’s define the whole process as a function:
def avg_with_bg(tm, verbose=False): tm = tm.copy() tail_names = [d.name for d in raw_dims] # compute the tail mass for each coordinate (can be done once) tail_mass = tm.sum(tail_names) # now some things don't have any mass, but we still need to distribute background distributions. # solution: we impute the marginal tail distribution # first compute it tail_marg = tm.sum([d.name for d in avg_dims]) # then impute that where we don't have mass tm_imputed = xr.where(tail_mass > 0, tm, tail_marg) # and re-compute the tail mass tail_mass = tm_imputed.sum(tail_names) # and finally we compute the rescaled matrix tail_scale = tm_imputed / tail_mass del tm_imputed for case in avg_cases: # for deugging: get names known_names = [d.name for (d, known) in zip(avg_dims, case) if known] if verbose: print('processing known:', known_names) # Step 1: background bg = reduce(operator.mul, [ d.background for (d, known) in zip(avg_dims, case) if known ]) if not np.allclose(bg.sum(), 1.0): warnings.warn('background distribution for {} sums to {}, expected 1'.format(known_names, bg.values.sum())) # Step 2: selector sel = case_selector(case) # Steps 3: sum in preparation for normalization c_sum = tm[sel].sum() # Step 5: spread the background bg_spread = bg * tail_scale[sel] * c_sum if not np.allclose(bg_spread.sum(), c_sum): warnings.warn('rescaled background sums to {}, expected c_sum'.format(bg_spread.values.sum())) # Step 4 & 6: average with the background tm[sel] *= 0.5 bg_spread *= 0.5 tm[sel] += bg_spread if not np.allclose(tm[sel].sum(), c_sum): warnings.warn('target distribution for {} sums to {}, expected {}'.format(known_names, tm[sel].values.sum(), c_sum)) return tm
C.9.4 Computing Targets
We’re now ready to compute a multidimensional target. This works like the Task 1, with the difference that we are propagating work needed into the targets as well; the input will be series whose index is page IDs and values are the work levels.
def query_xalign(pages): # compute targets to average avg_pages = reduce(operator.mul, [d.align.loc[pages.index] for d in avg_dims]) raw_pages = reduce(operator.mul, [d.align.loc[pages.index] for d in raw_dims]) # weight the left pages pages.index.name = 'page' qpw = xr.DataArray.from_series(pages) avg_pages = avg_pages * qpw # convert to query distribution tgt = sum_outer(avg_pages, raw_pages) tgt /= qpw.sum() # average with background distributions tgt = avg_with_bg(tgt) # and return the result return tgt
C.9.5 Applying Computations
Now let’s run this thing - compute all the target distributions:
q_tgts = [query_xalign(qp_exp.loc[q]) for q in tqdm(q_ids)]
{"model_id":"825dff5cd101402e8910af2cb8a4abf7","version_major":2,"version_minor":0}
reSave this to NetCDF (xarray’s recommended format):
output.save_xarray(q_tgts, f'task2-{DATA_MODE}-int-targets')
C.10 Task 2B - Equity of Underexposure - NOT YET DONE
For 2022, we are using a diffrent version of the metric. Equity of Underexposure looks at each page’s underexposure (system exposure is less than target exposure), and looks for underexposure to be equitably distributed between groups.
On its own, this isn’t too difficult; averaging with background distributions, however, gets rather subtle. Background distributions are at the roup level, but we need to propgagate that into the page level, so we can compute the difference between system and target exposure at the page level, and then aggregate the underexposure within each group.
The idea of equity of underexposure is that we and . We then compute , and restrict it to be negative, and aggregate it by group; if is our page alignment matrix and , we compute the group underexposure by .
That’s the key idea. However, we want to use that has the equivalent of averaging group-aggregated with global target distributions . We can do this in a few stages. First, we compute the total attention of each group, and use that to compute the fraction of group global weight that should go to each unit of alignment:
\begin{align*} s_g & = \sum_d a_{dg} \ \hat{w}_g & = \frac{w_g}{s_g} \end{align*}
We’re going to reuse demo topic data from before:
reNow, let’s make a copy, and start building up a world target matrix that properly accounts for missing values:
reNow, let’s put in the known intersectional targets:
reNow we need the known-gender / unknown-geo targets:
W[0, 1:] = int_tgt.sum(axis=0) * W[0, 1:].sum()
reAnd the known-geo / unknown-gender targets:
W[1:, 0] = int_tgt.sum(axis=1) * W[1:, 0].sum()
reThe massive values are only where we have no relevant items, so they’ll never actually be used.
pWe can now compute the query-aligned target matrix.
qp_gt = (q_xa * (Wh * qp_exp.sum())).sum(axis=(1,2)).to_series()qp_gt.index.name = 'page_id'qp_gt
C.10.2 Setting Up Matrix
Now that we have the math worked out, we can create actual global target frames for each query.
def topic_page_tgt(qdf): pages = qdf['page_id'] pages = pages[pages.isin(page_xalign.indexes['page'])] q_xa = page_xalign.loc[pages.values, :, :] # now we need to get the exposure for the pages p_exp = qp_exp.loc[qdf.name] assert p_exp.index.is_unique # need our sums s_xg = q_xa.sum(axis=0) + 1e-10 # set up the global target W = s_xg / s_xg.sum() W[1:, 1:] = int_tgt * W[1:, 1:].sum() W[0, 1:] = int_tgt.sum(axis=0) * W[0, 1:].sum() W[1:, 0] = int_tgt.sum(axis=1) * W[1:, 0].sum() # per-unit global weights, de-normalized by total exposure Wh = W / s_xg Wh *= p_exp.sum() # compute global target gtgt = q_xa * Wh gtgt = gtgt.sum(axis=(1,2)).to_series() # compute average target avg_tgt = 0.5 * (p_exp + gtgt) avg_tgt.index.name = 'page' return avg_tgt
qp_tgt = qrels.groupby('id').progress_apply(topic_page_tgt)qp_tgt
save_table(qp_tgt.to_frame('target'), 'task2-all-page-targets')
train_qptgt = qp_tgt.loc[train_topics['id']].to_frame('target')eval_qptgt = qp_tgt.loc[eval_topics['id']].to_frame('target')
save_table(train_qptgt, 'task2-train-page-targets')save_table(eval_qptgt, 'task2-eval-page-targets')