1,720,965 research outputs found
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Gender
Abstract: Vrouwen en mannen kregen verschillende rollen binnen de Vlaamse beweging. De heersende verwachtingen spoorden echter niet altijd met de werkelijke genderverhoudingen en evolueerden bovendien door de tijd \u2013 parellel aan de veranderende invulling van het politieke project van de Vlaamse beweging
Towards a datafication of Antwerp street life? Co-creating a dataset of 100.000 pages of handwritten police reports (1876-1945)
Since the 1960s, when the so-called ‘new social historians’ started questioning mainstream historical narratives and experiences, new historical sources such as police and judicial sources came under scrutiny, putting the spotlight on social groups and ‘people without history’ who have left almost no paper trail of their own (Thompson 1963; Wolf 1982). These historical documents have been valuable for a myriad of approaches which are not necessarily crime-related: working-class history, the history of mentalities, gender and sexuality history, youth and family history, Alltagsgeschichte, and so on. Up until today, qualitative methodological approaches have dominated this research area, emphasising thick description and case-based research. Only recently, social historians have been discovering the potential of data-driven approaches towards criminal records (e.g. Van den Heuvel 2020; Saygi / Yasunaga 2021). One of the most straightforward explanations for this lag is the lack of machine-readable data. Nineteenth-century bureaucratisation processes have left historians of the modern period with an abundance of archival material, but only the tip of the iceberg is digitally available, let alone searchable. Grand digitisation initiatives such as Old Bailey Online remain scarce because they require a lot of time, money, and labour (Hitchcock et al. 2012). In response, crowdsourcing platforms have been popping up like mushrooms, but these are often ill-suited for researchers with limited means such as PhDs.
Over the past few years, the technique of Handwritten Text Recognition (HTR) – made accessible through e.g. Transkribus and eScriptorium – improved greatly, which broke new ground for small-scale research projects (Muehlberger et al. 2019; Kiessling et al. 2019). In this poster, I will present a dataset of approximately 100.000 HTRed pages of local police reports from the city of Antwerp (1876-1945). Apart from the dataset specifics, I will show the challenges for HTR-training, which relied on both gold standard expert transcriptions, and semi-gold standard material provided by undergraduate students. Both collections will be scrutinised, and solutions will be examined for bypassing noisy training data. I hope that this presentation not only encourages qualitative social historians to test their hypotheses on ‘big data’, but also inspires them to step by step unlock more underexplored paper archives for data-driven research.
Dataset specifics
Incident books
From the 1860s onwards Antwerp police officers were obliged to write a daily report after patrolling their respective neighbourhoods in the so-called ‘incident books’. One of the major historiographical benefits is the fact that these reports not only document prosecuted crimes, but non-prosecuted misdemeanours and other irregularities as well. Therefore, they give a very extensive account of all sorts of happenings on the Antwerp streets, such as neighbourhood quarrels, monkey tricks, lost children… The serial character of the incident books ensures a consistent registration of who (gender, profession, age, place of residence and birth) was when (date and time of day) and where (address) doing what (detailed description of incident).
Space and time
The dataset consists of 326 books which date from 1876 to 1945 and are located within 11 different police districts. The data is not spread evenly throughout space and time. Most pages contain police reports dedicated to the third and the sixth district (predominantly bourgeois districts), and the page number increases during the 1930s and the Second World War (a period of increasing police surveillance).
HTR-training results
The incident books were very challenging for HTR-training because they contain numerous different handwritings and are written in French as well as Dutch. Furthermore, the difficult and irregular layout of the reports made segmentation very time-consuming. I have worked with two different training sets: gold standard expert transcriptions, and semi-gold standard material provided by undergraduates. The former consists of 326 pages selected randomly from each incident book and the latter contains several samples (in total 3000 pages).
Despite the imperfect undergraduate work, the models trained on all the material performed best, but unfortunately only led up to a (still unsatisfactory) character error rate (CER) of 9.3%. After systematically analysing the word and character errors of this model, I managed to improve the CER up to 6.3% by eliminating skewed text and signature marks. Finally, I switched to a more relaxed evaluation of the model by lowercasing all text and excluding interpunction and redundant spacing characters.
The fully HTRed dataset consists of about 30,5 million words and will soon be published in open-access format.
References
Kiessling, Benjamin et al. (2019): “eScriptorium: An Open Source Platform for Historical Document Analysis”, in: International Conference on Document Analysis and Recognition Workshops (ICDARW), 2: 19–24.
Hitchcock, Tim et al. (2012): The Old Bailey Proceedings Online, 1674-1913, https://www.oldbaileyonline.org [4.11.2022];
Muehlberger, Guenter et al. (2019): “Transforming scholarship in the archives through handwritten text recognition: Transkribus as a case study”, I,: Journal of Documentation 75, 5: 954–976.
Saygi, Gamze / Yasunaga, Marie (2021): “The digital urban experience of a lost city using mixed methods to depict the historical street life of Edo/Tokyo”, in: Magazén 2, 2: 193-224.
Thompson, Edward Palmer (1963): The Making of the English Working Class. London: Victor Gollancz Ltd.
van den Heuvel, D. et al. (2020): “Capturing gendered mobility and street use in the historical city: a new methodological approach”, in: Cultural and Social History 17, 4: 515–536.
Wolf, Eric (1982): Europe and the People without History. Berkeley: University of California Press
'Den nieuwen mensch' als homo fascistus? ‘Fascisering’ van mannelijkheidsidealen op de Vlaams-nationalistische IJzerbedevaarten tijdens de late jaren dertig en WOII
De Eerste Wereldoorlog deed nationalistische bewegingen in Europa heropleven. De oorlog gaf bovendien een impuls aan nieuwe nationalistische bewegingen, zoals het Vlaams-nationalisme in België. De nationalistische herdenkingsculturen die ontstonden in de nasleep van de oorlog bevestigden het dominante beeld van oorlog als een mannelijk evenement en plaatsten mannelijkheidsidealen centraal in hun ideologie. De opkomst van het fascisme in de jaren dertig, als een ultranationalistische ideologie, maakte dat bestaande mannelijkheidsidealen tot een hoogtepunt werden opgevoerd. De verrechtsing van een deel van de Vlaams-nationalistische beweging in deze periode, doet de vraag rijzen naar de mate waarin ook Vlaams-nationalistische mannelijkheidsidealen tijdens de Tweede Wereldoorlog ‘gefasciseerd’ zijn
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
koamabayili/VECTRON-author-checklist: VECTRON author checklist
We have done our best to complete the author checklist relating to the use of animals in the hut study. Note that the objective for the hut study was to evaluate the IRS treatment applications for residual efficacy against Anopheles mosquitoes, including the local An. coluzzii mosquito population. Cows were only used to attract mosquitoes into the huts and no tests were carried out directly on the cows. The author checklist is intended for use with studies where experiments are carried out on animals, which is why we have had such difficulty in completing this for the hut study, as many of the questions do not relate to how the cows were used
- …
