DataverseNO
Not a member yet
2167 research outputs found
Sort by
Replication Data for: Diverging grammaticalization patterns across Spanish varieties: the case of 'perdón' in Mexican and Peninsular Spanish
Dataset description:
This dataset contains one data file (.csv) used to create the tables in the paper "Diverging grammaticalization patterns across Spanish varieties: the case of perdón in Mexican and Peninsular Spanish". The data used to investigate the contemporary uses of the apology marker perdón come from a sizeable, manually annotated corpus consisting of spoken spontaneous conversations and interviews, recorded during the last quarter of the 20th century and the first decades of the 21st century.
The following list of existing, spoken corpora were exploited: Corpus del Proyecto para el Estudio Sociolingüístico del Español de España y de América (PRESEEA), América y España español coloquial (AMERESCO), Corpus Sociolingüístico de la Ciudad de México (CSCM), Corpus del Habla de Baja California (CHBC) and Corpus Michoacano del Español (CME) for Mexican Spanish and Corpus Oral de Madrid (CORMA), Valencia Español Coloquial (Val.Es.Co.), Corpus Oral de Referencia del Español Contemporáneo (CORLEC), Corpus integrado de referencia en lenguas romances (C-ORAL-ROM) and Corpus del Proyecto para el Estudio Sociolingüístico del Español de España y de América (PRESEEA) for Peninsular Spanish.
The data file includes all occurrences of perdón ('sorry') together with its near-synonymous apologetic markers such as lo siento ('I am sorry') and forms derived from the performative verbs perdonar ('to forgive') and disculpar ('to apologize'). It contains a total of 769 occurrences: 363 cases for Mexican Spanish and 406 for Peninsular Spanish. The data is annotated for (i) Spanish variety, (ii) corpus, (iii) form, (iv) type of offense, (v) face affected (positive face/negative face) and (vi) orientation of the face (speaker/hearer).Article abstract:
This study investigates the contemporary grammaticalized uses of perdón (‘sorry’) in two varieties of Spanish, namely Mexican and Peninsular Spanish. Methodologically, the investigation is based on a taxonomy of offenses, organized around the concept of face and based on spoken data of Spanish from Mexico and Spain. This taxonomy turns out to be a fruitful methodological tool for the analysis of apologetic markers: it does not only offer usage-based evidence for previous theorizing concerning the grammaticalization process of apologetic markers, but also leads to a refinement of these previous results from a contrastive point of view. Evidence from both corpora suggests a more advanced stage in the grammaticalization process of perdón in Mexican Spanish, where it can be used not only as a self-face-saving device geared towards the positive face of the speaker, but also in turn-taking contexts oriented towards the negative face of the interlocutor. Peninsular Spanish, on the other hand, resorts to a more varied gamut of apologetic markers in these contexts. </p
Background data for: Sprachliches Place-Making. Eine sprachwissenschaftliche Analyse der diskursiven Konstruktion von Wissen über Raum
This dataset contains corpus statistical calculations that were used to investigate patterns of linguistic place-making in the German language. Patterns are defined here primarily as typical words and word combinations (collected as ngrams and collocations). Those were obtained from a text corpus specifically created for this analysis (not published). The text corpus has a size of 2,630,095 tokens and contains only texts in which one or more geographical places form the main topic. It turned out that these are primarily texts from tourism contexts. The corpus therefore includes 34 travel guides, 145 texts from place/travel blogs, 65 journalistic texts and 73 (city) marketing texts. The analyses were carried out with #LancsBox.
Initial analyses quickly made it clear that patterns that refer specifically to places can only be distinguished from other patterns if the corpus is semantically annotated. In this annotation process, nouns that refer to places were tagged (German: Placebezeichnungen, therefore abbreviated as PB). This makes it possible to evaluate their frequency and distribution in the corpus. On the other hand, it is possible to calculate their collocations. In other words: It is possible to calculate which words typically appear together with which words. Detailed analyses were then carried out for words with which place-referring nouns collocate particularly intensively (so-called collocation profiles). In this way, typical usage contexts can be determined and patterns of place-making can be identified.
The dataset contains data on the frequency, distribution and typical collocations of place-referring nouns, as well as the collocation profiles of words that collocate particularly frequently with place-referring nouns. The analysis and the resulting conclusions on the patterns of linguistic place-making can be found in the related publication.
Abstract of the related publication: Places are geographical areas or spots that are meaningful to people. The members of a discourse community can clearly identify places and name certain characteristics. The process of place-making, which has been theorized in sociology and geography, is thus – viewed from a linguistic perspective - a process of knowledge transfer: if people know something about a geographical area or spot, it becomes a place for them. This knowledge is negotiated and communicated linguistically in discourses.
The study deals with the questions of (a) which knowledge is relevant in place-making processes and (b) how it is verbalized. To answer these questions, linguistic patterns are collected in a representative corpus. In this process, common nouns that refer to places (so-called Placebezeichnungen) play a special role - both in terms of content and methodology. They are semantically annotated and form the anchor points of collocation analyses. </p
Replication Data for: Long term effects of forest management on forest structure and dead wood in mature boreal forests
This dataset contains data on forest structure, volumes of coarse woody debris and microclimate at 24 Picea abies-forests in southeastern Norway. The 24 sites are in 12 pairs where one of sites are a mature previously clear-cut forest and the other is a forest that never have been clear-cut. We investigated long-term effects of clear-cutting on forest structure and dead wood volumes
Replication Data for: Long-term leaching kinetics and solution chemistry of aqueous BaTiO3 powder suspensions: A numerical model supported experiment
The dataset contains the files and raw data for a submitted manuscript named: Long-term leaching kinetics and solution chemistry of aqueous BaTiO3 powder suspensions: A numerical model supported experiment. The purpose of the dataset is to allow other researchers to replicate the study and model output.
The data has been generated from the exposure of BaTiO3 powders made in-house to aqueous HCl solutions of different pH for varying periods of time up to 31 days. The data gives a comprehensive description of the dissolution/leaching behavior of the perovskite BaTiO3 in aqueous environments. The dataset contains data from: Inductively coupled plasma mass spectrometry (ICP-MS) from the liquid solution after centrifugating out the BaTiO3 powders upon completion of an individual experiment. pH data obtained over the course of each individual liquid exposure experiment. Scanning Electron Microscopy (SEM) micrographs of BaTiO3 powders before, and after exposure to the different solutions. X-Ray Diffraction (XRD) diffractograms of BaTiO3 powders before, and after exposure to the different solutions. Energy Dispersive X-ray Spectrometry (EDXS) analysis of BaTiO3 powder after exposure to aqueous solution. Particle Size Distribution (PSD) analysis of BaTiO3 powder before exposure to liquid. Brunauer-Emmett-Teller (BET) analysis of BaTiO3 powder before exposure to liquid. Plain text model output data from a numerical model of the chemical kinetics of the system in the proprietary software COMSOL Multiphysics. The model files are also included
Replication Data for: Are nuclear masks all you need for improved out-of-domain generalization? A closer look at cancer classification in histopathology
This dataset is a processed version of the CAMELYON17 dataset used in the NeurIPS 2024 paper "Are nuclear masks all you need for improved out-of-domain generalization? A closer look at cancer classification in histopathology". It consists of patches / tiles from 50 Whole Slide Images (WSIs) (10 WSIs from each of the 5 hospitals) in the CAMELYON17 dataset that have tumour segmentation available. Tiles were picked such that each hospital has equal number of tumourous and non-tumours tiles. Each tile is of size 270x270 pixels. A tile is considered tumourous if the centre region of tile (90x90 pixels in size) has at least 1 pixel that lies inside the tumour segmentation map.
The dataset also contains nuclear segmentation masks for all the tiles. Masks were generated using HoVer-Net trained on the CoNSeP dataset
Flow loop plugging experiments: experimental logs
Flow loop sensor logs for plugging flow experiments which are described in the associated articles. The logs present experimental pressure, temperature, and mass flow rate measurements conducted at different places of the multiphase flow loop loaded with ice-in-decane slurry. The experiments were performed for various flow rates, particle concentrations, and medium temperatures
Supporting Data for: Fluctuations in somatic cell count and their impact on individual goat milk quality throughout lactation
A cohort of 40 goats was selected from the Livestock Production Research Centre (SHF at the Norwegian University of Life Sciences Ås, Norway. Throughout the study, a total of 360 milk samples were collected from the sampler of the milking machine.The somatic cell count (SCC) is used by the dairy industry as an indicator of milk quality and udder health. However, in goats, its reliability is significantly masked by non-infectious variables such as the milk secretion process and physiological lactation changes. Additionally, notable individual variability between goats might exist. This study aimed to investigate fluctuations in SCC in individual goats during an entire lactation and to examine the relationship between SCC and milk parameters such as bacterial count and chemical composition. Individual milk samples from forty Norwegian dairy goats from the University herd were collected monthly across an entire lactation including the pasture period. The goats were categorized based on SCC levels to analyze patterns in chemical components within these groups. Notably, goats exhibited increased SCC and decreased bacterial counts when moved to summer grazing pastures. At that stage milk samples from goats with the highest SCC (>2 000 000 cells/mL) showed a distinct decrease in individual bacterial counts (IBC). Several milk components, especially lactose, protein, and various minerals were affected by the SCC level, as well as pH. These effects were further amplified when considering the interaction with the lactation stage, influencing a broader range of variables. The results presented in this study provide new insights into the SCC`s effect on milk composition and its critical role. These insights underscore the necessity for establishing lactation-specific thresholds for interpretation, as well as for making adjustments in quality payments
Background data for: “I regret lying” VS. “I regret that I lied”: Variation in the clausal complementation profile of REGRET in American and British English
This dataset contains tabular files recording occurrences of the verb REGRET complemented by a that- or (S) -ing-complement clause (CC) in the GloWbE corpus. Tokens were retrieved using the online interface (https://www.english-corpora.org/glowbe/) and manually annotated for several syntactic and semantic variables (variety, text type, finiteness, subject in the main clause (MC), voice of the CC, meaning of the verb in the CC, subject in the CC, animacy of the subject in the CC, words in the CC, coreferentiality, intervening material, negation in the CC, temporal relation). See ReadMe file for more details. Related publication: Romasanta, Raquel P. 2022. “I regret lying” VS. “I regret that I lied”: Variation in the clausal complementation profile of REGRET in American and British English. Miscelánea 65: 37-5
Zooplankton behavioral data polar night 2023
The dataset is data of behavioral experiments conducted on diverse zooplankton collected in Svalbard in January 2023. It includes video files of the experiments, stimulus light information and experimental protocol. The dataset is collected as a part of the RCN project 315728 "Light as a Cue for Life in the Arctic and Northern Seas (LightLife)"
Replication Data for: The many faces of "možno" in Russian and across Slavic. Corpus investigation of constructions with the modal možno (Chapter 3)
This dataset encompasses data from the Main corpus of the Russian National Corpus (RNC, ruscorpora.ru) used for analysis provided in Chapter 3 of the Introductory Chapter in the doctoral dissertation "The many faces of "možno" in Russian and across Slavic. Corpus investigation of constructions with the modal možno".
Chapter 3 presents a study of 500 examples of Russian constructions with the modal word možno ‘can, be possible’. The query consisted of a single word možno without specification of a time period. The search returned 361 755 examples, 5000 examples were downloaded in the .xlsx format, pseudorandomized, and then the first 500 examples were extracted for the analysis. The data in the spreadsheet 01DataTheManyFacesOfMozno comprises these 500 examples. The data was collected in March 2023 from the RNC. All of the examples are semantically and syntactically annotated by hand based on the syntactic analyses given in the corpus. The syntactic and morphological categories used in the corpus are explained here https://ruscorpora.ru/corpus/main.</p