University of Tartu

DSpace at Tartu University Library
Not a member yet
    108495 research outputs found

    DUDU: A Treebank for Ottoman Turkish in UD Style

    No full text
    This paper introduces a recently released Ottoman Turkish (ota) treebank in Universal Dependencies (UD) style, DUDU. The DUDU Treebank consists of 1,064 automatically annotated and manually corrected sentences. The texts were manually collected from various academic or literary sources available on the Internet. Following preprocessing, the sentences were annotated using a MaCHAMP-based neural network model utilizing the large language model (LLM) architecture and manually corrected. The treebank became publicly available with the 2.14 release, and future steps involve expanding the treebank with more data and refining the annotation scheme. The treebank is the first and only treebank that utilizes the IJMES transliteration alphabet. The treebank not only gives insight on Ottoman Turkish lexically, morphologically, and syntactically, but also provides a small but robust test set for future computational models for Ottoman Turkish

    FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering

    Get PDF
    Data quality is crucial for training Large Language Models (LLMs). Traditional heuristic filters often miss low-quality text or mistakenly remove valuable content. In this paper, we introduce an LLM-based line-level filtering method to enhance training data quality. We use GPT-4o mini to label a 20,000-document sample from FineWeb at the line level, allowing the model to create descriptive labels for low-quality lines. These labels are grouped into nine main categories, and we train a DeBERTa-v3 classifier to scale the filtering to a 10B-token subset of FineWeb. To test the impact of our filtering, we train GPT-2 models on both the original and the filtered datasets. The results show that models trained on the filtered data achieve higher accuracy on the HellaSwag benchmark and reach their performance targets faster, even with up to 25\% less data. This demonstrates that LLM-based line-level filtering can significantly improve data quality and training efficiency for LLMs. We release our quality-annotated dataset, FinerWeb-10BT, and the codebase to support further work in this area

    Perspectives on Forests and Forestry in Finnish Online Discussions - A Topic Modeling Approach to Suomi24

    Get PDF
    This paper explores how forests and forest industry are perceived on the largest online discussion forum in Finland, Suomi24 ('Finland24'). Using 30,636 posts published in 2014–2020, we investigate what kind of topics and perspectives towards forest management can be found. We use BERTopic as our topic modeling approach and evaluate the results of its different modular combinations. As the dataset is not labeled, we demonstrate the validity of our best model through illustrating some of the topics about forest use. The results show that a combination of UMAP and K-means leads to the best topic quality. Our exploratory qualitative analysis indicates that the posts reflect polarized discourses between the forest industry and forest conservation adherents

    Differences in values, conflict beliefs, and resolution between Estonians and Estonian Russians

    No full text
    Uurimistöö eesmärk on välja selgitada erinevused eestlaste ja eesti venelaste konflikti hoiakutes ja lahendusviisides ning hinnata seoseid väärtuste, konflikti hoiakute ja konflikti lahendusviiside vahel. Uuringus osales 222 täiskasvanut eestlast ja vene emakeelega Eesti elanikku. Andmeid koguti veebiküsimustikuga, mis hindas osalejate seotuse ja autonoomia väärtusi, konstruktiivseid konflikti hoiakuid ja ebaefektiivseid konflikti lahendamise viise. Uurimistöö tulemused näitavad, et eestlastel on eesti venelastega võrreldes konstruktiivsemad hoiakud konflikti ning eesti venelased kasutavad rohkem ebaefektiivseid konflikti lahendamisviise. Seotuse väärtuste ja konstruktiivsete hoiakute vahel leiti negatiivne seos ning konstruktiivsete konflikti hoiakute ja ebaefektiivsete konflikti lahendusviiside vahel on negatiivne seos. Uurimistöö tulemused võimaldavad kultuurilisi erinevusi märgata ja seeläbi edendada mitmekultuurilises ühiskonnas mõistvamat suhtlust

    Läänemere makrovetikate lahustunud orgaanilise süsiniku dünaamika: tootmine, biojuurdepääsetavus ja ökosüsteemi mõju

    Get PDF
    Väitekirja elektrooniline versioon ei sisalda publikatsiooneRannikualad on ühed bioloogiliselt produktiivseimad piirkonnad Maal, olles olulised elurikkuse keskused ja mängides olulist rolli globaalses süsinikuringes. Nendes ökosüsteemides on makrovetikad ehk merevetikad võtmetähtsusega organismid, mis muudavad anorgaanilise süsiniku lahustunud orgaaniliseks süsinikuks (DOC), mis toetab mikroobide kooslusi ja kõrgemaid toiduahela tasemeid. Hoolimata nende olulisusest on makrovetikate poolt toodetud DOC roll mikroobide ahelas ja selle potentsiaal märkimisväärse sinise süsiniku allikana jäänud vähe uurituks. Samal ajal seisavad rannikualad silmitsi üha suurenevate ohtudega inimtegevuse tõttu, nagu toitainete reostus, linnastumine ja kliimamuutused, mis põhjustavad muutusi makrovetikate koosseisus. Eriti murettekitav on tõsiasi, et tõhusad süsinikusidujad nagu pruunvetikametsad asenduvad kiirekasvuliste niitvetikatega, mis vabastavad süsinikku kiiremini. See häirib energiavoogu mikroobide kooslustele ja nõrgendab toiduahela alust. Need muutused ohustavad ökosüsteemi üldist tootlikkust ja vastupanuvõimet, millel on ulatuslikud mõjud nii üksikutele liikidele kui ka tervetele rannikualadele ning inimkogukondadele, kes sõltuvad neist toidujulgeoleku, elatusvahendite ja looduslike tormibarjääride jaoks. Uurides seoseid makrovetikate DOC dünaamika, mikroobiprotsesside ja ökosüsteemi muutuste vahel, annab käesolev doktoritöö uusi teadmisi selle kohta, kuidas rannikualad panustavad mere süsinikureservi. Need teadmised on hädavajalikud kaitsemeetmete väljatöötamiseks, mis tagavad nende ökosüsteemide vastupanuvõime, võimaldades neil jätkuvalt toetada elurikkust, reguleerida süsinikuringet ja pakkuda bioloogilisi teenuseid.Coastal ecosystems are among the most biologically productive areas on Earth, serving as vital hubs for biodiversity and playing a crucial role in global carbon cycling. Within these systems, macroalgae, or seaweeds, are key contributors, transforming inorganic carbon into dissolved organic carbon (DOC) that fuels microbial communities and sustains higher levels of the food web. Despite their importance, the role of macroalgae-derived DOC in the microbial loop and its potential as a significant source of blue carbon remains understudied. At the same time, coastal ecosystems face growing threats from human activities such as nutrient pollution, urban development, and climate change, which are driving shifts in macroalgal community composition. Notably, kelp forests – efficient carbon storers – are being replaced by fast- growing filamentous algae that release carbon more rapidly, disrupting energy transfer to microbial communities and weakening the foundation of the food web. These changes threaten overall ecosystem productivity and resilience, with cascading effects that extend beyond individual species to entire coastal ecosystems and the human communities that rely on them for food security, livelihoods, and natural storm barriers. By investigating the connections between macroalgal DOC dynamics, microbial processes, and ecosystem changes, this thesis enhances our understanding of how coastal systems contribute to the marine carbon pool. These insights are critical for developing conservation strategies that protect the resilience of these ecosystems, ensuring their continued role in supporting biodiversity, regulating carbon cycles, and providing biological services.https://www.ester.ee/record=b573156

    Learning from external crisis: changes in Estonia’s civil protection framework in the light of Russo-Ukrainian war

    No full text
    For Ukraine, the war started in 2014, but only after the large-scale invasion by the Russian Federation did Estonia start to see the bigger picture and think about the protection of its own population within the civil protection framework. The civil protection framework of Estonia has changed due to the war in Ukraine. These changes are the awareness of citizens about the public bomb shelters, sirens, and the support of the internal security volunteers. In this thesis the policy learning framework is used to assess whether a policy change took place and if it was the result of a policy learning in a context of an external crisis. Policy learning does not occur from nowhere, therefore there are triggers that start the process. This study demonstrates that Russia’s use of hybrid measures, what constituted an external crisis for Estonia, contributed to the emergence of policy learning and as a result policy change. The biggest change triggered by the external crisis was the comprehensive model for the evacuation as a part of the civil protection framework. It comprises of a system of notification such as sirens, guidance for evacuation, and safe structures for taking cover. In addition to the aforementioned aspects, the transition to Estonian-language education has also been a significant change as an effort to curb the spread of Russian propaganda. If Estonia was afraid to deal with the previous aspects before the Ukrainian war, expecting a response from the Russian Federation, the situation where the Russian Federation has significantly fewer resources to respond allows for new opportunities. All these factors suggest that the policy learning process has started in Estonia. The findings reveal that perceived as an external crisis, Russian use of hybrid measures in Ukraine triggered policy learning in the area of civil protection in Estonia.https://www.ester.ee/record=b5733937*es

    Adding Metadata to Existing Parliamentary Speech Corpus

    No full text
    Parliamentary proceedings are convenient data sources for creating corpora for speech technology. Given its public nature, there is an abundance of extra information about the speakers that can be legally and ethically harvested to enrich this kind of corpora. This paper describes the methods we have used to add speaker metadata to the Stortinget Speech Corpus (SSC) containing over 5,000 hours of Norwegian speech with non-verbatim transcripts but without speaker metadata. The additional metadata for each speech segment includes speaker ID, gender, date of birth, municipality of birth, and counties represented. We also infer speaker dialect from their municipality of birth using a manually designed mapping between municipalities and Norwegian dialects. We provide observations on the SSC data and give suggestions for how it may be used for tasks other than speech recognition. Finally, we demonstrate the utility of this new metadata through a dialect identification task. The described methods can be adapted to add metadata information to parliamentary corpora in other languages

    Majandustsüklite ja finantskriiside prognoosimine masinõppe abil

    No full text
    Doktoritöö uurib masinõppe meetodite rakendamist makromajanduslike näitajate prognoosimises, keskendudes äristsüklite ja süsteemsete finantskriiside ennustamisele. Uuring ühendab kaks perspektiivi – andmepõhise meetodi, mis kasutab suurt hulka muutujaid mittelineaarsete mustrite tabamiseks majandusnäitajates, ning teooriapõhise lähenemise, mis valib muutujaid mudelitesse makroökonoomika teoreetiliste põhimõtete alusel, – kombineerides teooriapõhist muutujate valikut puhtalt andmetest tuvastatud ulatusliku ennustajate (nt kriisieelsete märkide) kogumiga. Uurimus rõhutab tasakaalustamata andmete käsitlemise, asjakohaste tunnuste valiku ja mittelineaarsuste arvessevõtmise tähtsust masinõppe mudelite rakendamisel. Tulemused viitavad sellele, et võimendamis-meetodid võivad suurte andmekogumite puhul olla tõhusad, kuid need nõuavad hoolikat ettevalmistust, eriti kui andmete maht on piiratud. Doktoritöö toob esile ka ajalooliste andmete põhjal kriiside ennustamise keerukuse ning rõhutab vajadust edasiste uuringute järele mudelite optimeerimise vallas. Doktoritöö demonstreerib, kuidas masinõpe saab täiustada makromajanduslike näitajate prognoosimist, ühendades kaasaegseid masinõppetehnikaid traditsiooniliste ökonomeetriliste lähenemistega. Samuti rõhutatakse töös andmete ja kasutatavate meetodite kriitiliselt hindamise tähtsust tagamaks, et masinõppemudelid täiendaksid, mitte ei asendaks traditsioonilisi ökonomeetrilisi meetodeid. Kuigi suurandmete valdkonna areng on suurendanud masinõppe atraktiivsust, ei ole mõistlik sellele pimesi toetuda ega traditsioonilisi meetodeid kõrvale heita. Selle asemel annavad masinõppemudelid parimaid tulemusi traditsiooniliste ökonomeetriliste tehnikatega kombineerituna. Mõlemal on ühised statistilised alused, mistõttu on paljud uued meetodid traditsioonilistega lähemalt seotud kui esialgu tunduda võib.This thesis explores the application of machine learning methods in macroeconomic forecasting, focusing on business cycles and the prediction of systemic financial crises. It investigates two key approaches: The study integrates a data-driven method that leverages a large set of variables to capture economic nonlinearities and a theory-based approach that selects variables based on macroeconomic principles by combining theory-driven variable selection with a broad range of predictors. The research underscores the significance of addressing imbalanced data, selecting relevant features, and accounting for nonlinearities in machine learning models. It suggests that boosting methods can be effective when dealing with large datasets, although they require careful preparation, especially when the number of data is limited. The thesis also highlights the difficulties of crisis prediction using historical data and emphasizes the need for further research on model optimization. Overall, the thesis demonstrates how machine learning can enhance macroeconomic forecasting by merging modern machine learning techniques with traditional econometric approaches. It stresses the importance of critically evaluating data and methodologies to ensure that machine learning models complement rather than replace conventional methods. While the rise of Big Data has increased the appeal of machine learning, relying on it blindly or discarding traditional approaches is not advisable. Instead, machine learning models tend to perform best when combined with econometric techniques, as they share common statistical foundations, making many new methods more closely linked to traditional ones than they initially appear.https://www.ester.ee/record=b574399

    OpusDistillery: A Configurable End-to-End Pipeline for Systematic Multilingual Distillation of Open NMT Models

    Get PDF
    In this work, we introduce OpusDistillery, a novel framework to streamline the Knowledge Distillation (KD) process of multilingual NMT models. OpusDistillery's main features are the integration of openly available teacher models from OPUS-MT and Hugging Face, comprehensive multilingual support and robust GPU utilization tracking. We describe the tool in detail and discuss the individual contributions of its pipeline components, demonstrating its flexibility for different use cases. OpusDistillery is open-source and released under a permissive license, aiming to facilitate further research and development in the field of multilingual KD for any sequence-to-sequence task. Our code is available at https://github.com/Helsinki-NLP/OpusDistillery

    60,617

    full texts

    108,495

    metadata records
    Updated in last 30 days.
    DSpace at Tartu University Library is based in Estonia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇