University of Tartu

DSpace at Tartu University Library
Not a member yet
    108495 research outputs found

    Got Compute, but No Data: Lessons From Post-training a Finnish LLM

    No full text
    As LLMs gain more popularity as chatbots and general assistants, methods have been developed to enable LLMs to follow instructions and align with human preferences. These methods have found success in the field, but their effectiveness has not been demonstrated outside of high-resource languages. In this work, we discuss our experiences in post-training an LLM for instruction-following for English and Finnish. We use a multilingual LLM to translate instruction and preference datasets from English to Finnish. We perform instruction tuning and preference optimization in English and Finnish and evaluate the instruction-following capabilities of the model in both languages. Our results show that with a few hundred Finnish instruction samples we can obtain competitive performance in Finnish instruction-following. We also found that although preference optimization in English offers some cross-lingual benefits, we obtain our best results by using preference data from both languages. We release our model, datasets, and recipes under open licenses at https://huggingface.co/LumiOpen/Poro-34B-chat-OpenAssistant

    Associations Between Estonian Children's Anxiety Symptoms and Social Relationships Based on the Data of the Estonian Children's Mental Health Study

    No full text
    Uurimistöö eesmärgiks oli uurida seoseid Eesti laste ärevuse näitajate ning nende sotsiaalsete suhete vahel. Sotsiaalsete suhete all mõeldakse antud töös noorte pere- ja sõprussuhteid. Uurimistöö toetub andmete osas Tervise Arengu Instituudi, Tartu ülikooli ja Turu-uuringute AS poolt läbiviidavale Laste vaimse tervise uuringule (LVTU). Uuringus osales 471 kooliõpilast vanuses 11 - 17 eluaastat. Noorema grupi moodustasid 11 - 14-aastased õpilased ning vanema grupi õpilased vanusega 15 ja sellest vanemad. Uurimustöö tulemused näitasid, et tüdrukutel esines ärevuse ilminguid rohkem kui poistel. Lisaks leiti, et kvaliteetsed sõprussuhted on negatiivses seoses noorte ärevuse näitajatega ning vägivaldsed peresuhted on ärevuse skooriga positiivses seoses. Samuti selgus, et kvaliteetsetel sõprussuhetel oli antud uurimuses suurim ennustamisvõime ärevuse üldskoorile. Tulemustest järeldub, et noorte ärevuse ja sotsiaalsete suhete vahel on seos ning noorte heaolu toetamiseks tuleb valida vastavad toetus- ning sekkumismeetodid

    A ‘bridge of science’ across the Gulf of Finland. Scientific relations between Estonia and Finland from 1918 to 1940

    No full text
    Eesti ja Soome haritlaste vahelisest kultuurikoostööst maailmasõdadevahelisel ajal on aastakümnete vältelt korduvalt ja põhjalikult kirjutatud, kuid sissevaated maadevahelistesse teadussidemetesse on valdavalt keskendunud koostööle humanitaarteadustes. Seda, millist koostööd esines teistes valdkondades, sh loodus-, põllumajandus-, arsti- ja muudes teadustes, on aga käsitletud märgatavalt vähem. Käesoleva väitekirja eesmärgiks on tõestada, et Eesti ja Soome vahel oli 1920. ja 1930. aastatel oluliselt mitmekülgsem teaduslik suhtlus ja koostöö kui seda on näha senisest ajalookirjutusest. Selleks vaadeldakse, millised olid kommunikatsioonikanalid ehk peamised koostöömoodused, mida Eesti ja Soome teadlased teabe ja kogemuste jagamiseks kasutasid, mil moel koostöömooduste aktiivsus ajas muutus, millistes valdkondades oli teadussuhtlus aktiivsem ning kumb osapool sai suhetest rohkem kasu. Väitekiri tugineb suurel hulgal mõlema maa mäluasutustest ning avaldatud kirjandusest leiduval teabel, millest osa pole maadevaheliste teadussuhete kontekstis varem kasutatud. Kommunikatsioonikanalite vaatlemisel on aluseks võetud soome ajaloolase Marjatta Hietala mudel, mille kohaselt olid peamisteks teabe ja kogemuste jagamise moodusteks välismaiste asjatundjate värbamine/kasutamine, välismaal õppimine ja isiklikud kontaktid, õppe- ja tööreisid välismaale, uurimistööde avaldamine ja vahetamine ning rahvusvahelised kongressid ja näitused. Doktoritööst selgub, et Eesti ja Soome vahel oli 1920. ja 1930. aastatel arvukalt teaduslikke sidemeid ning neid oli nii humanitaar-, loodus-, põllumajandus-, arsti-, loomaarsti- kui usuteadustes. Samal ajal kinnitab väitekiri, et kõige tihedamad sidemed olid humanitaarteadustes. Kõige mitmekülgsemateks koostöömoodusteks osutusid töö- ja uurimisreisid ning rahvusvahelised kongressid, kuid kõigi mooduste puhul on võimalik tuua näiteid mitmest erinevast teadusvaldkonnast. Maadevahelised teaduslikud sidemed olid arvukamad ja tugevamad 1920. aastate esimesel poolel ja 1930. aastate teisel poolel ning üldises plaanis said teadussuhetest enam kasu Eesti teadlased.During the past decades, there have been numerous studies into cultural ties and cooperation between Estonian and Finnish intellectuals between the two world wars, but the study of scientific cooperation between the two countries has mostly concentrated on connections in the fields of humanities. There has been less research into what ties and connections there were in natural, agricultural, medical, and other sciences. The aim of this dissertation is to show that scientific communication and cooperation between Estonia and Finland during the 1920s and 1930s was considerably more diverse than suggested by previous historiography. The dissertation analyses what were the channels of communication that Estonian and Finnish scientists used to share knowledge and experiences, whether there were changes in the level of activity in said channels, what were the fields of research where cooperation was more active, and which party benefited more from the connections. This research is based on materials found in the archives and published literature of both countries and some of the information has not been previously used in the context of scientific connections between Estonia and Finland. The analysis of channels of communication is based on a model developed by the Finnish historian Marjatta Hietala, according to whom the main channels of communication were employment of foreign experts, study abroad and personal contacts, work and research trips abroad, publication and exchange of research papers, literature, journals etc., and international congresses and exhibitions. The dissertation reveals that there were numerous connections in a wide range of sciences, including humanities, natural, agricultural, medical, veterinarian sciences, and theology. Simultaneously, the dissertation confirms that, as suggested by previous studies, the most active cooperation took place in humanities. The most diverse channels of communication were revealed to have been trips abroad and scientific congresses, but examples from a wide range of fields of research could be seen in all channels of communication. Scientific ties between Estonia and Finland were stronger and more numerous during the early 1920s and late 1930s and, in general, the Estonian scientists benefited more from the connections between the two countries.https://www.ester.ee/record=b573403

    Infoturbe ja privaatsuse haldamise meetod nutilahendustes

    No full text
    Kujutage ette uut nutikat parkimislahendust, mis võimaldab teil parkida oma sõidukit mobiilirakenduse abil. Selline rakendus võimaldab teil parkida parklatesse, mis kuuluvad ükskõik millisele ettevõttele riigis. Samas rakenduses saate teatada ka teiste sõidukite ebaseaduslikust parkimisest ning vaadata üle oma parkimispiletid ja trahvid valesti parkimise eest. Sinu kui lõppkasutaja jaoks näeb rakendus välja nagu üks süsteem. Sellegipoolest koosneb see mitmest parklaomaniku, rakenduse arendusettevõtte ja kohaliku parkimiskorraldaja hallatavast süsteemist. Kuna igal ettevõttel on juba enda arendatud süsteemid olemas, kasutavad nad neid ja lepivad kokku ainult nende integreerimises. Kuigi eeldatakse, et iga eraldi ehitatud süsteem on turvline ja kaitseb andmeid, muudab iga eraldiseisva süsteemi integreeriminesüsteemi keerulisemaks ja uutele turvaohtudele ja andmeleketele vastuvõtlikumaks. Senised traditsioonilised turvameetmed koostati teistsuguste olukordade jaoks, mistõttu on vaja rohkem panusada, et kaitsta arendatavaid nutisüsteeme. Me pakume selle lünga täitmiseks välja infoturbe ja privaatsuse haldamise meetodi nutilahendustes. See meetod peaks aitama ettevõtteid, kes soovivad pakkuda oma infosüsteemi uue nutika lahenduse komponendina. Seega peaks meetod tagama, et vastloodud nutika lahenduse süsteem kaitseb nii kasutajate kui ka ettevõtete tundlikke andmeid. Kavandatav meetod aitab ettevõtteid kolmes etapis. Esiteks aitab see määratleda, kuidas ettevõtted oma süsteemis teavet kaitsevad ja kas nutilahenduse integratsioon võib olla vastuolus olemasolevate eeldustega. Teiseks näitame, kuidas ettevõtted saavad kasutada olemasolevaid tööriistu, et kontrollida, kas nende süsteemid ikka vastavad integratsiooni puhul kohalikele privaatsusregulatisoonide nõuetele. Kolmandaks pakume välja kaks identiteedihaldussüsteemi kavandit, mis võimaldavad mittetraditsioonilistel usalduslikel eeldustel kaitsta vahetatud andmeid partneritega.Imagine a new smart parking solution that allows you to park your vehicle using a mobile app. Such an app will enable you to park in lots owned by any company in the country. In the same app, you can also report illegal parking of others' vehicles and review your parking tickets and fines for wrong parking. For you as an end user, the app looks like a single system. Yet, it comprises multiple systems managed by parking lot owners, the app development company, and the local parking enforcement agency. As each of these companies already has its built systems, they only agree on the integration to allow you to use a smart parking solution. Even though each separately built system is expected to be secure and protect data, the integration into a new complex system makes each separate one prone to new security threats and data leakages. Traditional information security approaches struggle to support securing collaborative environments, such as those found in smart solutions, so more is needed to protect evolving intelligent solution systems. We propose a method for information security and privacy management in smart solutions to bridge this gap. This method should help organisations which want to provide their information systems as components for a new smart solution. This method should ensure that the newly established smart solution protects both users' and the companies' sensitive data. The proposed method should help companies in the three stages. First, it helps define how companies protect information in their system and how integration into a smart solution may contradict the used assumptions. Second, we show how companies can utilise existing open-source tools to check that their systems comply with local privacy laws in the case of integration. Third, we propose two identity management system designs which allow non-traditional trust assumptions to protect the exchanged data with partners.https://www.ester.ee/record=b574375

    Vähimaast lahtiütlemine: rinnavähist räägitavate lugude muutmine

    No full text
    Käesolevas doktoritöös “Vähimaast lahtiütlemine: Rinnavähist räägitavate lugude muutmine" uurin Ameerika Ühendriikides rinnavähki haigestunud naiste autobiograafilisi narratiive. Rinnavähki neoliberaalsest ja individualistlikust perspektiivist vaatlevad lood moodustavad olulise kaasaegse kultuurinähtuse ning on osa ulatuslikust objektide, struktuuride ja tähenduste võrgustikust, mis määravad inimeste arusaamu rinnavähist ning selle üle toimuvaid arutelusid laiemalt. Uurimistöös selgus, et peavoolu rinnavähinarratiivid asetavad enamasti rõhu ellujäämisele ja positiivsele mõtlemisele, isiklikule vastutusele ja heteronormatiivsetele/keskklassi väärtustele, varjutades seega tegelikkuse ja eksistentsi mitmekesisust ning takistades inimestel kaalumast teistsuguseid jutustamisvõimalusi, mis tõstaksid esile taoliste jutustamispraktikate seoseid majandus-poliitiliste huvide ja olukordadega ning kutsuksid esile eetilisemaid ja kogukonnale suunatud lähenemisi. Lisaks peavoolu rinnavähinarratiivide kriitikale ja diskussioonile erinevatest teguritest, mis neid kujundavad ja neile kaalu annavad, panen ette pöörduda kriitiliste autobiograafiliste rinnavähilugude ehk vastandnarratiivide poole, mis tõrguvad vastu “võitjate“ ehk “sangarite“ loojutustamise mustrile. Analüüsin kolme sellist narratiivi ning vaatlen neid seoses nii nende eripära kui ka ’aktivistlike’ omadustega. Väidan, et erinevad lood (mis ei rõhuta sidusust ja kangelase enesearengut; ei räägi rangelt individualistlikust vaatepunktist; ei järgi kirjutamiskursuste ja käsiraamatute juhiseid, kuidas kirjutada tänapäeval edukat memuaari) võivad aeglaselt ja järk-järgult kuuldavaks teha marginaliseeritute hääled ja aja jooksul viia eetilisemate eksistentsiviisideni.In this dissertation, Breaking Free of Cancerland: Changing the Stories We Tell About Breast Cancer, I examine autobiographical narratives written by women with breast cancer in the U.S. These stories, looking at breast cancer from a neoliberal, individualist perspective, constitute a big contemporary cultural phenomenon as well as part of an extensive network (of objects, structures, and meanings) that determines people’s perception of breast cancer and, consequently, what happens or does not happen on a broader level about it. In my research, I found that mainstream breast cancer narratives mostly emphasize survivorship and positive thinking, personal responsibility and heteronormative/middle class values of life. In doing so, they obscure different realities and modes of existence, and preclude people from considering different responses to this storytelling epidemic, such that might foreground its links to economic-political interests and circumstances, and elicit more ethical and community-oriented approaches. Alongside my critique of mainstream breast cancer stories and a discussion of various factors that shape them and keep them in currency, I suggest turning to counter-narratives – critical autobiographical breast cancer stories that resist the storytelling pattern of winners’/heroes’ tales. I analyze three such narratives that stood out for me and I look at them in connection to their own specific features and with respect to their activist qualities. I maintain that different stories (not emphasizing coherence and the hero’s self-development, not told from a strictly individualist point of view, not following the instructions of writing courses and manuals on how to write a successful memoir today) can slowly and gradually make the minoritarian voices heard and, over time, lead to more ethical ways of existence.https://www.ester.ee/record=b572044

    Aligning Language Models for Icelandic Legal Text Summarization

    Get PDF
    The integration of language models in the legal domain holds considerable promise for streamlining processes and improving efficiency in managing extensive workloads. However, the specialized terminology, nuanced language, and formal style of legal texts can present substantial challenges. This study examines whether preference-based training techniques, specifically Reinforcement Learning from Human Feedback and Direct Preference Optimization, can enhance models' performance in generating Icelandic legal summaries that align with domain-specific language standards and user preferences. We compare models fine-tuned with preference training to those using conventional supervised learning. Results indicate that preference training improves the legal accuracy of generated summaries over standard fine-tuning but does not significantly enhance the overall quality of Icelandic language usage. Discrepancies between automated metrics and human evaluations further underscore the importance of qualitative assessment in developing language models for the legal domain

    Assessed and Annotated Vowel Lengths in Spoken Icelandic Sentences for L1 and L2 Speakers: A Resource for Pronunciation Training

    No full text
    We introduce a dataset of time-aligned phonetic transcriptions focusing on vowel length (quantity) in Icelandic. Ultimately, this aims to support computer assisted pronunciation training (CAPT) software, to automatically assess length and possible errors in Icelandic learners' pronunciations. The dataset contains a range of long and short vowel targets, including the first acoustic description of quantity in non-native Icelandic. Evaluations assess how manual annotations and automatic forced alignment characterise quantity contrasts. Initial analyses also imply partial acquisition of phonologically conditioned quantity alternations by non-native speakers

    Eesti kirjanduse leksikon. Kava ja kokkuvõtted

    No full text

    "I Need More Context and an English Translation": Analysing How LLMs Identify Personal Information in Komi, Polish, and English

    Get PDF
    Automatic identification of personal information (PI) is particularly difficult for languages with limited linguistic resources. Recently, large language models (LLMs) have been applied to various tasks involving low-resourced languages, but their capability to process PI in such contexts remains under-explored. In this paper we provide a qualitative analysis of the outputs from three LLMs prompted to identify PI in texts written in Komi (Permyak and Zyrian), Polish, and English. Our analysis highlights challenges in using pre-trained LLMs for PI identification in both low- and medium-resourced languages. It also motivates the need to develop LLMs that understand the differences in how PI is expressed across languages with varying levels of availability of linguistic resources.https://aclanthology.org/2025.resourceful-1.0

    Towards large-scale speech foundation models for a low-resource minority language

    Get PDF
    Modern ASR systems require massive amounts of training data. While ASR training data for most languages are scarce and expensive to transcribe, a practical solution is to collect huge amounts of raw untranscribed speech and pre-train the ASR model in a self-supervised manner. Unfortunately, for many low-resource minority languages, even untranscribed speech data are scarce. In this paper, we propose a solution for the Northern Sámi language with 22,400 hours of speech extracted from the Finnish radio and television archives. We evaluated the model performance with different decoding algorithms and examined the models' internal behavior with interpretation-based techniques

    60,617

    full texts

    108,495

    metadata records
    Updated in last 30 days.
    DSpace at Tartu University Library is based in Estonia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇