4926 research outputs found
Sort by
Kt 93/k 550 (X-Ray Tomography 3D data of an Enveloped Clay Tablet, Museum of Anatolian Civilizations, Ankara)
Written Artefact Metadata
Object type: Enveloped clay tablet
Material: Clay
Writing: Cuneiform writing
Language: Old Assyrian
Nature of the text: Loan contract
Provenience: Kanesh (mod. Kültepe)
Period/Date: Old Assyrian (ca. 1920-1850 BCE)
Country of discovery: Türkiye
Dimensions: Height: 5 cm; Width: 4.3 cm; Depth: 2.7 cm
Museum/Collection: Museum of Anatolian Civilizations, Ankara, Türkiye
Excavation number: Kt 93/k 550
Publication number (envelope): forthcoming
Transliteration of the tablet inside the envelope
Ob.11 ma-na kù-babbar li-tí 2i-ṣé-er Ša-ak-du-nu-/a 3A-šur-ták-lá-ku 4i-šu 15 gín kù-babbar 55 na-ru-uq še-um 6ù 5 na-ru-uq gig 7ṣí-ba-sú : a-na 8ha-ar-pé : i-ša-qál lo.e.9iš-tù ha-muš-tim rev.10ša A-šur-i-mì-tí 11iti-kam : ša Sá!(KI)-ra-tim 12li-mu-um 13E-na-ma-nim 141 ma-na kù-babbar 15ší-im-tám a-na e-tí-šu 16i-ša-qal : šu-ma 17lá iš-qul ki-a-ma 18ṣí-ib-tám u.e.19ú-ṣa-áb igi Bu-li-/na le.e.20igi Pè-er-wa 21 igi Puzur4-A-na
Translation of the tablet in the envelope
Aššur -taklāku has loaned 1 mina of litum-quality silver to Šakdunuwa. He will pay 15 shekels of silver, 5 sacks of barley and 5 sacks of wheat, its interest, at the harvest time. From the week of Aššur-imittī, month ša Sarratim (ii), eponym Ennamānum. He will pay the amount agreed on at his term. If he has not paid, he will add the interest accordingly. In the presence of Bulina, of Perwa, and of Puzur-Ana.
Structure and data
00_Photos.zip:
Photos of the enveloped tablet from all sides.
Kt93k550aLe.jpg
Kt93k550aLoe.jpg
Kt93k550aOb.jpg
Kt93k550aRe.jpg
Kt93k550aReBis.jpg
Kt93k550aRev.jpg
Kt93k550aUe.jpg
01_VolumeData.zip:
3D tomographic reconstruction volume data.
017_93k550_reconstruction_crop.nxs
02_Visualisations.zip:
Geometry files of the tablet and envelope (.ply) and internal format files (.exa).
017_93k550_3.exa
017_93k550_3_medianw_0.015_32_8_-0.30_5.exa
017_93k550_3_medianw_0.015_32_8_-0.30_5_32_0.5_2.exa
017_93k550_envelope.ply
017_93k550_looseParts.ply
017_93k550_tablet.ply
03_Scripts.zip:
Scripts to run the extraction and visualisation software Exavis42.
017_93k550_extract.sh
017_93k550_visualise.s
INEL Kalmyk Corpus
Corpus citation
Baranova, Vlada. 2025. INEL Kalmyk Corpus. Archived at Universität Hamburg. Version 1.0. Publication date 2025-07-17. https://hdl.handle.net/11022/0000-0007-FFB1-2. Archived at Universität Hamburg. In: The INEL Corpora of Indigenous Northern Eurasian Languages. https://hdl.handle.net/11022/0000-0007-F45A-1.
Corpus Description
The INEL Kalmyk Corpus has been created within the long-term INEL project ("Grammatical Descriptions, Corpora and Language Technology for Indigenous Northern Eurasian Languages"), 2016–2033.
The corpus consists of transcribed audio recordings collected in the Republic of Kalmykia between 2007 and 2018 in the Ketchenerovsky District (Derbet and Torgut dialect).
All texts in the corpus are provided with interlinear morpheme-by-morpheme glosses and translation into English and Russian. All texts for which the audio recordings were accessible are time-aligned with them.
Corpus Size
The corpus contains 55 texts, 2,076 sentences, and 19,742 tokens. The total duration of the audio recordings is 4 hours and 23 minutes.
Funding
The corpus has been produced in the context of the joint research funding of the German Federal Government and Federal States in the Academies’ Programme, with funding from the Federal Ministry of Education and Research and the Free and Hanseatic City of Hamburg. The Academies’ Programme is coordinated by the Union of the German Academies of Sciences and Humanities.
Contributions / Acknowledgements
Native speakers generously shared their knowledge of Kalmyk, making the creation of this corpus possible. Zamira Xejchieva and Galina Cabdy`rova assisted with oral transcription and the Russian translation of the audio materials.
Part of the materials were recorded during joint expeditions of St. Petersburg University and the Institute for Linguistic Studies of the Russian Academy of Sciences in 2007–2008, under the direction of Elena Perekhvalskaya and Sergey Say.
This corpus primarily follows the transcription system and partially adopts the glossing conventions developed by a research team led by Sergey Say, with input from other expedition participants.
Searching the corpus
The corpus can be downloaded from the ZFDM Repository using the links provided below and browsed or searched locally using the EXMARaLDA software or, alternatively, ELAN.
Online search with Tsakorpus platform is available at https://inel.corpora.uni-hamburg.de/KalmykCorpus/search.
Remote search with EXMARaLDA is also possible without downloading all the files (see https://inel.corpora.uni-hamburg.de/portal/help/en/index.php).
See the user documentation (section 3) for details on transcription, annotation tiers and annotation tags.
Find further information and links on the Kalmyk Corpus page at the INEL Resources portal: https://inel.corpora.uni-hamburg.de/portal/corpora/kalmyk/
ScriptSight_v1.5
ScriptSight is a software tool that can be used to explore and sort collections of document-page images using their AI-generated ‘computational visual catalogues’. You can find a test sample of such catalogue here. Multiple visual attributes, such as text orientation, used writing implements (pencil versus ink), and text colour, can be filtered concurrently to isolate pages matching the specified criteria. Matching pages can be previewed as thumbnails with optional overlays indicating detected text regions, and and saved as organised, date-stamped folders for review or further processing.
What’s new in this version:
Several bug fixes
Several performance enhancements
Several filtering and viewing speed improvements
New colour filtering based on LAB space
Additional parameter controls in the config file
Watch the accompanying video for an illustration of the basic functions.
The research for this work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy - EXC 2176 ‘Understanding Written Artefacts: Material, Interaction and Transmission in Manuscript Cultures’, project no. 390893796. The research was conducted within the scope of the Centre for the Study of Manuscript Cultures (CSMC) at Universität Hamburg.
I thank Dr. Quang-Vinh Dang for his significant contributions to developing the AI models used to generate the "computational visual catalogues". In addition, I thank Hui Xu for testing this tool and providing a useful feedback
The PHOENIX/1D NewEra model atmosphere grid: Access software & low resolution synthetic spectra
Software to access NewEra spectrum files (DOI 10.25592/uhhfdm.16727) from python and an example reader.
get_NewEra_from_FDR.py get a single model from the data repository.
example_read_HSR_H5.py reads data from a single model h5 file.
example_read_structure_from_HSR_H5.py reads and parses model structure data (radii, temperatures etc.)
example_read_gaia_fmt.py read the first spectrum of one of the GAIA format archives
list_of_available_NewEraV2_models.txt: Version 2.0 of the NewEra spectra (for Teff>=5000K). Use these files!
list_of_available_NewEra_models.txt is a list of all available models, MD checksums, file sizes and download links.
Readme.PHOENIX.gaia_fmt.txt explains the format of the GAIA archive files
PHOENIX-NewEra-LowRes-SPECTRA.tar.gz archive with all low resolution spectra
PHOENIX-NewEra-JWST-SPECTRA.tar.gz archive with all spectra in the JWST spectral range and resolution
NewEra_for_GAIA_DR4.tar Synthetic spectra, colors and BCs in GAIA DR4 format as tar file, includes Readme
Curation Policy für das Forschungsdatenrepositorium der Universität Hamburg – Community des Hamburger Zentrums für Sprachkorpora (HZSK)
Diese Curation Policy legt die Kriterien und Verfahren für die Aufnahme, Pflege und langfristige Bereitstellung von Forschungsdaten in der HZSK-Community des Repositoriums der Universität Hamburg fest. Ziel ist eine nachhaltige, interoperable und nachvollziehbare Archivierung sowie die Unterstützung der Forschungscommunity in der Sprachwissenschaft
Material Analysis of the Graz and Leipzig Collections
“Material Analysis of the Graz and Leipzig Collections”, Manuscript Cultures in the Caucasus, Centre for the Study of Manuscript Cultures, Hamburg
The International Composition of Linguistic Landscapes at Japanese Universities _ Minoh Campus
This is the dataset for my master's thesis: The International Composition of Linguistic Landscapes at Japanese Universities.
It includes all linguistic landscape elements that were used for annotation and an excel sheet with annotative data
Full metadata for Jerusalem guestbook of Miryam and Moshe Ya'akov Ben-Gavriêl (1927-1966)
The dataset contains all metadata collected and created in the context of a digital edition of the Jerusalem guestbook (1927-1966) of Miryam and Moshe Ya'akov Ben-Gavriêl, kept at the National Library of Israel, call no. ARC. Ms. Var. 365 1 11. The Excel file contains two spreadsheets: "entries" contains information on almost 1,200 individual entries in the guestbook and "visitors" contains information on the visitors that have been identified by name.
More information on the guestbook and direct access to the exploration environment can be found at https://csmc.demo.hcds.uni-hamburg.de/ The digital edition is part of the research project RFE22 "A Fresh Look. Visualising Digitised German-Jewish Archives" at the Cluster of Excellence Understanding Written Artefacts (UWA) and has been realized in close cooperation with the Hub of Computing and Data Science (HCDS)
The DLA-RMR dataset: Annotated subset of RMR notebooks for CVC development
This dataset is structured into four components, each serving a distinct role in the development of a document analysis system.
Word-level annotations are provided in the file word_annotations_for_cropped_images.json. These annotations describe the images contained in the cropped_images folder. Each entry specifies the location of a word as a polygon, together with its orientation (horizontal, vertical, or tilted) and the type of writing implement used (ink or pencil). Additional metadata, such as bounding boxes and segmentation areas, is also included.
Cropped images are stored in the cropped_images folder. This set comprises 50 images, each containing only the primary page extracted from the corresponding full notebook scans.
Full images are located in the full_images folder. This collection also contains 50 items, representing the complete notebook scans in which the primary page appears alongside other material.
Page-level annotations are contained in the page_annotations folder. These are provided in YOLO format, with a single class (page) defined in classes.txt. Each annotation file specifies the bounding box of the primary page within the corresponding image in the full_images folder.
Examples illustrate the annotation structure. In the JSON file, a typical word annotation records polygon coordinates, the attribute "orientation": "horizontal", and "writing_tool": "pencil". In the YOLO annotations, a sample entry such as 0 0.499023 0.500776 0.777344 0.816912 denotes the normalised coordinates of the primary page bounding box.
Acknowledgement:
The research for this work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy - EXC 2176 ‘Understanding Written Artefacts: Material, Interaction and Transmission in Manuscript Cultures’, project no. 390893796. The research was conducted within the scope of the Centre for the Study of Manuscript Cultures (CSMC) at Universität Hamburg.
The images are taken from notebook pages of Rainer Maria Rilke, from the Deutsche Literaturarchiv Marbach (DLA), A:Rilke-Archiv Gernsbach.
We thank Hui Xu for her support in annotating the images
Beta maṣāḥǝft working papers 5: The manuscript Berlin, Staatsbibliothek, Petermann II Nachtrag 24, and its collection of qǝne-poems
Among various additional texts, the manuscript Berlin, Staatsbibliothek, Petermann II Nachtrag 24 includes a number of qəne-poems that have not been previously addressed by scholars. The working paper offers their preliminary assessment