Centrum Wiskunde & Informatica

CWI's Institutional Repository
Not a member yet
    26838 research outputs found

    craniac-swistak/s-covers

    No full text

    YagaoLiu/SFSS

    No full text

    How well do LLMs reason over tabular data, really?

    Get PDF
    Large Language Models (LLMs) excel in natural language tasks, but less is known about their reasoning capabilities over tabular data. Prior analyses devise evaluation strategies that poorly reflect an LLM's realistic performance on tabular queries. Moreover, we have a limited understanding of the robustness of LLMs towards realistic variations in tabular inputs. Therefore, we ask: Can general-purpose LLMs reason over tabular data, really?, and focus on two questions 1) are tabular reasoning capabilities of general-purpose LLMs robust to real-world characteristics of tabular inputs, and 2) how can we realistically evaluate an LLM's performance on analytical tabular queries? Building on a recent tabular reasoning benchmark, we first surface shortcomings of its multiple-choice prompt evaluation strategy, as well as commonly used free-form text metrics such as SacreBleu and BERT-score. We show that an LLM-as-a-judge procedure yields more reliable performance insights and unveil a significant deficit in tabular reasoning performance of LLMs. We then extend the tabular inputs reflecting three common characteristics in practice: 1) missing values, 2) duplicate entities, and 3) structural variations. Experiments show that the tabular reasoning capabilities of general-purpose LLMs suffer from these variations, stressing the importance of improving their robustness for realistic tabular inputs

    Clifford testing: algorithms and lower bounds

    Get PDF
    We consider the problem of Clifford testing, which asks whether a black-box nn-qubit unitary is a Clifford unitary or at least ε\varepsilon-far from every Clifford unitary. We give the first 4-query Clifford tester, which decides this problem with probability poly(ε)\mathrm{poly}(\varepsilon). This contrasts with the minimum of 6 copies required for the closely-related task of stabilizer testing. We show that our tester is tolerant, by adapting techniques from tolerant stabilizer testing to our setting. In doing so, we settle in the positive a conjecture of Bu, Gu and Jaffe, by proving a polynomial inverse theorem for a non-commutative Gowers 3-uniformity norm. We also consider the restricted setting of single-copy access, where we give an O(n)O(n)-query Clifford tester that requires no auxiliary memory qubits or adaptivity. We complement this with a lower bound, proving that any such, potentially adaptive, single-copy algorithm needs at least Ω(n1/4)\Omega(n^{1/4}) queries. To obtain our results, we leverage the structure of the commutant of the Clifford group, obtaining several technical statements that may be of independent interest

    Metadata matters in dense table retrieval

    Get PDF
    Recent advances in Large Language Models (LLMs) have enabled powerful systems that perform tasks by reasoning over tabular data [9, 10, 13, 7, 4]. While these systems typically assume relevant data is provided with a query, real-world use cases are mostly open-domain, meaning they receive a query without context regarding the underlying tables. Retrieving relevant tables is typically done over dense embeddings of serialized tables [5]. Yet, there is limited understanding of the effectiveness of different inputs and serialization methods for using such offthe-shelf text-embedding models for table retrieval. In this work, we show that different serialization strategies result in significant variations in retrieval performance. Additionally, we surface shortcomings in commonly used benchmarks applied in open-domain settings, motivating further study and refinement

    Invited talk: "How AI will kill us"

    No full text

    Human machine interaction and AI transparency

    No full text
    Abdallah El Ali, a Human-Computer Interaction (HCI) researcher with a background in cognitive science discusses trustworthy AI, explainability and transparency with Ahmad Tafti from the University of Pittsburgh and Humanitarian AI Today’s Producer, Brent Phillips

    FireScore, a framework for incident risk evaluation, simulation, coverage optimization and relocation experiments

    Get PDF
    This paper introduces fireSCore, an open source framework for incident risk evaluation, simulation, coverage optimization, and relocation experiments. The visualization front-end provides a live view of current coverage for the most common fire department units. Manually changing a unit status allows for a view into future coverage as it triggers an immediate recalculation of prognosed response times and coverage using the Open Source Routing Machine. The back-end provides the controller and model, and implements various algorithms, e.g. a relocation algorithm that optimizes coverage during major incidents. The databroker handles communication with data sources and provides data for the front- and back-end. An optional simulator adds an environment in which various scenarios, models and algorithms can be tested and aims to drive current and future organizational developments within the Dutch national fire service

    Zo wordt zonne-energie weer 'made in Holland'

    No full text

    13,690

    full texts

    26,838

    metadata records
    Updated in last 30 days.
    CWI's Institutional Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇