arXiv.org e-Print Archive

arXiv.org e-Print Archive
Not a member yet
    623509 research outputs found

    Unlocking the Theory Behind Scaling 1-Bit Neural Networks

    No full text
    Recently, 1-bit Large Language Models (LLMs) have emerged, showcasing an impressive combination of efficiency and performance that rivals traditional LLMs. Research by Wang et al. (2023); Ma et al. (2024) indicates that the performance of these 1-bit LLMs progressively improves as the number of parameters increases, hinting at the potential existence of a Scaling Law for 1-bit Neural Networks. In this paper, we present the first theoretical result that rigorously establishes this scaling law for 1-bit models. We prove that, despite the constraint of weights restricted to {1,+1}\{-1, +1\}, the dynamics of model training inevitably align with kernel behavior as the network width grows. This theoretical breakthrough guarantees convergence of the 1-bit model to an arbitrarily small loss as width increases. Furthermore, we introduce the concept of the generalization difference, defined as the gap between the outputs of 1-bit networks and their full-precision counterparts, and demonstrate that this difference maintains a negligible level as network width scales. Building on the work of Kaplan et al. (2020), we conclude by examining how the training loss scales as a power-law function of the model size, dataset size, and computational resources utilized for training. Our findings underscore the promising potential of scaling 1-bit neural networks, suggesting that int1 could become the standard in future neural network precision

    Collective Dissipation of Oscillator Dipoles Strongly Coupled to 1-D Electromagnetic Reservoirs

    No full text
    We study the collective dissipative dynamics of dipoles modeled as harmonic oscillators coupled to 1-D electromagnetic reservoirs. The bosonic nature of the dipole oscillators as well as the reservoir modes allows an exact numerical simulation of the dynamics for arbitrary coupling strengths. At weak coupling, apart from essentially recovering the dynamics expected from a Markovian Lindblad master equation, we also obtain non-Markovian effects for spatially separated two-level emitters. In the so called ultrastrong coupling regime, we find the dynamics and steady state depends on the choice of the reservoir which is chosen as either an ideal cavity with equispaced, unbounded dispersion or a cavity array with a bounded dispersion. Moreover, at even higher coupling strengths, we find a decoupling between the light and matter degrees of freedom attributable to the increased importance of the diamagnetic term in the Hamiltonian. In this regime, we find that the dependence of the dynamics on the separation between the dipoles is not important and the dynamics is dominated by the occupation of the polariton mode of lowest energy

    MamT4^4: Multi-view Attention Networks for Mammography Cancer Classification

    No full text
    In this study, we introduce a novel method, called MamT4^4, which is used for simultaneous analysis of four mammography images. A decision is made based on one image of a breast, with attention also devoted to three additional images: another view of the same breast and two images of the other breast. This approach enables the algorithm to closely replicate the practice of a radiologist who reviews the entire set of mammograms for a patient. Furthermore, this paper emphasizes the preprocessing of images, specifically proposing a cropping model (U-Net based on ResNet-34) to help the method remove image artifacts and focus on the breast region. To the best of our knowledge, this study is the first to achieve a ROC-AUC of 84.0 ±\pm 1.7 and an F1 score of 56.0 ±\pm 1.3 on an independent test dataset of Vietnam digital mammography (VinDr-Mammo), which is preprocessed with the cropping model.The crop model is available here: https://github.com/ispras/mammo_cro

    Perceiving and Countering Hate: The Role of Identity in Online Responses

    No full text
    This study investigates how online counterspeech, defined as direct responses to harmful online content with the intention of dissuading the perpetrator from further engaging in such behavior, is influenced by the match between a target of the hate speech and a counterspeech writer\u27s identity. Using a sample of 458 English-speaking adults who responded to online hate speech posts covering race, gender, religion, sexual orientation, and disability status, our research reveals that the match between a hate post\u27s topic and a counter-speaker\u27s identity (topic-identity match, or TIM) shapes perceptions of hatefulness and experiences with counterspeech writing. Specifically, TIM significantly increases the perceived hatefulness of posts related to race and sexual orientation. TIM generally boosts counter-speakers\u27 satisfaction and perceived effectiveness of their responses, and reduces the difficulty of crafting them, with an exception of gender-focused hate speech. In addition, counterspeech that displayed more empathy, was longer, had a more positive tone, and was associated with higher ratings of effectiveness and perceptions of hatefulness. Prior experience with, and openness to AI writing assistance tools like ChatGPT, correlate negatively with perceived difficulty in writing online counterspeech. Overall, this study contributes insights into linguistic and identity-related factors shaping counterspeech on social media. The findings inform the development of supportive technologies and moderation strategies for promoting effective responses to online hate.28 page

    A shooting-Newton procedure for solving fractional terminal value problems

    No full text
    In this paper we consider the numerical solution of fractional terminal value problems (FDE-TVPs). In particular, the proposed procedure uses a Newton-type iteration which is particularly efficient when coupled with a recently-introduced step-by-step procedure for solving fractional initial value problems (FDE-IVPs), able to produce spectrally accurate solutions of FDE problems. Some numerical tests are reported to make evidence of its effectiveness.23 pages, 4 figures, 7 tables, one typo fixe

    CmdCaliper: A Semantic-Aware Command-Line Embedding Model and Dataset for Security Research

    No full text
    This research addresses command-line embedding in cybersecurity, a field obstructed by the lack of comprehensive datasets due to privacy and regulation concerns. We propose the first dataset of similar command lines, named CyPHER, for training and unbiased evaluation. The training set is generated using a set of large language models (LLMs) comprising 28,520 similar command-line pairs. Our testing dataset consists of 2,807 similar command-line pairs sourced from authentic command-line data. In addition, we propose a command-line embedding model named CmdCaliper, enabling the computation of semantic similarity with command lines. Performance evaluations demonstrate that the smallest version of CmdCaliper (30 million parameters) suppresses state-of-the-art (SOTA) sentence embedding models with ten times more parameters across various tasks (e.g., malicious command-line detection and similar command-line retrieval). Our study explores the feasibility of data generation using LLMs in the cybersecurity domain. Furthermore, we release our proposed command-line dataset, embedding models\u27 weights and all program codes to the public. This advancement paves the way for more effective command-line embedding for future researchers

    An automorphic description of the zeta function of the basic stratum of certain Kottwitz varieties

    No full text
    We derive formulas for the number of points on the basic stratum of certain Kottwitz varieties in terms of automorphic representations and certain explicit polynomials, for which we present efficient algorithms for computation. We obtain our results using the trace formula, base change, representations of general linear groups over p-adic fields, and a truncation of the formula of Kottwitz for the number of points on Shimura varieties over finite fields.arXiv admin note: text overlap with arXiv:1111.6830 by other author

    RLE: A Unified Perspective of Data Augmentation for Cross-Spectral Re-identification

    No full text
    This paper makes a step towards modeling the modality discrepancy in the cross-spectral re-identification task. Based on the Lambertain model, we observe that the non-linear modality discrepancy mainly comes from diverse linear transformations acting on the surface of different materials. From this view, we unify all data augmentation strategies for cross-spectral re-identification by mimicking such local linear transformations and categorizing them into moderate transformation and radical transformation. By extending the observation, we propose a Random Linear Enhancement (RLE) strategy which includes Moderate Random Linear Enhancement (MRLE) and Radical Random Linear Enhancement (RRLE) to push the boundaries of both types of transformation. Moderate Random Linear Enhancement is designed to provide diverse image transformations that satisfy the original linear correlations under constrained conditions, whereas Radical Random Linear Enhancement seeks to generate local linear transformations directly without relying on external information. The experimental results not only demonstrate the superiority and effectiveness of RLE but also confirm its great potential as a general-purpose data augmentation for cross-spectral re-identification. The code is available at \textcolor{magenta}{\url{https://github.com/stone96123/RLE}}.Accepted to NeurIPS 202

    Local structure of tame symmetric algebras of period four

    No full text
    In this paper we study the structure of Gabriel quivers of tame symmetric algebras of period four. More precisely, we focus on algebras having Gabriel quiver {\it biregular}, i.e. the numbers of arrows starting and ending at any vertex are equal, and do not exceed 22. We describe the local structure of biregular Gabriel quivers of tame symmetric algebras of period four, including certain idempotent algebras. The main result of this paper shows that, in fact, these Gabriel quivers have local structure exactly as Gabriel quivers of so called {\it weighted surface algebras}, which partially extends known characterization of algebras of generalized quaternion type

    Efficient Sparse Training with Structured Dropout

    No full text
    Dropout is a common regularisation technique in deep learning that improves generalisation. Even though it introduces sparsity and thus potential for higher throughput, it usually cannot bring speed-ups on GPUs due to its unstructured nature. In this project, I experiment with SparseDrop, a structured, hardware-friendly variant of dropout that can exploit such sparsity. I provide a CUDA implementation of SparseDrop, achieving speed-ups against its dense counterpart even at low sparsity levels. The empirical results demonstrate that SparseDrop provides similar, or sometimes even better, regularisation properties as standard dropout. This suggests its potential as a drop-in replacement to standard dropout with faster training speeds. The source code is available at https://github.com/andylolu2/sparse-dropou

    375,182

    full texts

    623,509

    metadata records
    Updated in last 30 days.
    arXiv.org e-Print Archive is based in United States
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇