5632 research outputs found
Sort by
AutoBiasTest: Controllable Sentence Generation for Automated and Open-Ended Social Bias Testing in Language Models
Social bias in Pretrained Language Models (PLMs) affects text generation and other downstream NLP tasks. Existing bias testing methods rely predominantly on manual templates or on expensive crowd-sourced data. We propose a novel AutoBiasTest method that automatically generates sentences for testing bias in PLMs, hence providing a flexible and low-cost alternative. Our approach uses another PLM for generation and controls the generation of sentences by conditioning on social group and attribute terms. We show that generated sentences are natural and similar to human-produced content in terms of word length and diversity. We illustrate that larger models used for generation produce estimates of social bias with lower variance. We find that our bias scores are well correlated with manual templates, but AutoBiasTest highlights biases not captured by these templates due to more diverse and realistic test sentences. By automating large-scale test sentence generation, we enable better estimation of underlying bias distributions
VoxFormer: Sparse Voxel Transformer for Camera-based 3D Semantic Scene Completion
Humans can easily imagine the complete 3D geometry of occluded objects and scenes. This appealing ability is vital for recognition and understanding. To enable such capability in AI systems, we propose VoxFormer, a Transformer-based semantic scene completion framework that can output complete 3D volumetric semantics from only 2D images. Our framework adopts a two-stage design where we start from a sparse set of visible and occupied voxel queries from depth estimation, followed by a densification stage that generates dense 3D voxels from the sparse ones. A key idea of this design is that the visual features on 2D images correspond only to the visible scene structures rather than the occluded or empty spaces. Therefore, starting with the featurization and prediction of the visible structures is more reliable. Once we obtain the set of sparse queries, we apply a masked autoencoder design to propagate the information to all the voxels by self-attention. Experiments on SemanticKITTI show that VoxFormer outperforms the state of the art with a relative improvement of 20.0% in geometry and 18.1% in semantics and reduces GPU memory during training by ~45% to less than 16GB. Our code is available on this https://github.com/NVlabs/VoxFormer
Eventual Discounting Temporal Logic Counterfactual Experience Replay
Linear temporal logic (LTL) offers a simplified way of specifying tasks for policy optimization that may otherwise be difficult to describe with scalar reward functions. However, the standard RL framework can be too myopic to find maximally LTL satisfying policies. This paper makes two contributions. First, we develop a new value-function based proxy, using a technique we call eventual discounting, under which one can find policies that satisfy the LTL specification with highest achievable probability. Second, we develop a new experience replay method for generating off-policy data from on-policy rollouts via counterfactual reasoning on different ways of satisfying the LTL specification. Our experiments, conducted in both discrete and continuous state-action spaces, confirm the effectiveness of our counterfactual experience replay approach
Smoothed Online Optimization with Unreliable Predictions
We examine the problem of smoothed online optimization, where a decision maker must sequentially choose points in a normed vector space to minimize the sum of per-round, non-convex hitting costs and the costs of switching decisions between rounds. The decision maker has access to a black-box oracle, such as a machine learning model, that provides untrusted and potentially inaccurate predictions of the optimal decision in each round. The goal of the decision maker is to exploit the predictions if they are accurate, while guaranteeing performance that is not much worse than the hindsight optimal sequence of decisions, even when predictions are inaccurate. We impose the standard assumption that hitting costs are globally α-polyhedral. We propose a novel algorithm, Adaptive Online Switching (AOS), and prove that, for a large set of feasible δ > 0, it is (1+δ)-competitive if predictions are perfect, while also maintaining a uniformly bounded competitive ratio of 2^(O̅(1/(αδ))) even when predictions are adversarial. Further, we prove that this trade-off is necessary and nearly optimal in the sense that any deterministic algorithm which is (1 + δ)-competitive if predictions are perfect must be at least 2^(O̅(1/(αδ)))-competitive when predictions are inaccurate. In fact, we observe a unique threshold-type behavior in this trade-off: if δ is not in the set of feasible options, then no algorithm is simultaneously (1 + δ)-competitive if predictions are perfect and ζ-competitive when predictions are inaccurate for any ζ < ∞. Furthermore, we discuss that memory is crucial in AOS by proving that any algorithm that does not use memory cannot benefit from predictions. We complement our theoretical results by a numerical study on a microgrid application
Connectome-constrained deep mechanistic networks predict neural responses across the fly visual system at single-neuron resolution
We can now measure the connectivity of every neuron in a neural circuit, but we are still blind to other biological details, including the dynamical characteristics of each neuron. The degree to which connectivity measurements alone can inform understanding of neural computation is an open question. Here we show that with only measurements of the connectivity of a biological neural network, we can predict the neural activity underlying neural computation. We constructed a model neural network with the experimentally determined connectivity for 64 cell types in the motion pathways of the fruit fly optic lobe but with unknown parameters for the single neuron and single synapse properties. We then optimized the values of these unknown parameters using techniques from deep learning, to allow the model network to detect visual motion. Our mechanistic model makes detailed experimentally testable predictions for each neuron in the connectome. We found that model predictions agreed with experimental measurements of neural activity across 24 studies. Our work demonstrates a strategy for generating detailed hypotheses about the mechanisms of neural circuit function from connectivity measurements. We show that this strategy is more likely to be successful when neurons are sparsely connected---a universally observed feature of biological neural networks across species and brain regions
Maximum Mutational Robustness in Genotype-Phenotype Maps Follows a Self-similar Blancmange-like Curve
Phenotype robustness, defined as the average mutational robustness of all the genotypes that map to a given phenotype, plays a key role in facilitating neutral exploration of novel phenotypic variation by an evolving population. By applying results from coding theory, we prove that the maximum phenotype robustness occurs when genotypes are organised as bricklayer's graphs, so called because they resemble the way in which a bricklayer would fill in a Hamming graph. The value of the maximal robustness is given by a fractal continuous everywhere but differentiable nowhere sums-of-digits function from number theory. Interestingly, genotype-phenotype (GP) maps for RNA secondary structure and the HP model for protein folding can exhibit phenotype robustness that exactly attains this upper bound. By exploiting properties of the sums-of-digits function, we prove a lower bound on the deviation of the maximum robustness of phenotypes with multiple neutral components from the bricklayer's graph bound, and show that RNA secondary structure phenotypes obey this bound. Finally, we show how robustness changes when phenotypes are coarse-grained and derive a formula and associated bounds for the transition probabilities between such phenotypes
A dual sgRNA library design to probe genetic modifiers using genome-wide CRISPRi screens
The ability to map genetic interactions has been essential for determining gene function and defining biological pathways. Therefore, a system to readily perform genome-wide genetic modifier screens in human cells is a powerful platform for dissecting complex processes in mammalian cells, where redundancy and adaptation commonly mask the phenotype of a single genetic perturbation. Here, we report a CRISPR interference (CRISPRi) based platform, compatible with Fluorescence Activated Cell Sorting (FACS)-based reporter screens, that can be used to query epistatic relationships at scale. This is enabled by a flexible dual-sgRNA library design that allows for the simultaneous delivery and selection of a fixed sgRNA and a second randomized guide, comprised of a genome-wide library, with a single transduction. As a proof of principle, we apply our approach to study the pathways that mediate tail-anchored (TA) protein insertion at the endoplasmic reticulum (ER). We show that this dual-guide library approach can be successfully coupled with FACS-based reporter screening, to identify genetic epistasis and thereby place TA biogenesis factors in their respective parallel pathways. We demonstrate that this dual-guide approach is both more sensitive and specific than traditional growth screening approaches, and is ideally suited for dissecting the complex interplay between factors in human cells
Interactive computational and experimental approaches improve the sensitivity of periplasmic binding protein-based nicotine biosensors for measurements in biofluids
To develop more sensitive fluorescent protein sensors for nicotine, we combined computational protein design, site-saturated, site-directed, and combinatorial mutagenesis with fluorescence assays, molecular dynamics simulations, and absorbance measurements. The data showed that the resulting molecules, iNicSnFR11 and iNicSnFR12, have higher sensitivity to nicotine than previously reported constructs. In the linear portion of the dose-response relation at sub-μM [nicotine] for iNicSnFR12, ∆F/F₀ increased with a proportionality constant (S-slope) of 2.6 μM⁻¹, representing a 6.5-fold higher sensitivity than iNicSnFR3a. Molecular dynamics calculations enabled identification of a binding pose for nicotine previously indeterminate from experimental data. Further comparative simulations based on this model revealed a tilt in helix 4 in the optimized sensor, likely altering allosteric networks involving the ligand binding site. The absorbance data showed that the fluorescence activation results from increased absorption rather than increased quantum yield for fluorescence. iNicSnFR12 resolved nicotine in diluted mouse and human serum at the peak concentration (100-200 nM) that occurs during smoking or vaping, but also at the decaying concentrations (< 100 nM) during the intervals between smoking or vaping sessions. NicSnFR12 was roughly as sensitive to varenicline or acetylcholine as to nicotine; the sensitivity to choline was at least one order of magnitude less. None of these drugs would markedly distort measurements in human biofluids such as sweat and interstitial fluid. Therefore, iNicSnFR12 is a promising candidate as the molecular sensor that could underlie a continuous nicotine monitor for human biofluids
Mechanistic modeling with a variational autoencoder for multimodal single-cell RNA sequencing data
We motivate and present biVI, which combines the variational autoencoder framework of scVI with biophysically motivated, bivariate models for nascent and mature RNA distributions. In simulated benchmarking, biVI accurately recapitulates key properties of interest, including cell type structure, parameter values, and copy number distributions. In biological datasets, biVI provides a route for the identification of the biophysical mechanisms underlying differential expression. The analytical approach outlines a generalizable strategy for representing multimodal datasets generated by single-cell RNA sequencing
Engineering View of Gravitation
We describe here an internally-consistent, Quantum coupled treatment of gravitation and electromagnetism. The electromagnetic part is put forth in Collective Electrodynamics—Quantum Foundations of Electromagnetism [44], hereinafter referred to as CE, and described briefly in Section 1.3.1. The Gravitational theory described in this document, which we call G4v, is a direct extension of Einstein’s 1911/12 approach. It differs from previous attempts in a number of important ways:
• Neither the electromagnetic nor gravitational fields are quantized.
The wave functions and four-potentials are continuous functions of space and time.
Quantization results from the interaction of matter and field wave functions.
• The theory is based on Mach’s Principle and provides a conceptual base for the Equivalence Principle.
• It is not a metric theory; it is formulated in flat space-time.
Lengths are constant and do not vary with gravitation potential
• The speed of light c is equal to the gravitational scalar potential.
It is not constant, but varies with position and time.
• The quantity of matter coupled gravitationally is not the mass m.
It is the Compton wave number k0 = mc/~.
• The theory is based on four-vector coupling.
It is thus locally Lorentz-invariant in regions where the speed of light can be considered constant.
• The source of the electrical four-potential is the charge–current density four-vector, and that for the
gravitational four potential is the energy-momentum four-vector. Both quantities are defined for the wave
function of the source matter, and appear as terms in the affected matter wave function.
The concept of force is not necessary, but can be computed if desired