1,721,005 research outputs found
Advancing machine learning in astrophysics
This thesis explores four projects applying supervised deep learning to help answer astrophysical questions.
I first consider faint tidal features. Tidal features are a long-lasting signature of galaxy mergers, making them useful for measuring merger rates. However, current automated methods struggle to detect faint tidal features in complex galaxies. I use convolutional neural networks to identify galaxies with tidal features in the CFHTLS- Wide Survey, improving on previous methods applied to the same dataset. I show that my networks can identify which pixels are associated with tidal features, potentially enabling researchers to not only identify but also characterise tidal features.
I then turn to Galaxy Zoo, a citizen science project measuring galaxy morphology. Galaxy Zoo is being gradually outpaced by the increasing scale of new surveys. Au- tomated classifiers can be trained using volunteer responses; however, such classifiers are often unable to consider uncertainty in either volunteer responses or predictions, leading to wasted volunteer effort and overconfident classifications. I introduce a probabilistic approach that allows classifiers to flexibly express uncertainty. I use this probabilistic approach to build a machine learning system that ‘asks’ volunteers to label the galaxies it could best learn from. I relaunch Galaxy Zoo with images from the Dark Energy Camera Legacy Survey and run my system live, collecting 1.8 million volunteer responses. My final models are around 99% accurate on every question for galaxies with confident volunteer answers and are otherwise correctly uncertain.
Next, I help the Canadian Hydrogen Intensity Mapping Experiment (CHIME) detect fast radio bursts. CHIME only attempts to detect FRB above a signal-to- noise threshold of σ = 8.5, in part for lack of expert time to review candidates. I created and launched a citizen science project to classify the 7.8 ≤ σ < 8.5 signal-to- noise candidates detected by CHIME each week. Candidates found by this project may be the most distant fast radio bursts ever detected, which I hope will serve as useful cosmological probes of the intergalactic medium.
Finally, I show that neural network emulation can efficiently recover posteriors of galaxy parameters from photometry. Galaxy SED simulators are too slow to use MCMC inference on large samples. I train a neural network to emulate an SED simulator, providing both faster likelihood evaluations and known gradients. These gradients can then be used for efficient Hamiltonian Monte Carlo inference.
Together, these projects show how deep learning can help astronomers make ef- fective use of limited and uncertain data
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Finding radio transients
Modern radio telescopes are data-intensive machines, producing many TB of data every night. Amongst this deluge of data are transient and variable phenomena, whose study can shed new light on processes as varied as stellar dynamos and the accretion discs in supermassive black holes. In this work I demonstrate the applicability of different methods to the discovery of these astrophysical transients and variables coming from telescopes such as MeerKAT.
I first consider a standard approach to discovering transients by characterising their variability. By making use of even modest sampling with the high sensitivity and wide field of view of MeerKAT, I demonstrate how we are now able to uncover new transients almost by accident - if we exclude the vast amount of time spent planning, building and operating excellent telescopes, efficient pipelines and well- crafted observing proposals. In this work I found a stellar flare from a nearby M dwarf, which was then followed up and complemented by optical and X-ray photometry and spectroscopy, providing new insights on the system.
Next I built a citizen science platform in order to perform such transient searches at scale, making use of a wide range of data available in the MeerKAT archive. I detail the process of review and beta-testing that resulted in the final design of the Bursts from Space: MeerKAT project. Over 1000 volunteers took part, demonstrating a healthy appetite for further Zooniverse data releases. Volunteers discovered or recovered a wide range of phenomena, from flare stars and pulsars to scintillating AGN and transient OH maser emission. I was also able to use the known transients in our fields to understand some reasons why interesting sources may be missed and will fold this learning through to future iterations of the project. This is the first demonstration of volunteers finding radio transients in images.
Finally, I show how anomaly detection, an unsupervised machine learning approach, is a suitable tool for finding these variable phenomena at scale, as is required for modern astronomical surveys. I use three feature sets as applied to two anomaly detection techniques in the Astronomaly package and analyse anomaly detection performance by comparison with citizen science labels. By using transients found by citizen scientists as a ground truth I demonstrate that anomaly detection techniques can recall over half of the radio transients within 10% of the sample dataset. I find that the choice of feature set is crucial, especially when considering available resources for human inspection and follow-up. I find that active learning on ∼2% of the data improves recall by up to 10%, depending on the feature-model pair. The best performing feature-model pairs result in a factor of 5 times fewer sources requiring vetting by humans. This is the first effort to apply anomaly detection techniques to finding radio transients and shows great promise for application to other datasets, a real-time transient detection system and upcoming large surveys
Using the penguin watch network to examine phenology and breeding success in pygoscelis penguins
Remote time-lapse cameras can facilitate high-frequency monitoring of wild animal populations over large spatial and temporal scales, yet the volume of imagery they produce can be overwhelming for small research teams. As ecological studies transition towards a ‘big data’ approach, we must develop ways to deal with the ‘data deluge’, in the form of integrated processing frameworks. In this thesis, I present and examine data-processing pipelines for the Penguin Watch project – a network of approximately 80 time-lapse cameras positioned in penguin colonies across the Scotia Arc region of the Southern Ocean. I show that both the Penguin Watch online citizen science project and the Pengbot computer vision algorithm are efficient and reliable image annotation tools, which can be used in combination to maximise output and minimise error. I present new methodologies to extract biologically meaningful metrics from citizen science annotations, and show that spatial data – for example ‘nearest neighbour distances’ – can be used to examine behaviour, such as chick huddling.
Seabirds, including penguin species, are often considered indicators of marine health. As such, effective monitoring of their populations offers an invaluable opportunity to assess the health of the wider marine ecosystem. An essential part of this process is obtaining baseline data, and here I use the Penguin Watch project to extract phenological dates for a number of new Chinstrap penguin (Pygoscelis antarcticus) colonies and seasons – a species which has historically been understudied. Through a ‘virtual’ mark-recapture study I examine breeding success in Gentoo penguin (P. papua) colonies across a substantial part of their range in the Atlantic sector of the Southern Ocean. Investigation of the impact of three key anthropogenic stressors (precipitation, the krill fishery, and tourism) on breeding success revealed no significant effects over the study period. However, there is evidence that the relationship between chick survival and precipitation could be highly non-linear, with severe precipitation events associated with high chick mortality. This has implications for a number of Southern Ocean bird species, and with precipitation levels expected to increase under climate-change scenarios, continued monitoring is essential. Repetition of this study on Adélie (P. adeliae) and Chinstrap penguins – species which have a higher dependency on krill than Gentoos – is a key future research priority
People powered planet hunting with TESS
To date, over 5000 exoplanets have been confirmed, and studies of their characteris- tics have unveiled an extremely wide range of masses, sizes, system architectures, and orbital periods. More than two-thirds of these known planets were identified using dedicated planet detection algorithms that search for transit events in data obtained by space-based survey satellites, such as the Transiting Exoplanet Survey Satellite (TESS). These automated transit detection algorithms are, however, strongly biased towards the detection of short-period planets. For example, 90% of the pipeline- detected TESS planet candidates known to date have orbital periods shorter than 27 days. In this thesis, I show that we can address this period bias by visually inspecting TESS data with the help of a global community of over 30 000 volunteers.
I present the Planet Hunters TESS (PHT) citizen science project that I designed, built, and continue to manage. Throughout this thesis I demonstrate that with visual vetting we are sensitive to detecting both short- and longer-period planets. This makes PHT an effective tool to help populate under-explored regions of parameter space that have limited prospects for being studied by automated searches that typically require a minimum of two transit events for detection. To that end, the PHT planet sample can benefit studies of planet occurrence rates, as well as inform theories of physical processes involved with the formation and evolution of different types of exoplanets. I also present LATTE; a new open-source vetting suite that includes numerous diagnostic tests to aid in the detection and characterisation of planetary and stellar signals. The tool, which is regularly used by both professional astronomers and citizen scientists, forms part of the PHT pipeline and helps identify promising candidates for ground-based follow-up.
The large-scale visual vetting with PHT, combined with the LATTE vetting suite, has led to the discovery of 139 new planet candidates in the two-minute cadence data from the first three years of TESS data alone. In addition to discussing this ensemble sample of new planet candidates, I present a detailed analysis of two of these systems. First, I discuss a Saturn-sized planet orbiting around a bright subgiant on an ~84-day orbit. The planet’s relatively long orbital period combined with the evolved nature of the host star places this planet in a relatively under explored region of parameter space and is, therefore, an exciting target for further characterisation. Second, I present a two-planet system orbiting around a bright G dwarf. Preliminary mass estimates derived from ground-based radial velocity observations, combined with detailed transit modelling, suggest that both planets in this second system have low bulk densities and therefore extended H/He atmospheres. This makes both planets prime candidates for future atmospheric characterisation and comparative planetology.
Finally, I show that visual inspection of large amounts of data via PHT yields many scientifically valuable by-products, including the identification of 4584 eclipsing binary candidates and multi-stellar systems. I discuss one of these systems - a massive compact hierarchical triple star system - in more detail and present a dynamical analysis of its past, present, and future. Overall, this thesis demonstrates the scientific value of citizen science for identifying unique signals that are often missed by automated searches and, therefore, shows that people-powered planet hunting can play an important role in a world that is becoming increasingly automated
Interstellar objects in a galactic context
Interstellar objects (ISOs) form a vast Galaxy-spanning population with an number density ∼ 1015 pc−3 in the Solar neighbourhood. Though only two members of this population have been observed, 1I/‘Oumuamua and 2I/Borisov, the Vera C. Rubin’s Legacy Survey of Space and Time (LSST) which begins later this year is expected to increase our sample of known ISOs by an order of magnitude. In this thesis I model the Milky Way’s ISO population, across the Galactic disk, in the solar neighbourhood and streaming through the Solar System, in order to predict the properties of the LSST ISO sample. These predictions combine models of planetary and Galactic processes, and using a Bayesian framework I develop my predictions can be compared to the LSST ISO sample that is actually observed. By doing this, inferences can be drawn about any process modelled, allowing the study of phenomena across astrophysical scales with an entirely different set of biases to traditional observational methods.
Using debiased data from the the APOGEE Milky Way red giant survey I build a model of what the Galactic stellar population would be if stars did not die, naming this the sine morte population. By combining this with a protoplanetary disk chemical model I predict the ISO water mass fraction distribution and how this changes with Galactocentric radius, finding that the known stellar metallicity gradient has a matching ISO composition gradient. I demonstrate the Bayesian method of comparing my predictions to observed ISOs by placing constraints, albeit weak ones, on the metallicity dependence of ISO production using the measured composition of 2I/Borisov.
Next I predict the joint composition and velocity distribution, or chemodynamics, of ISOs in the solar neighbourhood. I debias a sample of stars from the Gaia survey within 200 pc, reconstruct the sine morte stellar population, then from this predict the ISO distribution using the same protoplanetary disk chemical model. I demonstrate that, like the stellar velocity distribution, the velocity distribution of ISOs is complex, with furrows and overdensities caused by orbital resonances with the Galactic spiral arms and bar, as well as correlated with properties such as age and composition. Both known ISOs have typical velocities within the boundaries of my predicted overdensities, and they share their velocities with stars with a range of ages and compositions. Finally I demonstrate that my predicted velocity distribution could be distinguished from simpler Gaussian approximations used in previous works with ISO sample sizes within the capability of the LSST.
I devise a fast method of sampling observable ISO orbits through the Solar System, for use in simulation of surveys discovering ISOs. The innovations of this method significantly speed up survey simulation by avoiding sampling orbits which have no chance of being observed, allowing longer surveys and wider ranges of ISO sizes to be simulated, as well as increasing the statistical quality of the results. I also present the results of this survey simulation: the LSST ISO sample will consist of 6–51 ISOs dependant on the size distribution, most of them larger than 1I, with a slight discovery bias towards slower velocities relative to the Sun.
Changing tack slightly, I present a hypothesis that collisions between ISOs and neutron stars may be the cause of a significant fraction of fast radio bursts (FRBs), millisecond-length radio transients with many suggested origins. Such collisions are expected to cause FRB-like signals, and I demonstrate that they occur across the local Universe at a comparable volumetric rate to observable FRBs. I further argue that the rate of these collisions will have the same evolution with redshift as is observed in FRBs, and discuss other implications of the hypothesis.
Then changing tack quite significantly, I finish by using my APOGEE-based sine morte stellar population model to predict the distribution of single stellar black holes in the Milky way. I predict that though the Milky Way black hole population will show a strong gradient in occurrence relative to stars, this will only have a marginally detectable effect on the yield of the Nancy Grace Roman Space Telescope’s microlensing survey relative to larger uncertainties such as in the fraction of stars born which form black holes on their death
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
