165 research outputs found
Time-based alignment of video recordings: a service and an application
Video content is uploaded and shared by users on video sharing websites in vast scale. It is estimated, for example, that every minute 24 hours of video is uploaded to YouTube. Many of the clips are captured at live events, for instance, a U2 concert in Giants Stadium. Increasingly, then, individuals attending the same events upload related content: there are over three hundred YouTube clips from the said concert. This overload makes relevant and interesting videos harder to find, and the event content harder to view and understand. An innovative approach to present the video content is thus necessary. In this thesis, we propose a solution to tackle the problem above. We use the audio fingerprinting algorithm to find overlapping video clips within events, and time-align these overlapping videos based on their audio tracks. We developed a highly interactive video player that organizes and presents the time-aligned video content. The player integrates social data like view counts to help people create a better understanding of the event content, and improve the viewing experience and seeking of clips. We conducted user study sessions to understand user’s interaction and get user feedback for the video player. The web is adopting open standard these days. The work in this thesis follows this trend by providing API services. We designed and built a set of standard API services to expose the underlying audio fingerprinting. In this way other developers can utilize our system remotely and programmatically. They can also build applications using time-align data.M.S.Includes bibliographical referencesby Zihao Y
Online resource allocation for service platforms and bandit experiments
Online resource allocation problems have attracted considerable attention in today’s data-driven world, where decision-making processes often rely on real-time information and samples. Effective algorithms must strike a delicate balance between exploration and exploitation. In this thesis, we contribute decision-making algorithms for two online resource allocation problems. In the first problem, online service platform manages heterogeneous servers to sequentially serve online customers with unknown reward. We integrate multi-armed bandit algorithms into resource allocation to explore and exploit the customer reward information under service resource constraint. Based on the new algorithms, we propose a three-stage architecture for the platform’s long-run operation. In the second problem, an agent sequentially allocates measurement efforts to select the best set of arms with the highest means in a multi-armed bandit experiment. We characterize the necessary and sufficient conditions for the optimal allocation problem. We reveal the connection between the Karush–Kuhn–Tucker conditions and top-two algorithm design principle, initially proposed for best-arm-identification. We propose a simple and effective selection rule dubbed information-directed selection (IDS) that selects one of the top-two candidates based on a measure of information gain. Extensive numerical experiments show the superior performance of the proposed top-two algorithms with IDS.</p
Nanoscale self-assembly and nonequilibrium transformation of colloidal nanoparticles visualized by liquid-phase TEM
Liquid-phase transmission electron microscopy (TEM) has been widely used for probing solution-phase nanoscale dynamics, such as nanoparticle growth/corrosion, electrochemical processes, and aggregation of soft materials (e.g., micelles, proteins). Among them, one particularly interesting direction is to resolve and understand the self-assembly pathways of nanosized building blocks, namely individual one of them diffusing and interacting with each other to form into functional superstructures. Complicated, nonclassical pathways occur as originated from the nonadditivity of nanoscale interactions and the resultant complex free-energy landscapes for nanosized entities. During my Ph. D., my research has provided insights in charting the energy diagram, fundamental interactions, formation of prenucleation precursors and interfacial energy of crystals using in-situ liquid-phase TEM. Such scientific understanding is enabled by a statistical mechanics based conceptual framework to extract spatiotemporal information from the liquid-phase TEM movies, which can be generalizable to a broad range of phase behavior studies for materials at the nanoscale.
We used a model system of nanoparticles to study their crystallization pathways into well-shaped supracrystals. We first quantitatively studied the electron induced beam effect in aqueous environment and developed a protocol to control the self-assembly dynamics of nanoparticles inside liquid-phase TEM. Then, the full transition of crystallization was mapped from single nanoparticle towards a large-scale supracrystal, and the density and structure took place separately in time, following a nonclassical two-step crystallization. I elucidated the origin of this two-step crystallization by experimentally measuring the free-energy barrier for nucleation based on the statistical distribution of transient clusters at initial stage. During the growth process, the crystal surface fluctuates due to thermally excited capillary waves, based on which we were able to measure the interfacial energy of such supracrystals, again enabled by our in-situ liquid-phase TEM imaging. The supracrystal surface has a roughness driven by the thermal capillary waves, which we validated for the first time at the nanoscale.
In summary, we have demonstrated the capability of utilizing liquid-phase TEM to investigate various spatiotemporal fluctuating phenomena at the nanoscale. Fast diffusion of nanoparticles enables investigation over various collective behaviors, such as glass transition, grain-boundary migration and aggregate coalescence, which have been challenging to probe due to the extremely long relaxation time. Apart from inorganic nanoparticles, the protocol can be extended to study other spatially heterogenous phenomena, where nanoscale dynamics play a critical role, such as the penetration of nanoparticles through a porous membrane and degradation dynamics of polymeric materials into micro/nanoparticles. Real-space investigations will offer insight into such physical processes, improving coating recipe engineering and efficiency of drinking water filtration. Such methodologies can also be applied to optical microscopy and impact the understanding of other materials systems, including a synthesis–imaging–analysis platform I developed to visualize the morphology transformation of metal–organic framework crystals during chemical etching.Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2022-05-01The student, Zihao Ou, accepted the attached license on 2020-05-01 at 18:55.The student, Zihao Ou, submitted this Dissertation for approval on 2020-05-01 at 19:07.This Dissertation was approved for publication on 2020-05-05 at 10:59.DSpace SAF Submission Ingestion Package generated from Vireo submission #15163 on 2020-08-25 at 17:42:31Made available in DSpace on 2020-08-27T00:50:16Z (GMT). No. of bitstreams: 3
OU-DISSERTATION-2020.pdf: 11176382 bytes, checksum: a36bf2c948ba55a3c4797adcca333cff (MD5)
LICENSE.txt: 4205 bytes, checksum: 60ee2828f9c70a9cf82da7bc36d0b199 (MD5)
PROQUEST_LICENSE.txt: 4551 bytes, checksum: cef8fb7b6c9d087a0bab8ca3b9ee1bc5 (MD5)
Previous issue date: 2020-05-05Embargo set by: Seth Robbins for item 115915
Lift date: 2022-08-27T00:50:22Z
Reason: Author requested closed access (OA after 2yrs) in Vireo ETD systemEmbargo set by: Seth Robbins for item 115915
Lift date: 2022-08-27T00:51:40Z
Reason: Author requested closed access (OA after 2yrs) in Vireo ETD systemAuthor requested closed access (OA after 2yrs) in Vireo ETD systemLimite
Adapting Subtitles:Towards Accessible Subtitled Media for People with Aphasia
Subtitles can be consumed via TVs, laptops, and smartphones, which play a key role in many people’s daily lives. However, this marginalises people with complex accessibility needs in the context of comprehension, such as people with aphasia, a communication and language impairment. To address this accessibility challenge with a human-centred design approach, I started by building rapport with the community of focus and involving them in all stages of the design process. This paper reports on ongoing doctoral work that explores how personalising and customising subtitled media can fulfil complex accessibility needs for communities with aphasia, by delineating the research background and motivation, the accomplished work, and the proposed next steps
Demystifying LLM Attacks And Defense: A Comprehensive Study with Improved Attack Technique
Large Language Models (LLMs) have emerged as pivotal in content generation, offering profound societal impacts. Previous research has highlighted their propensity to generate content that breaches societal norms. Misuse of LLMs poses significant ethical concerns, including misinformation spread, social unrest, and political manipulation. To mitigate such risks, safety training techniques have been employed, instructing LLMs to avoid generating harmful content during inference time. Nonetheless, securing LLMs against \textit{Prompt Injection} and \textit{Jailbreak} attacks remains challenging, as evidenced by recent studies and abundant malicious instructions available online. To make things worse, these attacks are normally transferable due to their format in natural language, posing substantial security threats, as people without AI could also use these attacks. Although various defense techniques exist, their effectiveness against diverse attacks is largely untested.This thesis, therefore, provides the first comprehensive evaluation of the interplay between attack techniques and defense techniques, focusing particularly on the \textit{Jailbreak} type. Our analysis encompasses nine different attack methodologies and seven defense techniques, applied to three unique LLMs: Vicuna, LLama, and GPT-3.5 Turbo, with the objective of assessing their efficacy. Our results indicate that white-box attacks are generally less effective than universal approaches and that the inclusion of particular tokens in the input can significantly influence the success rate of attacks.Moreover, we identify that research into the vulnerabilities presented by continuous embeddings has been scant, with prior approaches mainly relying on the addition of discrete or continuous suffixes to prompts. Our investigation introduces a new approach for direct attacks on LLM inputs that bypasses the necessity for suffix appending or posing specific questions, as long as the output is pre-specified. We also notice that improper initialization of random continuous input or an excessive amount of iterations can lead to overfitting scenarios. To address this, we suggest an effective method, termed \textbf{Clip}, to alleviate the issue of overfitting.In conclusion, we contribute to the field by conducting the first study of the interaction between attack and defense techniques and by presenting a benchmark through our shared datasets, an easily integrable testing framework, and an attack algorithm to encourage further investigation into the security of LLMs.Computer Scienc
ChatCAD+: Towards a Universal and Reliable Interactive CAD using LLMs
The integration of Computer-Aided Diagnosis (CAD) with Large Language Models
(LLMs) presents a promising frontier in clinical applications, notably in
automating diagnostic processes akin to those performed by radiologists and
providing consultations similar to a virtual family doctor. Despite the
promising potential of this integration, current works face at least two
limitations: (1) From the perspective of a radiologist, existing studies
typically have a restricted scope of applicable imaging domains, failing to
meet the diagnostic needs of different patients. Also, the insufficient
diagnostic capability of LLMs further undermine the quality and reliability of
the generated medical reports. (2) Current LLMs lack the requisite depth in
medical expertise, rendering them less effective as virtual family doctors due
to the potential unreliability of the advice provided during patient
consultations. To address these limitations, we introduce ChatCAD+, to be
universal and reliable. Specifically, it is featured by two main modules: (1)
Reliable Report Generation and (2) Reliable Interaction. The Reliable Report
Generation module is capable of interpreting medical images from diverse
domains and generate high-quality medical reports via our proposed hierarchical
in-context learning. Concurrently, the interaction module leverages up-to-date
information from reputable medical websites to provide reliable medical advice.
Together, these designed modules synergize to closely align with the expertise
of human medical professionals, offering enhanced consistency and reliability
for interpretation and advice. The source code is available at
https://github.com/zhaozh10/ChatCAD.Comment: Authors Zihao Zhao, Sheng Wang, Jinchen Gu, Yitao Zhu contributed
equally to this work and should be considered co-first author
RETRACTED ARTICLE: New Co(II)-organic framework for cyanosilylation reactions and treatment effect against esophageal cancer by inhibiting cancer cell viability, migration, and invasion
We, the Editors and Publisher of the Journal of Coordination Chemistry, have retracted the following article: Hui Zan, Yongdong Huang, Changgao Wang, Zihao Liu, Nan Du & Xiaoman Wang. New Co(II)-organic framework for cyanosilylation reactions and treatment effect against esophageal cancer by inhibiting cancer cell viability, migration, and invasion, Journal of Coordination Chemistry; 2020, VOL. 73, NO. 9, 1464–1477. DOI: 10.1080/00958972.2020.1788001 Since publication, concerns have been raised about the integrity of the data in the article. When approached for an explanation, the authors have been unable to address the concerns raised and have not been able to provide sufficient supporting information. As verifying the validity of published work is core to the integrity of the scholarly record, we are therefore retracting the article. The corresponding author listed in this publication has been informed. The authors do not agree with the retraction. We have been informed in our decision-making by our policy on publishing ethics and integrity and the COPE guidelines on retractions. The retracted article will remain online to maintain the scholarly record, but it will be digitally watermarked on each page as ‘Retracted’.</p
Empowering Communities of Aphasia through Adaptable Subtitled Media:The Altering Factors, Accessibility Barriers, and Inclusive Design Practices
Cross-layer design benchmark for throughput maximization with fairness and delay constraints in DCF systems
- …
