BieColl - Bielefeld eCollections
Not a member yet
    1004 research outputs found

    Cross-Modal Learning of Visual Categories using Different Levels of Supervision

    Get PDF
    Today's object categorization methods use either supervised or unsupervised training methods. While supervised methods tend to produce more accurate results, unsupervised methods are highly attractive due to their potential to use far more and unlabeled training data. This paper proposes a novel method that uses unsupervised training to obtain visual groupings of objects and a cross-modal learning scheme to overcome inherent limitations of purely unsupervised training. The method uses a unified and scale-invariant object representation that allows to handle labeled as well as unlabeled information in a coherent way. One of the potential settings is to learn object category models from many unlabeled observations and a few dialogue interactions that can be ambiguous or even erroneous. First experiments demonstrate the ability of the system to learn meaningful generalizations across objects already from a few dialogue interactions

    Learning Responses to Visual Stimuli: A Generic Approach

    Get PDF
    A general framework for learning to respond appropriately to visual stimulus is presented. By hierarchically clustering percept-action exemplars in the action space, contextually important features and relationships in the perceptual input space are identified and associated with response models of varying generality. Searching the hierarchy for a set of best matching percept models yields a set of action models with likelihoods. By posing the problem as one of cost surface optimisation in a probabilistic framework, a particle filter inspired forward exploration algorithm is employed to select actions from multiple hypotheses that move the system toward a goal state and to escape from local minima. The system is quantitatively and qualitatively evaluated in both a simulated shape sorter puzzle and a real-world autonomous navigation domain

    Presentation Agents That Adapt to Users' Visual Interest and Follow Their Preferences

    Get PDF
    This research proposes an interactive presentation system that employs eye gaze as an intuitive and unobtrusive input modality. Eye movements are an excellent clue to users' attention, visual interest, and preference. By analyzing and interpreting eye behavior in real-time, our system can adapt to the current (visual) interest state of the user, and thus provide a more personalized and 'attentive' experience of the presentation. The system implements a virtual presentation room, where research content is presented by a team of two highly realistic 3D agents in a dynamic and interactive way. A small preliminary study was conducted to investigate users' gaze behavior with a non-interactive version of the system. A demo video based on our system was awarded as the best application of life-like agents at the GALA event in 2006

    View Independent Face Detection Based on Combination of Local and Global Kernels

    Get PDF
    In this paper, local and global kernels are combined to use the detailed and rough similarities simultaneously. In recent years, many recognition methods based on local features were proposed. However, the combination of only local matching is not sufficient. Global viewpoint is also necessary to improve the generalization ability. Local feature matching measures the detailed similarity and global feature matching measures the rough similarity. Therefore, the error pattern is different in local and global features. If they are combined well, the generalization ability is improved. In the proposed method, local kernels and global kernel are combined by summation, and the combined kernel is used in SVM. The proposed method is applied to view independent face detection task. We confirm that the false positive is reduced by combining local and global kernels. The effectiveness of the proposed method is demonstrated by the comparison with only global and local kernel

    Decision Manifolds: Classification Inspired by Self-Organization

    Get PDF
    We present a classifier algorithm that approximates the decision surface of labeled data by a patchwork of separating hyperplanes. The hyperplanes are arranged in a way inspired by how Self-Organizing Maps are trained. We take advantage of the fact that the boundaries can often be approximated by linear ones connected by a low-dimensional nonlinear manifold. The resulting classifier allows for a voting scheme that averages over the classifiction results of neighboring hyperplanes. Our algorithm is computationally efficient both in terms of training and classification. Further, we present a model selection framework for estimation of the paratmeters of the classification boundary, and show results for artificial and real-world data sets

    Robust Camera Calibration and Evaluation Procedure Based on Images Rectification and 3D Reconstruction

    Get PDF
    This paper presents a robust camera calibration algorithm based on contour matching of a known pattern object. The method does not require a fastidious selection of particular pattern points. We introduce two versions of our algorithm, depending on whether we dispose of a single or several calibration images. We propose an evaluation procedure which can be applied for all calibration methods for stereo systems with unlimited number of cameras. We apply this evaluation framework to 3 camera calibration techniques, our proposed robust algorithm, the modified Zhang algorithm implemented by J. Bouguet and Faugeras-Toscani method. Experiments show that our proposed robust approach presents very good results in comparison with the two other methods. The proposed evaluation procedure gives a simple and interactive tool to evaluate any camera calibration method

    Fixed point rules for heteroscedastic Gaussian kernel-based topographic map formation

    Get PDF
    We develop a number of fixed point rules for training homogeneous, heteroscedastic but otherwise radially-symmetric Gaussian kernel-based topographic maps. We extend the batch map algorithm to the heteroscedastic case and introduce two candidates of fixed point rules for which the end-states, i.e., after the neighborhood range has vanished, are identical to the maximum likelihood Gaussian mixture modeling case. We compare their performance for clustering a number of real world data sets

    Transform Learning - Registration of medical images using self organization

    Get PDF
    A network model is introduced that allows a multimodal registration of two images. It can be used for a image-model or a model-model registration. The application of the network to registering tomographic to 3D ultrasonic data is introduced. Results on artificial and real ultrasound image data sets are discussed

    A Motion Calculation System Based on Background Motion Modeling

    Get PDF
    Motion calculation is often a necessary pre-processing step for moving object detection and tracking. It is a challenging task when the images are taken in outdoor scenes with cameras mounted on a moving vehicle. In this paper we present an accurate and efficient motion calculation system. The accuracy of the system is achieved by estimating background motions and eliminating those pixels that have similar motions to the background motion, and by calculating motion vectors using affine image transformation with Newton-Raphson style search method under subpixel resolution. Efficiency is achieved by concentrating on the regions of interests through a coarse-to-fine process

    A Reactive Vision System: Active-Dynamic Saliency

    Get PDF
    We develop an architecture for reactive visual analysis of dynamic scenes. We specify a minimal set of system features based upon biological observations. We implement feature on a processing network based around an active stereo vision mechanism. Active rectification and mosaicing allows static stereo algorithms to operate on the active platform. Foveal zero disparity operations permit attended object extraction and ensures coordinated stereo fixation upon visual surfaces. Active-dynamic inhibition of return, and task dependent biasing result in a flexible, preemptive and retrospective system that responds to unique visual stimuli and is capable of top-down modulation of attention towards regions and cues relevant to tasks

    696

    full texts

    1,004

    metadata records
    Updated in last 30 days.
    BieColl - Bielefeld eCollections
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇