Luận án theo dõi và lập chỉ mục khuôn mặt trong chuỗi video 3D - Telecom ParisTECH
Phân tích và chỉ mục khuôn mặt trong chuỗi video giúp xác định và theo dõi chuyển động khuôn mặt trong video.
Telecom ParisTech
Computer Vision
Luan An
Luận án tiến sĩ
Năm xuất bản
Số trang
131
Thời gian đọc
20 phút
Lượt xem
0
Lượt tải
0
Phí lưu trữ
40 Point
Tổng quan nhanh
- Chủ đề:
- Theo dõi Khuôn mặt 3D: Giải pháp cho Video 3D
- Số trang:
- 131 trang
- Trường:
- Telecom ParisTech
- Chuyên ngành:
- Computer Vision
- Tác giả:
- Ngoc-Trung Tran
- Năm:
- 2015
Tóm tắt nội dung luận án
I.Theo dõi Khuôn mặt 3D Giải pháp cho Video 3D
Theo dõi khuôn mặt trong chuỗi video là một vấn đề quan trọng trong thị giác máy tính. Ứng dụng rộng rãi trong giám sát video, giao diện người-máy, và sinh trắc học. Lĩnh vực này có nhiều tiến bộ nhưng vẫn còn thách thức. Đặc biệt khi ước tính đồng thời 6 bậc tự do (3D tịnh tiến và xoay) và các điểm mốc (để nắm bắt hoạt ảnh khuôn mặt). Khó khăn phát sinh từ biến đổi ánh sáng, xoay đầu rộng, biểu cảm, che khuất, và nền lộn xộn. Một khung theo dõi 3D được xây dựng cho camera đơn sắc. Khung này ước tính chính xác 6 bậc tự do và hoạt ảnh khuôn mặt. Nó cũng có khả năng hoạt động mạnh mẽ với xoay đầu rộng, kể cả ở góc nghiêng. Phương pháp này áp dụng mô hình khuôn mặt 3D. Mô hình này là một tập hợp các đỉnh 3D, cần thiết để tính toán các giá trị 6 bậc tự do liên tục trong chuỗi video.
1.1. Ước tính tư thế đầu và hoạt ảnh khuôn mặt
Ước tính tư thế đầu 3D và hoạt ảnh khuôn mặt đồng thời vẫn là một nhiệm vụ phức tạp. Các thách thức chính bao gồm sự thay đổi ánh sáng đột ngột, góc quay đầu rộng và các biểu cảm khuôn mặt đa dạng. Một khung theo dõi 3D được phát triển để giải quyết những vấn đề này. Nó sử dụng camera đơn sắc để ước tính chính xác 6 bậc tự do. Đồng thời, nó cũng theo dõi các chuyển động hoạt ảnh của khuôn mặt. Khung này được thiết kế để duy trì độ bền vững ngay cả khi khuôn mặt quay góc rộng hoặc ở tư thế nghiêng. Công nghệ quét sâu có thể hỗ trợ việc này trong tương lai.
1.2. Vận dụng Mô hình Khuôn mặt 3D cho Chuỗi Video
Mô hình khuôn mặt 3D được áp dụng để theo dõi hiệu quả. Mô hình này biểu diễn khuôn mặt dưới dạng một tập hợp các đỉnh 3D. Các đỉnh này rất quan trọng để tính toán các giá trị liên tục của 6 bậc tự do trong suốt chuỗi video. Cách tiếp cận này giúp xử lý các tư thế đầu 3D rộng. Việc này cũng giảm thiểu sự biến đổi về dữ liệu. Xử lý đám mây điểm từ mô hình 3D tạo ra thông tin chi tiết về hình dạng khuôn mặt. Điều này cải thiện độ chính xác của việc theo dõi.
II.Vượt qua Thách thức Nhận dạng Khuôn mặt 3D
Việc theo dõi tư thế đầu 3D rộng gặp nhiều khó khăn đáng kể trong thu thập dữ liệu. Để xây dựng các mô hình thống kê về hình dạng và ngoại hình, thường yêu cầu dữ liệu chính xác về tư thế hoặc số lượng lớn điểm mốc. Dữ liệu tư thế ground-truth rất đắt đỏ để thu thập. Nó đòi hỏi các thiết bị hoặc cảm biến chuyên dụng. Việc chú thích thủ công các điểm tương ứng trên cơ sở dữ liệu lớn có thể tẻ nhạt và dễ gây lỗi. Hơn nữa, vấn đề chú thích điểm tương ứng trở nên khó khăn. Đó là do khó xác định các điểm bị che khuất trên khuôn mặt ở góc nghiêng. Kết quả là, một phương pháp tổng quát giải quyết các xoay ngoài mặt phẳng thường phải sử dụng mô hình dựa trên góc nhìn hoặc mô hình thích ứng. Vấn đề này trở nên quan trọng khi số lượng điểm mốc khác nhau giữa các góc nhìn. Ví dụ, góc nghiêng và chính diện, gây ra khoảng cách về khả năng theo dõi khi sử dụng một mô hình duy nhất. Dữ liệu luồng video cũng đặt ra thách thức tương tự.
2.1. Khó khăn trong thu thập dữ liệu tư thế đầu rộng
Thu thập dữ liệu chính xác cho tư thế đầu 3D rộng là một thách thức lớn. Dữ liệu ground-truth đòi hỏi các thiết bị chuyên biệt, làm tăng chi phí. Việc chú thích thủ công các điểm mốc trên các tập dữ liệu lớn rất tốn thời gian và dễ xảy ra lỗi. Đặc biệt, việc xác định các điểm bị che khuất trên khuôn mặt ở góc nghiêng là một vấn đề. Các góc nhìn khác nhau có số lượng điểm mốc khác nhau. Điều này gây khó khăn cho việc sử dụng một mô hình duy nhất để theo dõi hiệu quả. Vấn đề này ảnh hưởng đến khả năng Nhận dạng khuôn mặt 3D toàn diện.
2.2. Giải pháp dùng dữ liệu tổng hợp quy mô lớn
Dữ liệu tổng hợp quy mô lớn là giải pháp cho vấn đề thu thập dữ liệu. Dữ liệu này bao phủ toàn bộ phạm vi tư thế đầu. Kết hợp với thông tin trực tuyến tương quan giữa các khung hình, phương pháp này tăng cường độ bền vững của quá trình theo dõi. Việc này giúp vượt qua các hạn chế của dữ liệu thực. Nó cải thiện hiệu suất trong các kịch bản phức tạp. Xử lý đám mây điểm từ dữ liệu tổng hợp mang lại mô hình chính xác. Công nghệ quét sâu cũng đóng vai trò quan trọng trong việc tạo ra dữ liệu tổng hợp chất lượng cao.
III.Phương pháp Xử lý Dữ liệu Luồng Video 3D Hiệu quả
Để nghiêm ngặt trong việc theo dõi đồng thời hoạt ảnh khuôn mặt, các xoay ngoài mặt phẳng và hướng tới môi trường thực, một sự kết hợp giữa dữ liệu tổng hợp và dữ liệu thực được đề xuất. Mục tiêu là đào tạo mô hình theo dõi hồi quy tầng. Trong bối cảnh hồi quy tầng, mô tả cục bộ rất quan trọng để đạt hiệu suất cao. Do đó, phương pháp học tính năng được sử dụng. Điều này cho phép biểu diễn các bản vá cục bộ trở nên phân biệt hơn so với các mô tả được tạo thủ công. Hơn nữa, phương pháp này giới thiệu một số sửa đổi đối với cách tiếp cận tầng truyền thống ở các giai đoạn sau. Điều này nhằm nâng cao chất lượng của các điểm mốc hoặc điểm tương ứng được tìm thấy. Đây là bước tiến quan trọng trong Phân tích video 3D.
3.1. Giảm biến đổi ngoại hình bằng tính năng cục bộ
Tính năng cục bộ được áp dụng để giảm sự biến đổi cao của ngoại hình khuôn mặt. Điều này đặc biệt quan trọng khi học từ một tập dữ liệu lớn. Các tính năng này tập trung vào các chi tiết nhỏ trên khuôn mặt. Chúng giúp hệ thống duy trì độ chính xác. Điều này xảy ra ngay cả khi có sự thay đổi về biểu cảm, ánh sáng hoặc góc nhìn. Việc này là chìa khóa để phát hiện đặc điểm khuôn mặt bền vững. Nó cũng góp phần vào sự phát triển của Thị giác máy tính 3D.
3.2. Mô hình theo dõi hồi quy tầng và học tính năng
Sự kết hợp của các tập dữ liệu tổng hợp và thực được đề xuất. Chúng được sử dụng để đào tạo mô hình theo dõi hồi quy tầng. Cách tiếp cận này đảm bảo độ chính xác khi theo dõi đồng thời hoạt ảnh khuôn mặt và các xoay ngoài mặt phẳng. Học tính năng là một yếu tố quan trọng. Nó cho phép biểu diễn các bản vá cục bộ trở nên phân biệt hơn. Các sửa đổi đối với phương pháp tầng truyền thống giúp cải thiện chất lượng của các điểm mốc. Đây là một bước tiến quan trọng trong Lập chỉ mục khuôn mặt và Theo dõi khuôn mặt 3D.
IV.Tối ưu Lập chỉ mục Khuôn mặt Dữ liệu Tổng hợp Thực
Bộ dữ liệu Unconstrained 3D Pose Tracking (U3PT), là các bản ghi riêng, được đề xuất để đánh giá theo dõi khuôn mặt trong môi trường thực. Bộ dữ liệu này thu thập trên mười đối tượng (năm video/đối tượng) trong môi trường văn phòng với nền lộn xộn. Những người được quay video di chuyển thoải mái trước camera. Họ xoay đầu theo ba hướng (Yaw, Pitch và Roll), thậm chí 90 độ của Yaw. Họ cũng thực hiện một số hành động che khuất hoặc biểu cảm. Ngoài ra, các video này được ghi lại trong các điều kiện ánh sáng khác nhau. Với dữ liệu ground-truth có sẵn về tư thế 3D được tính toán bằng hệ thống camera hồng ngoại, độ bền vững và độ chính xác của khung theo dõi có thể được đánh giá chính xác. Khung này nhằm mục đích hoạt động trong môi trường không ràng buộc. Điều này rất quan trọng đối với Dữ liệu luồng video thực tế.
4.1. Xây dựng bộ dữ liệu U3PT cho môi trường thực
Bộ dữ liệu U3PT được tạo ra để đánh giá theo dõi khuôn mặt trong môi trường thực. Bộ dữ liệu này bao gồm các bản ghi từ mười đối tượng. Mỗi đối tượng có năm video, quay trong môi trường văn phòng với nền phức tạp. Các đối tượng được quay di chuyển tự nhiên, xoay đầu theo nhiều hướng (Yaw, Pitch, Roll), bao gồm cả góc Yaw 90 độ. Các tình huống che khuất và biểu cảm khuôn mặt cũng được ghi lại. Các video này được thu dưới nhiều điều kiện ánh sáng khác nhau. Đây là một nguồn tài nguyên quý giá cho Lập chỉ mục khuôn mặt.
4.2. Đánh giá độ bền và chính xác trong điều kiện thực
Dữ liệu ground-truth về tư thế 3D được tính toán bằng hệ thống camera hồng ngoại. Dữ liệu này cho phép đánh giá chính xác độ bền vững và độ chính xác của khung theo dõi. Khung này được thiết kế để hoạt động trong các điều kiện 'in-the-wild' không bị ràng buộc. Việc ước tính tư thế đầu chính xác là rất quan trọng. Điều này giúp đảm bảo hiệu suất mạnh mẽ trong các ứng dụng thực tế của Phân tích video 3D và Theo dõi khuôn mặt 3D.
V.Đánh giá Độ chính xác Thị giác Máy tính 3D Thử nghiệm
Các công bố quan trọng đã được thực hiện trong quá trình nghiên cứu này. Chúng đóng góp vào sự phát triển của Thị giác máy tính 3D và các lĩnh vực liên quan. Các bài báo này thể hiện những tiến bộ trong việc tạo bộ dữ liệu và cải thiện các thuật toán theo dõi khuôn mặt. Chúng cung cấp các phương pháp mới để đối phó với những thách thức phức tạp trong theo dõi khuôn mặt 3D.
5.1. Bộ dữ liệu mới cho theo dõi tư thế 3D không ràng buộc
Công trình "U3PT: A New Dataset for Unconstrained 3D Pose Tracking Evaluation" đã được công bố tại International Conference on Computer Analysis of Images and Patterns (CAIP), 2015. Công trình này giới thiệu bộ dữ liệu U3PT. Đây là một tài nguyên quan trọng cho việc đánh giá Theo dõi khuôn mặt 3D trong các điều kiện không bị ràng buộc. Bộ dữ liệu này giúp cộng đồng nghiên cứu đánh giá và so sánh các thuật toán một cách công bằng. Nó thúc đẩy sự phát triển của các hệ thống Thị giác máy tính 3D mạnh mẽ hơn.
5.2. Cải tiến theo dõi căn chỉnh khuôn mặt bằng hồi quy
Bài báo "Cascaded Regression of Learning Feature for Face alignment" đã được trình bày tại Advanced Concepts for Intelligent Vision Systems (ACIVS), 2015. Nghiên cứu này tập trung vào việc cải thiện căn chỉnh khuôn mặt. Nó sử dụng hồi quy tầng và học tính năng. Việc này giúp các mô tả cục bộ trở nên phân biệt hơn. Phương pháp này nâng cao chất lượng của các điểm mốc được tìm thấy. Điều này cải thiện độ chính xác và độ bền vững của Phát hiện đặc điểm khuôn mặt và Nhận dạng khuôn mặt 3D trong các ứng dụng thực tế.
Tải xuống file đầy đủ để xem toàn bộ nội dung
Tải đầy đủ (131 trang)Trích đoạn nội dung luận án
Tải xuống để đọc toàn bộFace Tracking and Indexing In Video Sequences Reporters: M. Farid MELGANI Examiners: Mme. Bernadette DORIZZI Student: M. Malik MALLEM Ngoc-Trung Tran M.
Kevin BAILLY Supervisors: M. Maurice CHARBIT A thesis submitted in fulfilment of the requirements for the degree of Doctor of Philosophy in the LTCI Department Telecom ParisTECH July 2015 Abstract Face tracking in video sequences is an important problem in computer vision because of multiple applications in many domains, such as: video surveillance, human computer in- terface, biometrics. Although there are many recent advances in this field, it still remains a challenging topic, in particular if 6 Degrees of Freedom (DOF) - three dimensional (3D) translation and rotation - or fiducial points (to capture the facial animation) needs to be estimated in the same time. Its challenge comes mainly from following factors: il- lumination variations, wide head rotation, expression, occlusion, cluttered background, etc.
In this thesis, contributions are made to address two of mentioned major difficul- ties: the expression and wide head rotation. We aim to build a 3D tracking framework on monocular cameras, which is able to estimate accurate 6 DOF and facial animation while being robust with wide head rotation, even profile. Our method adopt the 3D face model as a set of 3D vertices needed to compute continuous values of 6 DOF in video sequences. To track wide 3D head poses, there are significant difficulties regarding to the data collection, where the pose and/or a large number of fiducial points is generally required in order to build the statistical shape and appearance models.
The pose ground-truth is expensive to be collected because of requirement of some specific devices or sensors, while the manual annotation of correspondences on large databases can be tedious and error prone. Moreover, the problem of correspondence annotation is the difficulty of locating the hidden points of self-occlusion faces at the profile. The result is a general approach that wants to tackle the out-of-plane rotations have to usually use a view-based or adaptive models. The problem matters since the different numbers of fiducial points on views, for example, profile and frontal, that causes a gap in term of tracking between them using a single model.
The first main contribution of our thesis is to use of a large synthetic data to overcome this problem. Leveraging on the such data covering the full range of head pose and the on-line information correlating between frames, we propose the combination between them to robustify the tracking. The local features are adopted in our thesis to reduce the high variation of facial appearance when learning on a large dataset. ii To be rigorous with simultaneously tracking of the facial animation, out-of-plane ro- tations and towards to in-the-wild, the combination of synthetic and real datasets is proposed to train using cascaded-regression tracking model.
In the perspective of cas- caded regression, the local descriptor is very significant for high performance. Hence, we utilize the method of feature learning in order to allow the representation of local patches more discriminative than hand-crafted descriptors. Furthermore, the proposed method introduces some modifications of the traditional cascaded approach at later stages in order to boost the quality of the found fiducial points or correspondences. Lastly, Unconstrained 3D Pose Tracking (U3PT) dataset, our own recordings, would be nominated for the evaluation of in-the-wild face tracking in the community.
The dataset captured on ten subjects (five videos/subject) in the office environment with the cluttered background. People, who are captured, moving comfortably in front of the camera, rotating their heads in three direction (Yaw, Pitch and Roll), even 90 degree of Yaw, doing some occlusion or expression. In addition, these videos are recorded in different light conditions. With the available ground-truth of 3D pose computed using the infrared camera system, the robustness and accuracy of tracking framework, which aims to work in-the-wild, could be accurately evaluated.
Keywords: head pose tracking, pose estimation, 3D face tracking, in-the-wild face tracking, Bayesian tracking, face alignment, cascaded regression, synthetic data. iii Publications During this study, the following papers were published or under submission: U3PT: A New Dataset for Unconstrained 3D Pose Tracking Evaluation Ngoc Trung Tran, Fakhreddine Ababsa and Maurice Charbit, International Conference on Computer Analysis of Images and Patterns (CAIP), 2015. Cascaded Regression of Learning Feature for Face alignment Ngoc Trung Tran, Fakhreddine Ababsa, Sarra Ben Fredj and Maurice Charbit, Advanced Concepts for Intelligent Vision Systems (ACIVS), 2015. Towards Pose-Free Tracking of Non-Rigid Face using Synthetic Data Ngoc Trung Tran, Fakhreddine Ababsa and Maurice Charbit, International Conference on Pattern Recognition Applications and Methods (ICPRAM), 2015.
A Robust Framework for Tracking Simultaneously Face Pose and Animation using Synthesized Faces Ngoc Trung Tran, Fakhreddine Ababsa and Maurice Charbit, Pattern Recognition Letter, 2014. 3D Face Pose and Animation Tracking via Eigen-Decomposition based Bayesian Approach Ngoc-Trung Tran, Fakhreddine Ababsa, Jacques Feldmar, Maurice Charbit, Dijana Petrovska- Delacretaz and Gerard Chollet International Symposium on Visual Computing (ISVC), 2013. 3D Face Pose Tracking from Monocular Camera via Sparse Representation of Syn- thesized Faces Ngoc-Trung Tran, Jacques Feldmar, Maurice Charbit, Dijana Petrovska-Delacretaz and Gerard Chollet International Conference on Computer Vision Theory and Applications (VISAPP), 2013. Towards In-the-wild 3D Head Tracking Ngoc Trung Tran, Fakhreddine Ababsa and Maurice Charbit, Journal of Machine Vision and Applications (MVA), 2015.
(Submitted) Abbreviations 2D Two Dimensional 3D Three Dimensional AAM Active Appearance Model ASM Active Shape Model CLM Constrained Local Model DOF Degrees-of-Freedom GMM Gaussian Mixture Model IC Inverse Compositional IRLS Iteratively Re-weighted Least Squares LDM Linear Deformable Model KF Kalman Filter KLT Kanade–Lucas–Tomasi ML Maximum Likelihood MAE Mean Average Error PCA Principal Component Analysis RANSAC RANdom SAmple Consensus RGB Red Green and Blue RMS Root Mean Squared POSIT Pose from Orthography and Scaling with ITerations SDM Supervised Descent Method SfM Structure from Motion SSD Sum of Squared Differences SVD Singular Value Decomposition SVM Support Vector Machine 3DMM 3D Morphable Model Contents Contents v List of Figures viii List of Tables xi 1 Introduction 1 1. 8 2 State-of-the-art 10 2.1 Linear Deformable Models (LDMs) .1 Optical Flows Regularization .3 Multi-view Tracking .1 On-line adaptive models .2 View-based models .1 Off-line learning vs. On-line adaptation .2 Real data vs. 23 v Contents vi 3 A Baseline Framework for 3D Face Tracking 25 3.1 The General Bayesian Tracking Formulation .2 Face Tracking using Covariance Matrices of Synthesized Faces .3 Face Tracking using Feature Reconstruction through Sparse Representation 36 3.1 Reconstructed-Error Function .4 Face Tracking using Adaptive Local Models .1 Adaptive Local Models .3 Out-of-plane tracking.
47 4 Pose-Free Tracking of Non-rigid Face in Controlled Environment 49 4.2 Wide Baseline Matching as Initialization .3 Large-rotation tracking via Pose-wise Appearance Models .1 View-based Appearance Model .2 Matching Strategy by Keyframes .3 Rigid and Non-rigid Estimation using View-based Appearance Model .4 Flexible tracking via Pose-wise Classifiers .1 Templates for Matching using Support Vector Machine .2 Large off-line Synthetic Dataset .3 Local Appearance Models .4 Fitting via Pose-Wise Classifiers .3 Out-of-plane tracking. 69 5 Towards In-the-wild 3D Face Tracking 70 5.1 Cascaded Regression of Learning Features for Landmark Detection .3 Local Learning Feature .4 Coarse-to-Fine Correlative Cascaded Regression .1 Learing Feature on Inner vs Boundary points .2 Coarse-to-Fine Correlative Regression (CFCR) .2 In-the-wild 3D Face Tracking .2 Real and synthetic data for in-the-wild tracking .3 Unconstrained 3D Pose Tracking Dataset .2 Datasets for Visualization .3 Our dataset: Unconstrained Pose Tracking dataset. 99 6 Conclusion and Future Works 100 A Algorithms in Use 103 A.1 Nelder-Mead algorithm .4 Local Binary Patterns. 107 Bibliography 109 List of Figures 1.1 Three head orientations: Yaw, Pitch and Roll.2 Some challenging conditions of 3D face tracking: illumination, head rota- tion, clustered background in one sample video sequence.1 Some 3D face models from left to right: AAM, Cynlinder, Mesh and 3DMM.1 The frontal and profile views of Candide-3 model.2 One example of the perspective projection in our method to obtain 2D projected landmarks from the current model given the parameters.3 The diagram of baseline framework for 3D face tracking: data generation, training and tracking stages.4 Landmark initialization, 3D model fitting using POSIT and the rendering of synthesized images.5 The landmarks used in the framework, examples of mean and covariance matrices at the eye and mouth corners.6 The visualization of two and three first components of descriptor projec- tion using PCA learned through synthetic images of some feature points.7 Some video examples of BUFT videos at different views.8 The framework using reconstructed features via sparse representation with two modifications from the baseline.9 The sparse representation to estimate the coefficients of the positive (green) and negative (red) local patches of one corner of eyebrow.10 The construction of posistive and negative patches of eye corners.11 The framework using adaptive local models based on the baseline.12 An sample video of BUFT dataset using local adaptive model.13 Some tracking examples on BUFT dataset using local adaptive model (green) and FaceTracker (yellow).14 The visualization of our (Yaw, Pitch, Roll) estimation (blue) and ground- truth (red) on the video jam7.avi using local adaptive model.15 The visualization of our (Yaw, Pitch, Roll) estimatation (blue) and ground- truth (red) on the video vam2.avi using local adaptive model.16 The RMS error of 12 selected points for tracking in our framework (red) compared to [97] (blue).
The vertical axis is RMS error (in pixel) and the horizontal axis is the frame number.1 The basic idea of wide baseline matching. 50 viii List of Figures ix 4.2 Computing 3D points Lk as the 3D intersection points.3 From left to right: The 2D SIFT keypoints of the keyframe, SIFT match- ing, outliers removed by RANSAC.4 The pipeline of Pose-Wise Appearance Model based Framework.5 The structure of mapping table with the key as three orientation and the content as the local descriptors.6 The cross-validation to select two parameters: the number of nearest neighbors Nq and the threshold kHub. The validation of a) first row: the number of nearest poses, b) second row: kHub for |Y aw| ≤ 30◦ , and c) third row: kHub for |Y aw| > 30◦. The vertical axis is RMS error.7 Tracking examples on VidTimid and Honda/UCSD.8 The way how to pick up positive (blue) and negative samples (red), and the response map computed at the mouth corner after training.9 Some weight matrices or templates of local patches T (x) in our imple- mentation.10 From left to right of training process: 143 frontal images, landmark anno- tation and 3D model alignment, synthesized images rendering, and pose- wise SVMs training.11 a) The Candide-3 model with facial points in our method.
(b) The way to compute the response map at the mouth corner using three descriptors via SVM templates.12 The pipeline of tracking process from the frame t to t + 1.13 The RMS of our framework (red curve) and FaceTracker [97] (blue curve). The vertical axis is RMS error (in pixel) and the horizontal axis is the frame number.14 Our tracking method on some sample videos of VidTimid and Honda/UCSD.1 Some results of face alignment on some challenging images of 300-W dataset.2 The visualization of weight matrices when learning local descriptors of eye and mouth corners using RBM.3 The overview of our approach. In training, we learn sequentially RBM models, coarse-to-find regression (correlative and global regression).4 CED evaluation of local correlative regression on specific landmarks. i → j means using i-th landmark to detect j-th landmark.5 The cross-validation of feature size, the number of hiddens, the number of random samples and the number of regressors.6 The evaluation of the effect of inner and boundary points on cross-validation set of 300-W dataset.7 Some results on 300-W dataset.8 The annotation of 51 landmarks in synthetic images and its automatic rendering in different views.9 Examples of landmark annotation of Multi-PIE at different views.10 Examples of automatic rendering of synthetic images.11 Landmark detection on some large rotations using synthetic training data.
86 List of Figures x 5.12 Examples of landmark detection on real images using the model trained from the synthetic data.
Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ
Trích dẫn luận án này
Ngoc-Trung Tran (2015). Theo dõi và lập chỉ mục khuôn mặt trong chuỗi video 3D [Luận án tiến sĩ, Telecom ParisTech]. LuanAn.net. https://luanan.net/tai-lieu-khac/theo-doi-lap-chi-muc-khuon-mat-chuoi-video-3d
Câu hỏi thường gặp
Luận án "Theo dõi và lập chỉ mục khuôn mặt trong chuỗi video 3D" nghiên cứu về vấn đề gì?
Phân tích và chỉ mục khuôn mặt trong chuỗi video giúp xác định và theo dõi chuyển động khuôn mặt trong video.
Luận án "Theo dõi và lập chỉ mục khuôn mặt trong chuỗi video 3D" được bảo vệ tại trường nào?
Luận án này được bảo vệ tại Telecom ParisTech. Năm bảo vệ: 2015.
Luận án "Theo dõi và lập chỉ mục khuôn mặt trong chuỗi video 3D" thuộc chuyên ngành gì?
Luận án "Theo dõi và lập chỉ mục khuôn mặt trong chuỗi video 3D" thuộc chuyên ngành Computer Vision. Danh mục: Tài liệu khác.
Luận án "Theo dõi và lập chỉ mục khuôn mặt trong chuỗi video 3D" có bao nhiêu trang?
Luận án "Theo dõi và lập chỉ mục khuôn mặt trong chuỗi video 3D" có 131 trang. Bạn có thể xem trước một phần tài liệu ngay trên trang web trước khi tải về.
Cách tải luận án "Theo dõi và lập chỉ mục khuôn mặt trong chuỗi video 3D" về máy như thế nào?
Để tải luận án về máy, bạn nhấn nút "Tải xuống ngay" trên trang này, sau đó hoàn tất thanh toán phí lưu trữ. File sẽ được tải xuống ngay sau khi thanh toán thành công. Hỗ trợ qua Zalo: 0559 297 239.