Optimization algorithms for inference and classification of genet

Thuật toán tối ưu hóa cho suy luận và phân loại dữ liệu. Khám phá các phương pháp tiên tiến, hiệu quả cao.

Trường ĐH

Rowan University

Chuyên ngành

Electrical & Computer Engineering

Tác giả

Luan An

Thể loại

Luận văn thạc sĩ

Năm xuất bản

Số trang

99

Thời gian đọc

15 phút

Lượt xem

0

Lượt tải

0

Phí lưu trữ

40 Point

Tổng quan nhanh

Chủ đề:
Các thuật toán tối ưu hóa cho suy luận hồ sơ di truyền
Số trang:
99 trang
Trường:
Rowan University
Chuyên ngành:
Electrical & Computer Engineering
Tác giả:
Năm:

Tóm tắt nội dung luận án

I.Các thuật toán tối ưu hóa cho suy luận hồ sơ di truyền

Tài liệu này khám phá các thuật toán tối ưu hóa tiên tiến nhằm nâng cao khả năng suy luận và phân loại hồ sơ di truyền từ các phép đo thiếu mẫu. Các phương pháp này giải quyết các thách thức phức tạp trong phân tích dữ liệu sinh học, đặc biệt là khi dữ liệu bị hạn chế hoặc không đầy đủ. Nghiên cứu tập trung vào việc phát triển và cải tiến các kỹ thuật học máy để trích xuất thông tin có giá trị từ dữ liệu gen, tối ưu hóa các tham số mô hình. Mục tiêu chính là cải thiện độ chính xác của việc phân loại và độ ổn định của các mô hình trong môi trường dữ liệu thực tế. Các thuật toán tối ưu hóa được thảo luận bao gồm các phương pháp mở rộng của phân tích nhân tố ma trận không âm và các kỹ thuật hồi quy đa biến. Đây là những công cụ thiết yếu cho việc phân tích dữ liệu gen lớn, giúp giải quyết các bài toán về hàm mất mát và tối ưu hóa tham số một cách hiệu quả.

1.1. Mục tiêu và phạm vi nghiên cứu chính

Nghiên cứu này giải quyết ba vấn đề tối ưu hóa riêng biệt liên quan đến suy luận và phân loại hồ sơ di truyền. Các vấn đề bao gồm việc mở rộng các khung phân tích nhân tố ma trận không âm, phát triển phương pháp hồi quy đa biến cho dữ liệu kích thước cao nhưng mẫu nhỏ, và một thuật toán tham lam mới cho tối ưu hóa sparse. Mục tiêu là phát triển các kỹ thuật mới vượt trội so với các phương pháp hiện có về độ chính xác và hiệu quả tính toán. Các thuật toán tối ưu hóa được thiết kế để xử lý các đặc điểm dữ liệu gen cụ thể, bao gồm tính không âm và mối tương quan giữa các biến phản hồi.

1.2. Thách thức phân tích dữ liệu gen thiếu mẫu

Việc phân tích hồ sơ di truyền từ các phép đo thiếu mẫu đặt ra nhiều thách thức đáng kể. Dữ liệu gen thường có kích thước cao nhưng số lượng mẫu nhỏ, dẫn đến các vấn đề như nhiễu, tương quan biến cao và khó khăn trong việc ước tính đáng tin cậy. Các phương pháp tối ưu hóa truyền thống thường không hiệu quả trong những trường hợp này. Cần có các kỹ thuật mới để đối phó với sự phân kỳ của hàm mất mát và đảm bảo sự hội tụ của các mô hình. Các thuật toán tối ưu hóa phải có khả năng xử lý tính phức tạp của dữ liệu sinh học, cung cấp các giải pháp mạnh mẽ cho suy luận và phân loại.

II.PNMF Suy luận phân loại dữ liệu DNA Microarray

Phân tích nhân tố ma trận không âm xác suất (PNMF) là một mở rộng của khung phân tích nhân tố ma trận không âm (NMF) truyền thống, được phát triển để cải thiện việc phân cụm và phân loại dữ liệu DNA microarray. NMF là một thuật toán tối ưu hóa phổ biến trong học máy, nhưng phiên bản xác suất cung cấp sự ổn định và độ chính xác cao hơn. PNMF giải quyết các vấn đề liên quan đến tính ngẫu nhiên và biến động trong dữ liệu gen, cung cấp một khung làm việc mạnh mẽ hơn. Việc áp dụng PNMF cho dữ liệu microarray DNA đã chứng minh khả năng vượt trội so với NMF xác định và NMF sparse trong cả độ ổn định phân cụm và độ chính xác phân loại. Điều này làm cho PNMF trở thành một công cụ có giá trị để giải mã các mô hình gen phức tạp và phát hiện các dấu hiệu sinh học quan trọng. PNMF đại diện cho một bước tiến quan trọng trong việc ứng dụng các thuật toán tối ưu hóa vào sinh học tính toán.

2.1. Mở rộng từ NMF xác định sang PNMF xác suất

Nghiên cứu mở rộng khung NMF xác định sang trường hợp xác suất để xử lý tốt hơn sự không chắc chắn trong dữ liệu sinh học. PNMF được xây dựng trên nền tảng của NMF, nhưng tích hợp các yếu tố xác suất để mô hình hóa dữ liệu một cách linh hoạt hơn. Quá trình tối ưu hóa tham số trong PNMF liên quan đến việc tối thiểu hóa một hàm mất mát phù hợp, thường là Kullback-Leibler divergence. Việc này đảm bảo rằng mô hình có thể nắm bắt được các cấu trúc tiềm ẩn trong dữ liệu một cách hiệu quả hơn, dẫn đến kết quả phân cụm và phân loại đáng tin cậy hơn.

2.2. Ứng dụng PNMF trong phân tích Microarray gen

PNMF được áp dụng cụ thể để phân cụm và phân loại dữ liệu DNA microarray. Dữ liệu microarray thường có kích thước cao, và việc trích xuất các mẫu có ý nghĩa là rất quan trọng. PNMF giúp giảm chiều dữ liệu trong khi vẫn giữ được thông tin quan trọng, cho phép phát hiện các cụm gen có liên quan đến các trạng thái bệnh hoặc chức năng sinh học cụ thể. Các thuật toán tối ưu hóa được sử dụng để điều chỉnh các yếu tố trong ma trận, đảm bảo rằng việc phân loại gen đạt độ chính xác cao. PNMF đã chứng minh hiệu quả trong việc xử lý các tập dữ liệu gen thực tế.

2.3. Cải thiện độ ổn định và chính xác với PNMF

Kết quả thực nghiệm cho thấy PNMF vượt trội hơn hẳn NMF xác định và NMF sparse về độ ổn định phân cụm và độ chính xác phân loại. Độ ổn định đề cập đến khả năng của thuật toán tạo ra các cụm nhất quán trên các lần chạy khác nhau hoặc các tập dữ liệu tương tự. Độ chính xác phân loại được đo bằng khả năng của mô hình gán nhãn chính xác cho các hồ sơ gen. Sự cải tiến này là do khả năng của PNMF trong việc mô hình hóa tốt hơn sự biến động của dữ liệu và tối ưu hóa tham số một cách hiệu quả hơn, giảm thiểu hàm mất mát và tránh các điểm cực tiểu cục bộ.

III.SMURC Hồi quy đa biến và ước tính hiệp phương sai

Nghiên cứu đề xuất SMURC (Small-sample MUltivariate Regression with Covariance estimation) để giải quyết vấn đề hồi quy đa biến có kích thước cao nhưng số lượng mẫu thấp. Trong các trường hợp này, phương pháp khả năng hợp lý tối đa truyền thống thường không hiệu quả do hàm khả năng hợp lý phân kỳ. SMURC giới thiệu một chuẩn hóa cho hàm khả năng hợp lý, đảm bảo sự hội tụ và cung cấp các ước lượng đáng tin cậy. Vấn đề này thường gặp trong học máy khi phân tích các tập dữ liệu phức tạp như mạng lưới điều hòa gen. Các thuật toán tối ưu hóa được sử dụng để ước tính hiệp phương sai một cách hiệu quả, ngay cả khi số lượng biến vượt xa số lượng quan sát. SMURC cải thiện đáng kể khả năng phân tích và suy luận từ dữ liệu gen thách thức, nơi các phương pháp thông thường thất bại.

3.1. Thách thức trong hồi quy kích thước cao mẫu nhỏ

Hồi quy đa biến với số lượng biến lớn và số lượng mẫu nhỏ đặt ra một thách thức lớn. Trong bối cảnh này, ma trận hiệp phương sai mẫu trở nên suy biến, và ước lượng khả năng hợp lý tối đa có thể không tồn tại hoặc phân kỳ. Điều này làm cho việc áp dụng trực tiếp các phương pháp thống kê truyền thống trở nên không khả thi. Cần có các thuật toán tối ưu hóa mới để xử lý các vấn đề này một cách hiệu quả. Việc ước tính chính xác ma trận hiệp phương sai là chìa khóa để xây dựng các mô hình hồi quy đáng tin cậy và thực hiện suy luận chính xác.

3.2. Phương pháp Normalized Likelihood của SMURC

SMURC giới thiệu một cách tiếp cận Normalized-Likelihood để giải quyết vấn đề phân kỳ của hàm khả năng hợp lý. Bằng cách chuẩn hóa hàm khả năng hợp lý, thuật toán đảm bảo sự hội tụ của các ước lượng, ngay cả trong điều kiện dữ liệu khó khăn. Phương pháp này liên quan đến việc điều chỉnh hàm mất mát để tránh các ước lượng không ổn định. Các thuật toán tối ưu hóa được áp dụng để tìm các tham số mô hình tối ưu dưới chuẩn hóa này, cho phép ước tính ma trận hiệp phương sai một cách mạnh mẽ. Đây là một đóng góp quan trọng cho lĩnh vực học máy và thống kê.

3.3. Ứng dụng SMURC trong mạng lưới điều hòa gen

SMURC được áp dụng để phân tích các mạng lưới điều hòa gen, một lĩnh vực quan trọng trong sinh học tính toán. Các mạng lưới này thường liên quan đến một số lượng lớn gen và các tương tác phức tạp, nhưng dữ liệu thực nghiệm thường hạn chế. SMURC cung cấp một phương pháp đáng tin cậy để suy luận các mối quan hệ trong mạng lưới, ngay cả khi dữ liệu thiếu mẫu. Kết quả mô phỏng cho thấy SMURC vượt trội hơn các ước lượng khả năng hợp lý được điều chuẩn với ma trận hiệp phương sai đã biết và mô hình đồ họa Gaussian có điều kiện sparse (sCGGM) hiện đại. Điều này nhấn mạnh hiệu quả của SMURC trong việc giải quyết các vấn đề phức tạp trong học sâu và sinh học hệ thống.

IV.Tối ưu hóa sparse Thuật toán Tái cấu trúc Kernel mới

Nghiên cứu này trình bày một thuật toán tham lam mới, được gọi là Tái cấu trúc Kernel, cung cấp một giải pháp sparse chính xác cho bài toán tối ưu hóa l0. Các bài toán tối ưu hóa l0 rất quan trọng trong nhiều lĩnh vực, bao gồm học máy, xử lý tín hiệu và nén dữ liệu, nơi mục tiêu là tìm kiếm một giải pháp với số lượng phần tử khác 0 tối thiểu. Tuy nhiên, việc giải quyết các bài toán này thường rất phức tạp về mặt tính toán. Không giống như các phương pháp tham lam khác chỉ cung cấp các xấp xỉ, thuật toán Tái cấu trúc Kernel được chứng minh là tìm ra giải pháp tối ưu chính xác với thời gian tính toán ít hơn đáng kể. Đây là một bước tiến quan trọng trong việc phát triển các thuật toán tối ưu hóa hiệu quả cho các bài toán sparse.

4.1. Giải pháp tối ưu l0 và hạn chế hiện có

Bài toán tối ưu hóa l0 tìm kiếm vector sparse nhất thỏa mãn một ràng buộc nhất định. Giải pháp l0 lý tưởng là tìm ra tập hợp nhỏ nhất các thành phần để tái tạo một tín hiệu hoặc mô hình. Tuy nhiên, đây là một bài toán kết hợp khó, thường đòi hỏi thời gian tính toán theo cấp số nhân để tìm giải pháp chính xác. Các phương pháp tham lam hiện có như OMP (Orthogonal Matching Pursuit) hoặc CoSaMP (Compressive Sampling Matching Pursuit) thường chỉ cung cấp các giải pháp xấp xỉ. Các thuật toán tối ưu hóa này gặp khó khăn trong việc đảm bảo tính chính xác toàn cục khi xử lý hàm mất mát không lồi.

4.2. Cơ chế thuật toán Tái cấu trúc Kernel

Thuật toán Tái cấu trúc Kernel là một cách tiếp cận tham lam mới được thiết kế để tìm ra giải pháp tối ưu l0 chính xác. Cơ chế của nó khác với các phương pháp tham lam truyền thống ở chỗ nó sử dụng một chiến lược lựa chọn phần tử thông minh hơn, dựa trên việc tái cấu trúc các nhân. Thuật toán này giảm đáng kể thời gian tính toán, từ cấp số nhân xuống một mức hiệu quả hơn. Khả năng tìm ra giải pháp chính xác mà không phải hy sinh tốc độ là một ưu điểm lớn, đặc biệt đối với các ứng dụng yêu cầu độ chính xác cao trong học máy và xử lý dữ liệu lớn. Điều này tối ưu hóa tham số một cách hiệu quả.

Mục lục chi tiết luận án

List of Figures
List of Tables
1. Introduction
1.3. Organization
2. PNMF: Theory & Application To Microarray Data Analysis
2.2. Non-negative Matrix Factorization
2.3. Probabilistic Non-negative Matrix Factorization
2.4. PNMF-based Data Classification
2.5. Application to Gene Microarrays
2.6. Conclusion and Discussion
3. High-Dimension SMURC Estimation
3.2. The Normalized-Likelihood
3.3. Application: Genetic Regulatory Networks
3.4. Conclusion and Discussion
4. Kernel Reconstruction V.4 Conclusion
Bibliography
Appendix
Xem trước tài liệu
Tải đầy đủ để xem toàn bộ nội dung
Optimization algorithms for inference and classification of genet

Tải xuống file đầy đủ để xem toàn bộ nội dung

Tải đầy đủ (99 trang)

Trích đoạn nội dung luận án

Tải xuống để đọc toàn bộ

Rowan University Rowan Digital Works Theses and Dissertations 9-2-2014 Optimization algorithms for inference and classification of genetic profiles from undersampled measurements Belhassen Bayar Follow this and additional works at: https://rdw.edu/etd Part of the Electrical and Computer Engineering Commons Recommended Citation Bayar, Belhassen, "Optimization algorithms for inference and classification of genetic profiles from undersampled measurements" (2014). Theses and Dissertations.edu/etd/410 This Thesis is brought to you for free and open access by Rowan Digital Works. It has been accepted for inclusion in Theses and Dissertations by an authorized administrator of Rowan Digital Works. For more information, please contact graduateresearch@rowan.

OPTIMIZATION ALGORITHMS FOR INFERENCE AND CLASSIFICATION OF GENETIC PROFILES FROM UNDERSAMPLED MEASUREMENTS by Belhassen Bayar A Thesis Submitted to the Department of Electrical & Computer Engineering College of Engineering In partial fulfillment of the requirement For the degree of Master of Science at Rowan University June 2014 Thesis Chair: Nidhal Bouaynaya © 2014 Belhassen Bayar ACKNOWLEDGEMENTS I want to express my sincere gratitude to Dr. Nidhal Bouaynaya, my supervisor who has always bothered to offer me the best working conditions possible. I thank her for her wide availability, her high scientific qualifications and her guidance, illuminating discussions related to this work and beyond, encouragement, moral and financial support in this research. I express my appreciation and gratitude to Dr.

Roman Shterenberg , Associate Professor at the University of Alabama at Birmingham USA, for the time he spent with me, his availability even when he was abroad and the valuable advice he has given me throughout my research. I also would like to express my deep and sincere gratitude to Dr. Robi Polikar, Professor & Chair at the ECE Department, for the high quality courses he teaches, his availability and eagerness to provide the best learning experience for students at the department. Many thanks to all the students who accompanied me during these years and have continued to create a good working atmosphere within the laboratory.

Deepest thanks to my dear parents and grandmother to whom I owe so much. I would have neither the means nor the strength to accomplish this work without them. I also want to express my gratitude to my friends who have continued to give me the moral and intellectual support throughout my work during all the good and bad moments. They always say the best is for the end, that’s why I dedicate this project to my dear sister, my little light that gave me energy and courage.

iii Abstract Belhassen Bayar OPTIMIZATION ALGORITHMS FOR INFERENCE AND CLASSIFICATION OF GENETIC PROFILES FROM UNDERSAMPLED MEASUREMENTS 2014/06 Nidhal Bouaynaya, Ph. Master of Science in Electrical & Computer Engineering In this thesis, we tackle three different problems, all related to optimization tech- niques for inference and classification of genetic profiles. First, we extend the de- terministic Non-negative Matrix Factorization (NMF) framework to the probabilistic case (PNMF). We apply the PNMF algorithm to cluster and classify DNA microar- rays data.

The proposed PNMF is shown to outperform the deterministic NMF and the sparse NMF algorithms in clustering stability and classification accuracy. Sec- ond, we propose SMURC: Small-sample MUltivariate Regression with Covariance estimation. Specifically, we consider a high dimension low sample-size multivariate regression problem that accounts for correlation of the response variables. We show that, in this case, the maximum likelihood approach is senseless because the likeli- hood diverges.

We propose a normalization of the likelihood function that guaran- tees convergence. Simulation results show that SMURC outperforms the regularized likelihood estimator with known covariance matrix and the state-of-the-art sparse Conditional Graphical Gaussian Model (sCGGM). In the third Chapter, we derive a new greedy algorithm that provides an exact sparse solution of the combinatorial `0 - optimization problem in an exponentially less computation time. Unlike other greedy approaches, which are only approximations of the exact sparse solution, the proposed greedy approach, called Kernel reconstruction, leads to the exact optimal solution.

iv Table of Contents List of Figures vii List of Tables viii 1 Introduction 1 1.3 Organization 2 2 PNMF: Theory & Application To Microarray Data Analysis 4 2.2 Non-negative Matrix Factorization 10 2.3 Probabilistic Non-negative Matrix Factorization 13 2.4 PNMF-based Data Classification 15 2.5 Application to Gene Microarrays 19 2.6 Conclusion and Discussion 34 3 High-Dimension SMURC Estimation 36 3.2 The Normalized-Likelihood 41 3.3 Application: Genetic Regulatory Networks 53 3.4 Conclusion and Discussion 60 v 4 Kernel Reconstruction V.4 Conclusion 79 Bibliography 80 A Appendix 86 vi List of Figures 2.1 Clustering results for the Leukemia dataset 20 2.2 Metagenes expression patterns versus the samples for k = 4 21 2.3 Clustering results for the Medulloblastoma dataset 22 2.4 Clustering Percentage Error versus Nbr.5 The cophenetic coefficient versus the standard deviation 29 2.6 Cophenetic versus SNR in dB in Leukemia dataset 30 2.7 Cophenetic versus SNR in dB in Medulloblastoma dataset 31 3.1 Approximation of the optimization problem in Proposition 4 49 3.2 Approximation error ||S − S ∗ ||F /||S||F versus n 51 3.3 Performance comparison of SMURC with sCGGM and RMLE 53 3.4 The known undirected gene interactions in the Drosophila 57 3.5 Estimated gene regulatory networks of the Drosophila 57 4.1 Performance comparison of KR with `1 -based and `2 -based CS for N = 10 78 4.2 Performance comparison of KR with `1 -based and `2 -based CS for N = 20 78 vii List of Tables 2.1 Smallest SNR value for ρ ≥ 0.1 Detection of the known gene interactions in Flybase 60 viii Chapter 1 Introduction 1.1 Research Objectives We outline the goal of this research through the following objectives: 1. Study and analyse the Non-negative Matrix Factorization (NMF) and propose a probabilistic extension to NMF (PNMF) for data corrupted by noise. Build a PNMF-based classifier and apply it for tumor classification from gene expression data. Derive a convex optimization algorithm for the solution of an under-determined multivariate regression problem.

Apply the proposed algorithm to infer genetic regulatory networks from gene expression data. Derive a greedy algorithm for exact reconstruction of sparse signals from a limited number of observations.2 Research Contribution This work contributes to the field of computational bioinformatics and biology through the application of the signal processing algorithms aiming to study and analyze the microarray data. Our work shifts the focus of the genomic signal processing commu- nity from analyzing the genes expression patterns and samples clusters to considering 1 the mathematical aspect of the algorithm and deriving its application in the stochas- tic work. We also focus on solving under-determined multivariate regression systems in order to infer gene regulatory networks.

These networks are known to be sparse, therefore, we have a great interest in studying the compressive sensing approach which recovers sparse signal from linear model. Specific contributions of this work include: ˆ The improvement of the mathematical proof for the NMF algorithm by provid- ing a general evidence (see Appendix preposition 2). ˆ The development of a new NMF algorithm for the noisy Microarray data in order to improve the basic NMF approach and to predict some hidden data features. ˆ Solving under-determined multivariate regression systems to infer gene regula- tory networks using our new SMURC algorithm.

ˆ Recover k-sparse signal using our new approach, called Kernel Reconstruction, that guarantees an exact reconstruction and less computational time comparing to the `0 -based compressive sensing approach [18].3 Organization This thesis is organized as follows. In Chapter 2, we study and analyze the Non-negative Matrix Factorization and de- rive its probabilistic approach that we call PNMF algorithm and then we derive its corresponding update rules. The proof of the developed approaches is provided in 2 the Appendix chapter. We compare the performance of our PNMF approach with its homologues in clustering as well as classification.

In Chapter 3, we develop a new approach, called Small-sample MUltivariate Re- gression with Covariance Estimation (SMURC), to solve under-determined multivari- ate regression systems. We use this approach to infer gene regulatory networks. We compare our algorithm to other techniques cited in related works and using a syn- thetic data. Subsequently, we apply our approach to infer the know interactions in the Drosophila’s 11-gene wing muscle network.

Finally, in Chapter 4 we provide a complete review of the compressive sensing technique. We also come up with a new approach that performs an exact reconstruc- tion of a sparse signal. We call this approach, Kernel Reconstruction, and we compare it with what has been suggested in the related work. 3 Chapter 2 Probabilistic Non-negative Matrix Factorization: Theory and Application to Microarray Data Analysis 2.1 Introduction Extracting knowledge from experimental raw data and measurements is an important objective and challenge in signal processing.

Often data collected is high dimensional and incorporates several inter-related variables, which are combinations of underly- ing latent components or factors. Approximate low-rank matrix factorizations play a fundamental role in extracting these latent components [14]. In many applica- tions, signals to be analyzed are non-negative, e., pixel values in image processing, price variables in economics and gene expression levels in computational biology. For such data, it is imperative to take the non-negativity constraint into account in or- der to obtain a meaningful physical interpretation.

Classical decomposition tools, such as Principal Component Analysis (PCA), Singular Value Decomposition (SVD), Blind Source Separation (BSS) and related methods do not guarantee to maintain the non-negativity constraint. Non-negative matrix factorization (NMF) represents non-negative data in terms of lower-rank non-negative factors. NMF proved to be a powerful tool in many applications in biomedical data processing and analysis, 4 such as muscle identification in the nervous system [54], classification of images [29], gene expression classification [10], biological process identification [32] and transcrip- tional regulatory network inference [38]. The appeal of NMF, compared to other clustering and classification methods, stems from the fact that it does not impose any prior structure or knowledge on the data.

Brunet et al. successfully applied NMF to the classification of gene expression datasets [10] and showed that it leads to more accurate and more robust clustering than the Self-Organizing Maps (SOMs) and Hierarchical Clustering (HC). Analytically, the NMF method factors the original non-negative matrix V into two lower rank non-negative matrices, W and H such that V = W H + E, where E is the residual error. Lee and Seung [33] derived algorithms for estimating the optimal non-negative factors that minimize the Euclidean distance and the Kullback-Leibler divergence cost functions.

Their algorithms, guaranteed to converge, are based on multiplicative update rules, and are a good compromise be- tween speed and ease of implementation. In particular, the Euclidean distance NMF algorithm can be shown to reduce to the gradient descent algorithm for a specific choice of the step size [33]. Lee and Seung’s NMF factorization algorithms have been widely adopted by the community [6, 10, 19, 59]. The NMF method is, however, deterministic.

That is, the algorithm does not take into account the measurement or observation noise in the data. On the other hand, data collected using electronic or biomedical devices, such as gene expression profiles, are known to be inherently noisy and therefore, must be processed and analyzed by systems that take into account the stochastic nature of the data. Furthermore, the ef- fect of the data noise on the NMF method in terms of convergence and robustness has 5 not been previously investigated. Thus, questions about the efficiency and robustness of the method in dealing with imperfect or noisy data are still unanswered.

In this chapter, we extend the NMF framework and algorithms to the stochastic case, where the data is assumed to be drawn from a multinomial probability den- sity function. We call the new framework Probabilistic NMF or PNMF. We show that the PNMF formulation reduces to a weighted regularized matrix factorization problem. We generalize and extend Lee and Seung’s algorithm to the stochastic case; thus providing PNMF updates rules, which are guaranteed to converge to the optimal solution.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Trích dẫn luận án này

Belhassen Bayar (2014). Optimization algorithms for inference and classification of [Luận án tiến sĩ, Rowan University]. LuanAn.net. https://luanan.net/cong-nghe-thong-tin/tri-tue-nhan-tao/optimization-algorithms-for-inference-and-classification-of-genet

Câu hỏi thường gặp

Luận án "Optimization algorithms for inference and classification of" nghiên cứu về vấn đề gì?

Thuật toán tối ưu hóa cho suy luận và phân loại dữ liệu. Khám phá các phương pháp tiên tiến, hiệu quả cao.

Luận án "Optimization algorithms for inference and classification of" được bảo vệ tại trường nào?

Luận án này được bảo vệ tại Rowan University. Năm bảo vệ: 2014.

Luận án "Optimization algorithms for inference and classification of" thuộc chuyên ngành gì?

Luận án "Optimization algorithms for inference and classification of" thuộc chuyên ngành Electrical & Computer Engineering. Danh mục: Trí Tuệ Nhân Tạo.

Luận án "Optimization algorithms for inference and classification of" có bao nhiêu trang?

Luận án "Optimization algorithms for inference and classification of" có 99 trang. Bạn có thể xem trước một phần tài liệu ngay trên trang web trước khi tải về.

Cách tải luận án "Optimization algorithms for inference and classification of" về máy như thế nào?

Để tải luận án về máy, bạn nhấn nút "Tải xuống ngay" trên trang này, sau đó hoàn tất thanh toán phí lưu trữ. File sẽ được tải xuống ngay sau khi thanh toán thành công. Hỗ trợ qua Zalo: 0559 297 239.

Luận án liên quan

Chia sẻ tài liệu: Facebook Twitter