-
About the Course 0
LMT - SEBI GRADE A 2025 - IT This Course is Only Available
On Our AppDive in and start learning. Get offline access to all the course contents!
No items in this section -
Big Data Analytics [Videos] 36
-
M1:- Introduction to Big Data Analytics 07 minLecture2.1
-
M1:-Introduction to Hadoop Part #1 10 minLecture2.2
-
M1:-Introduction to Hadoop Part #2 10 minLecture2.3
-
M2:-Introduction to MapReduce 11 minLecture2.4
-
M2:-Map Reduce Word Count Problem PYQ NumericalsLecture2.5
-
M2:-Matrix Multiplication 16 minLecture2.6
-
M3:-Introduction to No SQL Database 08 minLecture2.7
-
M3:-Key-Value Stores 07 minLecture2.8
-
M3:-Column Store Database 06 minLecture2.9
-
M3:-Document Database 05 minLecture2.10
-
M3:-Graph Database 07 minLecture2.11
-
M4:-Data Stream Management System 08 minLecture2.12
-
M4:-Sampling Techniques – Part 1 07 minLecture2.13
-
M4:-Sampling Techniques – Part 2 07 minLecture2.14
-
M4:-Bloom Filtering 18 minLecture2.15
-
M4:-Bloom Filter Numerical in BDA 23 minLecture2.16
-
M4:-FM ( Flajolet Martin ) Algorithm PYQ Numerical//Lecture2.17
-
M4:-Flajolet Martin Algorithm 14 minLecture2.18
-
M4:-DGIM algorithm (Datar-Gionis-Indyk-Motwani Algorithm) 09 minLecture2.19
-
M5:-Distance Measure 09 minLecture2.20
-
M5:-Euclidean Distance 08 minLecture2.21
-
M5:-Jaccard Distance 06 minLecture2.22
-
M5:-Cosine Distance 07 minLecture2.23
-
M5:-Edit Distance 10 minLecture2.24
-
M5:-Hamming Distance 05 minLecture2.25
-
M5:-Cure Algorithm 11 minLecture2.26
-
M6:-Collaborative Filtering 20 minLecture2.27
-
M6:-Dead Ends 04 minLecture2.28
-
M6:-Clique and Community 14 minLecture2.29
-
M6:-Clique and Community Numerical 20 minLecture2.30
-
M6:-Content Based Recommendation System 18 minLecture2.31
-
M6:-Authority and hub 21 minLecture2.32
-
M6:-Determine the Communities Girvan Newman Sum VIMP 23 minLecture2.33
-
M6:-Determine Communities Girvan Newman Numerical 2 VVIMP 17 minLecture2.34
-
M6:Determine the Communities Girvan Newman PYQ #1//Lecture2.35
-
M6:Determine the Communities Girvan Newman PYQ #2//Lecture2.36
-
-
Big Data Notes and Importance 12
-
Module 1 – Introduction to Big Data and HadoopLecture3.1
-
Module 2 – Hadoop HDFS and Map ReduceLecture3.2
-
Module 3 – NoSqlLecture3.3
-
Module 4 – Mining Data StreamsLecture3.4
-
Module 5 – Real Time Big Data ModelLecture3.5
-
Module 6 – Data Analytics with RLecture3.6
-
Module 1 – Big Data analytics ImportanceLecture3.7
-
Module 2 – Hadoop HDFS and MapReduceLecture3.8
-
Module 3 – NO SQLLecture3.9
-
Module 4 – Mining Data StreamLecture3.10
-
Module 5 – Real Time Big Data ModelsLecture3.11
-
Module 6 – R ProgrammingLecture3.12
-
-
Big Data Analytics [viva] 6
-
Introduction to Big Data and HadoopLecture4.1
-
Hadoop HDFS and Map ReduceLecture4.2
-
NoSQLLecture4.3
-
Mining Data StreamsLecture4.4
-
Finding Similar Items and ClusteringLecture4.5
-
Real-Time Big Data ModelsLecture4.6
-
-
Natural Language Processing [Videos] 38
-
Introduction to NLP [Natural Language Processing] 12 minLecture5.1
-
Knowledge Required in NLP 11 minLecture5.2
-
Ambiguity in NLP 07 minLecture5.3
-
NLP Phases 08 minLecture5.4
-
Regular Expression 09 minLecture5.5
-
FSA 09 minLecture5.6
-
Language Model 10 minLecture5.7
-
Morphology Analysis 11 minLecture5.8
-
N-gram Model 04 minLecture5.9
-
Morphology Parsing 09 minLecture5.10
-
Design FSA for Word of English 1 – 99 05 minLecture5.11
-
NLP Bigram Numericals PYQ 29 minLecture5.12
-
POS Tagging 10 minLecture5.13
-
Syntax Analysis 03 minLecture5.14
-
Tag-set for English 12 minLecture5.15
-
Stochastic Part of Speech Tagging 08 minLecture5.16
-
Transformation Based Tagging 06 minLecture5.17
-
Multiple Tags ,Word and Unknown Words 04 minLecture5.18
-
Basic Concept of Grammar and Parse Tree 09 minLecture5.19
-
Parsing in NLP 06 minLecture5.20
-
Hidden Markov Model Part 1 10 minLecture5.21
-
Hidden Markov Model Part 2 07 minLecture5.22
-
Viterbi Algorithm 08 minLecture5.23
-
HMM Numerical-01 PYQ 36 minLecture5.24
-
HMM Numerical 2 – PYQ 29 minLecture5.25
-
Conditional Random Field CRF 20 minLecture5.26
-
Introduction to Semantic Analysis 13 minLecture5.27
-
Element of Semantic Analysis 06 minLecture5.28
-
Attachment for Fragment of English (Phrases #1) 09 minLecture5.29
-
Attachment for Fragment of English (Phrases #2) 05 minLecture5.30
-
Attachment for Fragment of English (Phrases #3) 04 minLecture5.31
-
WordNet 08 minLecture5.32
-
Word Sense Disambiguation (WSD) 09 minLecture5.33
-
Yarowsky Approach and HyperLex Approach 15 minLecture5.34
-
Machine Translation Introduction 14 minLecture5.35
-
Machine Translation Types 23 minLecture5.36
-
Question Answering System 20 minLecture5.37
-
Information Retrieval & the different steps in text processing for Information Retrieval 25 minLecture5.38
-
-
Natural Language Processing [Notes] 12
-
Introduction to NLP[Notes]Lecture6.1
-
Word Level Analysis [Notes]Lecture6.2
-
Syntax Analysis [Notes]Lecture6.3
-
Semantic Analysis [Notes]Lecture6.4
-
Pragmatics [Notes]Lecture6.5
-
Module 1 Introduction to NLPLecture6.6
-
Module 2 WORD LEVEL ANALYSISLecture6.7
-
Module 3 SYNTAX ANALYSISLecture6.8
-
Module 4 SEMANTIC ANALYSISLecture6.9
-
Module 5 Pragmatic & Discourse ProcessingLecture6.10
-
Module 6 Applications of NLPLecture6.11
-
Viva QuestionsLecture6.12
-
-
Deep Learning [Coming Soon] 0
No items in this section
Introduction to Big Data and Hadoop
1.What is the definition of Software Engineering?
Ans:Big Data in general is defined as high volume, velocity and variety information assets that demand cost-effective, innovative forms of information processing for enhanced insight and decision making.
2.What are the characteristics of Big Data ?
Ans:
i) Volume- vast 'volumes' of data is generated from many sources daily, such as business processes, machines, social media platforms, networks, human interactions, and many more.
ii) Variety- Big Data can be structured, unstructured, and semi-structured that are being collected from different sources.
iii) Velocity- Velocity creates the speed by which the data is created in real-time.The primary aspect of Big Data is to provide demanding data rapidly.
iv) Veracity- Veracity means how reliable the data is. It has many ways to filter or translate the data. Veracity is the process of being able to handle and manage data efficiently.
3.Distinguish between Traditional data and Big data
Ans:
Traditional Data:
● Traditional data is the structured data which is being majorly maintained by all types of businesses starting from very small to big organizations.
● In traditional database system a centralized database architecture used to store and maintain the data in a fixed format or fields in a file
Big Data:
● Big data can be considered as an upper version of traditional data.
● Big data deals with too large or complex data sets which is difficult to manage in traditional data-processing application software.
● It deals with large volumes of both structured, semi structured and unstructured data.
4.What is Hadoop? Why is Hadoop used in big data ?
Ans:
Hadoop is an Apache open source framework written in java that allows distributed processing of large datasets across clusters of computers using simple programming models.
Hadoop is used in Big data because it allows companies to manage huge volumes of data easily. It allows big problems to be broken down into smaller elements so that analysis could be done quickly and cost-effectively.
5.What are the components of Hadoop ?
Ans:
Hadoop has two major layers namely −
● Processing/Computation layer (MapReduce)
● Storage layer (Hadoop Distributed File System)
● Hadoop Common
● Hadoop YARN
6. Explain the components of Hadoop Architecture.
Ans:
MapReduce
MapReduce is a parallel programming model for writing distributed applications devised at Google for efficient processing of large amounts of data.
Hadoop Distributed File System (HDFS)
The Hadoop Distributed File System (HDFS) is based on the Google File System (GFS) and provides a distributed file system that is designed to run on commodity hardware.
Hadoop Common
These are Java libraries and utilities required by other Hadoop modules.
Hadoop YARN (Yet Another Resource Navigator)
This is a framework for job scheduling and cluster resource management.
7. Explain the Hadoop Ecosystem
Ans:
The Hadoop Ecosystem has the following stages:
i) Data Management
ii) Data Access
iii) Data Processing
iv) Data Storage
● HDFS: Hadoop Distributed File System
● YARN: Yet Another Resource Negotiator
● MapReduce: Programming based Data Processing
● Spark: In-Memory data processing
● PIG, HIVE: Query based processing of data services
● HBase: NoSQL Database
● Mahout, Spark MLLib: Machine Learning algorithm libraries
● Solar, Lucene: Searching and Indexing
● Zookeeper: Managing cluster
● Oozie: Job Scheduling
