Machine Learning Is One Of The Most Important Technologies Used In Modern Software Development, Data Analytics, And Artificial Intelligence Applications. Infosys Frequently Assesses Candidates On Fundamental Machine Learning Concepts, Algorithms, Data Preprocessing Techniques, Model Evaluation Metrics, And Real-World Applications During Technical Interviews. Freshers Are Expected To Understand Topics Such As Supervised Learning, Unsupervised Learning, Regression, Classification, Neural Networks, And Deep Learning Basics. Interviewers May Also Ask Questions Related To Python, Statistics, Data Handling, And Machine Learning Libraries. Preparing Common Machine Learning Interview Questions Helps Candidates Build Confidence And Improve Problem-Solving Skills. This Collection Of Infosys Machine Learning Interview Questions And Answers Will Help Freshers Understand Key Concepts And Prepare Effectively For Technical Interviews And AI-Related Roles.
1. What Is Machine Learning?
Ans:
- Machine Learning Is A Branch Of Artificial Intelligence That Enables Systems To Learn From Data.
- It Identifies Patterns Without Being Explicitly Programmed. Algorithms Improve Their Performance Through Experience. Machine Learning Is Used In Recommendation Systems, Fraud Detection, And Healthcare.
- It Can Handle Large Volumes Of Data Efficiently. Models Are Trained Using Historical Data. The Goal Is To Make Accurate Predictions Or Decisions.
2. What Are The Types Of Machine Learning?
Ans:
Machine Learning Is Divided Into Supervised, Unsupervised, And Reinforcement Learning. Supervised Learning Uses Labeled Data For Training. Unsupervised Learning Finds Hidden Patterns In Unlabeled Data. Reinforcement Learning Learns Through Rewards And Penalties. Each Type Serves Different Business Problems. Choosing The Right Approach Depends On Data Availability. These Methods Form The Foundation Of Modern AI Systems.
3. What Is Supervised Learning?
Ans:
Supervised Learning Trains A Model Using Input And Output Pairs. The Algorithm Learns The Relationship Between Features And Targets. It Is Commonly Used For Classification And Regression Tasks. Examples Include Spam Detection And House Price Prediction. Labeled Data Is Required For Training. The Model Attempts To Minimize Prediction Errors. Accuracy Improves As More Relevant Data Is Provided.
4. What Is Unsupervised Learning?
Ans:
- Unsupervised Learning Works With Data That Has No Labels. The Goal Is To Discover Hidden Structures And Patterns.
- Clustering And Dimensionality Reduction Are Common Techniques. Customer Segmentation Is A Popular Application.
- It Helps Organizations Understand Large Datasets Better. No Predefined Output Is Available During Training. Insights Generated Can Support Business Decisions.
5. What Is Reinforcement Learning?
Ans:
Reinforcement Learning Trains An Agent To Make Decisions Through Interaction. The Agent Receives Rewards Or Penalties Based On Actions. It Learns The Best Strategy Over Time. This Technique Is Used In Robotics And Gaming. Exploration And Exploitation Are Key Concepts. The Objective Is To Maximize Cumulative Rewards. Reinforcement Learning Mimics Human Learning Behavior.
6. What Is A Dataset In Machine Learning?
Ans:
A Dataset Is A Collection Of Data Used For Training And Testing Models. It Contains Features And Sometimes Labels. High-Quality Data Improves Model Performance. Datasets Can Be Structured Or Unstructured. Data Preparation Is Important Before Training Begins. Machine Learning Algorithms Depend On Relevant Information. Proper Datasets Lead To Better Predictions And Insights.
7. Write A Program To Predict Values Using Linear Regression
Ans:
Linear Regression Is Used To Predict Continuous Values Based On Input Features. It Finds The Best-Fit Line Between Variables. The Algorithm Learns Relationships From Historical Data.
- from sklearn.linear_model import LinearRegression
- model = LinearRegression()
- model.fit([[1],[2],[3]],[2,4,6])
- print(model.predict([[4]]))
8. What Is A Training Dataset?
Ans:
A Training Dataset Is Used To Teach A Machine Learning Model. The Algorithm Learns Patterns And Relationships From This Data. It Contains Historical Examples Relevant To The Problem. A Larger Dataset Often Improves Generalization. Data Quality Is More Important Than Quantity Alone. Training Is The First Major Step In Model Development. The Model Adjusts Parameters Based On Training Data.
9. What Is A Test Dataset?
Ans:
A Test Dataset Evaluates The Performance Of A Trained Model. It Contains Data Not Seen During Training. Testing Measures How Well The Model Generalizes. Metrics Such As Accuracy And Precision Are Calculated. Proper Testing Prevents Misleading Results. The Dataset Should Represent Real-World Scenarios. It Helps Determine Whether Deployment Is Appropriate.
10. What Is Overfitting?
Ans:
Overfitting Occurs When A Model Learns Training Data Too Well. It Captures Noise Instead Of General Patterns. Such Models Perform Poorly On New Data. Complex Models Are More Prone To Overfitting. Techniques Like Cross-Validation Can Reduce The Risk. Regularization Methods Also Help Improve Generalization. The Goal Is To Balance Accuracy And Flexibility.
11. What Is Underfitting?
Ans:
Underfitting Occurs When A Model Is Too Simple To Capture Data Patterns. It Performs Poorly On Both Training And Testing Data. Important Relationships Remain Unlearned During Training. Low Accuracy Is A Common Sign Of Underfitting. Increasing Model Complexity May Help Improve Results. Better Feature Engineering Can Also Enhance Performance. A Good Model Should Avoid Both Underfitting And Overfitting.
12. What Is Classification In Machine Learning?
Ans:
Classification Is A Supervised Learning Technique Used To Predict Categories. The Output Variable Belongs To A Specific Class. Examples Include Email Spam Detection And Disease Diagnosis. Algorithms Learn From Labeled Training Data. Predictions Are Made Based On Learned Patterns. Accuracy Is Commonly Used To Evaluate Classification Models. It Is One Of The Most Popular Machine Learning Tasks.
13. What Is Regression In Machine Learning?
Ans:
- Regression Is A Supervised Learning Method Used To Predict Continuous Values. The Output Is Numerical Rather Than Categorical.
- House Price Prediction Is A Common Example. Regression Models Identify Relationships Between Variables.
- Linear Regression Is A Widely Used Technique. Performance Is Evaluated Using Error Metrics. Regression Helps In Forecasting And Trend Analysis.
14. What Is Linear Regression?
Ans:
Linear Regression Is A Statistical Method For Predicting Continuous Values. It Establishes A Linear Relationship Between Variables. The Algorithm Fits A Straight Line To Data Points. It Is Easy To Implement And Interpret. Linear Regression Is Commonly Used In Forecasting Problems. The Goal Is To Minimize Prediction Errors. It Serves As A Foundation For Advanced Models.
15. What Is Logistic Regression?
Ans:
Logistic Regression Is A Classification Algorithm Despite Its Name. It Predicts Probabilities For Different Classes. The Output Is Usually Between Zero And One. It Is Commonly Used For Binary Classification Problems. Examples Include Fraud Detection And Medical Diagnosis. The Sigmoid Function Converts Outputs Into Probabilities. Logistic Regression Is Simple Yet Highly Effective.
16. What Is A Decision Tree?
Ans:
A Decision Tree Is A Supervised Learning Algorithm Used For Classification And Regression. It Splits Data Into Branches Based On Conditions. The Structure Resembles A Flowchart. Decision Trees Are Easy To Understand And Visualize. They Handle Both Numerical And Categorical Data. Overfitting Can Occur With Deep Trees. Pruning Techniques Help Improve Performance.
17. What Is Random Forest?
Ans:
- Random Forest Is An Ensemble Learning Method Based On Multiple Decision Trees. It Combines Predictions From Several Trees.
- This Approach Improves Accuracy And Stability. Random Forest Reduces The Risk Of Overfitting. It Works Well With Large Datasets.
- Feature Importance Can Also Be Measured. The Algorithm Is Widely Used In Real-World Applications.
18. What Is K-Nearest Neighbors (KNN)?
Ans:
K-Nearest Neighbors Is A Simple Supervised Learning Algorithm. It Classifies Data Based On Nearby Data Points. Similar Instances Tend To Belong To The Same Category. The Value Of K Determines The Number Of Neighbors Considered. KNN Is Easy To Understand And Implement. Performance Depends On Distance Measurement Techniques. It Is Commonly Used For Classification Problems.
19.Difference Between Bias And Variance
Ans:
| Aspect | Bias | Variance |
|---|---|---|
| Definition | Error Caused By Oversimplified Assumptions In The Model. | Error Caused By Excessive Sensitivity To Training Data |
| Effect On Model | Leads To Underfitting. | Leads To Overfitting |
| Learning Behavior | Fails To Capture Important Data Patterns.p | Learns Noise Along With Actual Patterns. |
| Training Accuracy | Usually Low | Usually High |
20. What Is Naive Bayes Algorithm?
Ans:
- Naive Bayes Is A Probabilistic Classification Algorithm Based On Bayes’ Theorem. It Assumes Features Are Independent Of Each Other.
- Despite This Assumption, It Performs Surprisingly Well. It Is Commonly Used For Spam Filtering. Naive Bayes Is Fast And Computationally Efficient.
- It Works Effectively With Large Datasets. The Algorithm Is Popular In Text Classification Tasks.
21. What Is Clustering In Machine Learning?
Ans:
Clustering Is An Unsupervised Learning Technique That Groups Similar Data Points Together. It Helps Discover Hidden Patterns Within Datasets. No Labeled Data Is Required For Training. Customer Segmentation Is A Common Application. Clustering Improves Data Understanding And Analysis. Different Algorithms Produce Different Grouping Results. It Is Widely Used In Business Intelligence Projects.
22. What Is K-Means Clustering?
Ans:
K-Means Is A Popular Clustering Algorithm Used To Partition Data Into Groups. It Assigns Data Points To The Nearest Cluster Center. The Number Of Clusters Is Specified In Advance. The Algorithm Iteratively Updates Cluster Centers. It Is Efficient And Easy To Implement. K-Means Works Best With Well-Separated Clusters. It Is Commonly Used For Customer Segmentation.
23. What Is Dimensionality Reduction?
Ans:
- Dimensionality Reduction Reduces The Number Of Input Variables In A Dataset. It Helps Simplify Complex Data Structures.
- Fewer Dimensions Improve Processing Speed And Efficiency. It Also Reduces Noise And Overfitting Risks.
- PCA Is A Popular Dimensionality Reduction Technique. Visualization Becomes Easier With Reduced Dimensions. The Process Improves Model Performance In Many Cases.
24. What Is Principal Component Analysis (PCA)?
Ans:
PCA Is A Dimensionality Reduction Technique Used In Machine Learning. It Converts Correlated Variables Into Uncorrelated Components. The First Components Capture Most Data Variance. PCA Helps Reduce Computational Complexity. It Improves Data Visualization And Interpretation. Noise Can Also Be Reduced Through PCA. It Is Widely Used In Data Preprocessing.
25. What Is Feature Engineering?
Ans:
Feature Engineering Involves Creating Or Transforming Variables For Better Predictions. It Enhances Model Accuracy And Performance. Domain Knowledge Plays An Important Role. New Features Can Reveal Hidden Data Patterns. Proper Feature Engineering Improves Learning Efficiency. It Is A Critical Step In Machine Learning Projects. Successful Models Often Depend On Quality Features.
26.Write A Program To Calculate Mean Of A Dataset
Ans:
Mean Is One Of The Most Important Statistical Measures In Machine Learning. It Represents The Average Value Of A Dataset. Data Scientists Use Mean To Understand Data Distribution
- data = [10, 20, 30, 40]
- mean = sum(data) / len(data)
- print(mean)
27. What Is Cross Validation?
Ans:
Cross Validation Evaluates Model Performance Using Multiple Data Splits. It Provides A More Reliable Accuracy Estimate. Training And Validation Occur On Different Data Segments. K-Fold Cross Validation Is Widely Used. It Reduces Bias In Performance Measurement. The Technique Helps Select Better Models. Cross Validation Improves Confidence In Results.
28. Write A Program To Normalize Data
Ans:
Normalization Scales Data Into A Common Range Such As Zero To One. It Prevents Features With Large Values From Dominating Others. Many Machine Learning Algorithms Benefit From Normalization.
- data = [10, 20, 30]
- norm = [(x-min(data))/(max(data)-min(data)) for x in data]
- print(norm)
29. What Is Precision?
Ans:
Precision Measures The Proportion Of Correct Positive Predictions. It Focuses On Prediction Quality Rather Than Quantity. High Precision Means Fewer False Positives. It Is Important In Fraud Detection Applications. Precision Helps Evaluate Classification Models Effectively. It Is Often Used With Recall Metrics. Together They Provide Better Insights Into Performance.
30. What Is Recall?
Ans:
Recall Measures The Ability To Identify Actual Positive Cases. It Calculates The Percentage Of Positives Correctly Detected. High Recall Means Fewer False Negatives. Medical Diagnosis Often Requires High Recall. It Helps Ensure Important Cases Are Not Missed. Recall Complements Precision Metrics. Both Are Important For Model Evaluation.Recall Is Especially Valuable In Applications Where Missing Positive Cases Can Have Serious Consequences.
31. What Is F1 Score?
Ans:
F1 Score Combines Precision And Recall Into A Single Metric. It Provides A Balanced Performance Measure. The Metric Is Useful For Imbalanced Datasets. High F1 Scores Indicate Strong Classification Performance. It Helps Compare Different Models Effectively. F1 Score Reduces Dependence On Accuracy Alone. It Is Widely Used In Machine Learning Evaluation.
32. What Is A Confusion Matrix?
Ans:
- A Confusion Matrix Summarizes Classification Results In Table Form. It Shows Correct And Incorrect Predictions.
- The Matrix Includes True Positives And False Positives. It Also Displays True Negatives And False Negatives.
- Performance Metrics Are Derived From It. The Matrix Provides Detailed Evaluation Insights. It Is Essential For Classification Analysis.
33. What Is Bias In Machine Learning?
Ans:
- Bias Refers To Errors Caused By Oversimplified Model Assumptions. High Bias Often Leads To Underfitting.
- The Model Fails To Capture Important Patterns. Bias Reduces Prediction Accuracy. Selecting Better Features Can Reduce Bias. Balanced Model Complexity Is Important.
- Understanding Bias Improves Model Development. Managing Bias Helps Create More Accurate And Reliable Machine Learning Models.
34. What Is Variance In Machine Learning?
Ans:
Variance Measures A Model’s Sensitivity To Training Data Changes. High Variance Often Causes Overfitting. The Model Learns Noise Instead Of Patterns. Predictions Become Unstable Across Datasets. Reducing Variance Improves Generalization. Ensemble Methods Can Help Address Variance Issues. A Balance Between Bias And Variance Is Necessary. Controlling Variance Ensures Consistent Performance On Unseen Data.
35. What Is The Bias-Variance Tradeoff?
Ans:
The Bias-Variance Tradeoff Balances Simplicity And Complexity In Models. High Bias Causes Underfitting. High Variance Causes Overfitting. The Goal Is To Achieve Optimal Performance. Proper Model Selection Is Important. Cross Validation Helps Find The Right Balance. This Concept Is Fundamental In Machine Learning. Achieving The Right Tradeoff Leads To Better Generalization And Accuracy.
36. What Is Gradient Descent?
Ans:
Gradient Descent Is An Optimization Algorithm Used For Training Models. It Minimizes Errors By Updating Parameters. The Algorithm Moves Toward Lower Cost Values. Learning Rate Controls Update Size. Proper Tuning Is Important For Convergence. Gradient Descent Is Used In Many Algorithms. It Plays A Key Role In Deep Learning. Efficient Optimization Helps Models Learn Faster And Perform Better.
37. What Is Learning Rate?
Ans:
- Learning Rate Determines The Step Size During Model Optimization. A Small Learning Rate Slows Training.
- A Large Learning Rate May Cause Instability. Proper Selection Improves Convergence Speed. It Directly Affects Model Performance.
- Learning Rate Is Critical In Gradient Descent. Careful Tuning Produces Better Results. Choosing An Appropriate Learning Rate Is Essential For Successful Training.
38. What Is Cost Function?
Ans:
A Cost Function Measures Prediction Errors In A Model. It Quantifies The Difference Between Actual And Predicted Values. Training Aims To Minimize Cost. Different Algorithms Use Different Cost Functions. Lower Costs Indicate Better Performance. Optimization Methods Reduce Cost Values. Cost Functions Guide The Learning Process. They Provide A Clear Measure Of How Well A Model Is Performing.
39. What Is Regularization?
Ans:
Regularization Prevents Overfitting By Adding Constraints To Models. It Reduces Model Complexity. L1 And L2 Are Common Techniques. Regularization Improves Generalization Ability. It Helps Models Perform Better On New Data. The Method Controls Excessive Parameter Growth. Regularization Is Important In Predictive Modeling. It Enhances Model Stability And Long-Term Performance.
40. What Is L1 Regularization?
Ans:
L1 Regularization Adds A Penalty Based On Absolute Coefficient Values. It Encourages Sparse Models. Some Features Receive Zero Weights. Feature Selection Occurs Automatically. The Technique Helps Simplify Models. L1 Reduces Overfitting Risks. It Is Also Known As Lasso Regression. L1 Regularization Improves Interpretability By Reducing Unnecessary Features.
41. What Is L2 Regularization?
Ans:
- L2 Regularization Adds A Penalty Based On Squared Coefficient Values. It Reduces Excessively Large Weights.
- Model Stability Improves Through This Method. L2 Helps Prevent Overfitting. Unlike L1, Features Usually Remain Present.
- It Produces Smooth Solutions. The Technique Is Known As Ridge Regression. L2 Regularization Encourages Better Generalization Across Different Datasets.
42. What Is Ensemble Learning?
Ans:
Ensemble Learning Combines Multiple Models To Improve Performance. It Reduces Prediction Errors And Variance. Different Algorithms Contribute To Final Results. Random Forest Is An Example. Ensemble Methods Improve Robustness. Accuracy Often Increases Compared To Single Models. They Are Widely Used In Industry Applications. Combining Multiple Models Often Produces More Reliable Predictions.
43. What Is Bagging?
Ans:
Bagging Stands For Bootstrap Aggregating In Machine Learning. Multiple Models Are Trained On Random Samples. Predictions Are Combined For Final Output. Bagging Reduces Variance And Overfitting. Random Forest Uses This Technique. It Improves Stability And Accuracy. The Method Works Well For Decision Trees. Bagging Creates Stronger Models By Leveraging Multiple Learners.
44. What Is Boosting?
Ans:
Boosting Builds Models Sequentially To Correct Previous Errors. Each New Model Focuses On Difficult Cases. Performance Improves Through Iterative Learning. AdaBoost And XGBoost Are Popular Examples. Boosting Often Produces High Accuracy. It Can Handle Complex Prediction Tasks. Proper Tuning Is Essential For Best Results. Boosting Is Highly Effective For Challenging Machine Learning Problems.
45. What Is XGBoost?
Ans:
- XGBoost Is An Advanced Boosting Algorithm Used For Prediction Tasks. It Provides High Accuracy And Efficiency.
- The Algorithm Handles Large Datasets Effectively. Regularization Reduces Overfitting Risks. XGBoost Supports Parallel Processing.
- It Is Popular In Data Science Competitions. Many Businesses Use It For Predictive Analytics. Its Speed And Performance Make It A Preferred Industry Choice
46. What Is Deep Learning?
Ans:
Deep Learning Is A Subset Of Machine Learning Based On Neural Networks. It Uses Multiple Hidden Layers For Learning. Complex Patterns Can Be Identified Automatically. Deep Learning Excels In Image And Speech Recognition. Large Datasets Are Usually Required. High Computational Power Is Often Necessary. It Powers Many Modern AI Applications. Deep Learning Continues To Drive Innovations In Artificial Intelligence.
47. What Is An Artificial Neural Network?
Ans:
An Artificial Neural Network Mimics The Structure Of The Human Brain. It Consists Of Input, Hidden, And Output Layers. Neurons Process And Transfer Information. Networks Learn By Adjusting Weights. Neural Networks Handle Complex Problems Effectively. They Are Used In Classification And Prediction Tasks. Deep Learning Relies On Neural Networks. These Networks Form The Foundation Of Modern AI Systems.
48. What Is An Epoch In Machine Learning?
Ans:
An Epoch Represents One Complete Pass Through The Training Dataset. Models Usually Require Multiple Epochs. Learning Improves After Repeated Passes. Too Many Epochs May Cause Overfitting. Training Progress Is Monitored Across Epochs. Optimization Occurs During Each Epoch. It Is A Common Deep Learning Term. Proper Epoch Selection Helps Achieve Better Model Performance.

49. What Is Batch Size?
Ans:
Batch Size Refers To The Number Of Samples Processed At Once. Training Occurs In Small Data Groups. Larger Batches Require More Memory. Smaller Batches Increase Training Iterations. Batch Size Affects Learning Speed. Proper Selection Improves Efficiency. It Is Important In Neural Network Training. Choosing The Right Batch Size Balances Speed And Accuracy.
50. What Is Transfer Learning?
Ans:
- Transfer Learning Reuses Knowledge From Previously Trained Models. It Reduces Training Time And Data Requirements.
- Pretrained Models Serve As Starting Points. This Technique Is Popular In Computer Vision. Performance Often Improves With Limited Data.
- Transfer Learning Saves Computational Resources. It Accelerates Machine Learning Development. It Enables Faster Deployment Of High-Performing AI Solutions.
51. What Is Data Preprocessing?
Ans:
Data Preprocessing Is The Process Of Preparing Raw Data For Machine Learning Models. It Includes Cleaning, Transforming, And Organizing Data. Missing Values And Errors Are Handled During This Stage. Proper Preprocessing Improves Model Accuracy And Reliability. It Helps Algorithms Learn More Effectively From Data. Feature Scaling And Encoding Are Common Techniques. Data Preprocessing Is A Critical Step In Every Machine Learning Project.
52. What Is Data Cleaning?
Ans:
- Data Cleaning Involves Detecting And Correcting Errors In Datasets. It Removes Duplicate Records And Inconsistent Values.
- Missing Data Is Also Addressed During Cleaning. Clean Data Improves The Quality Of Predictions.
- It Reduces Noise And Unwanted Variations. Accurate Data Leads To Better Model Performance. Data Cleaning Is Essential Before Model Training Begins.
53. What Is Missing Data?
Ans:
Missing Data Refers To Unavailable Values Within A Dataset. It Can Occur Due To Errors Or Incomplete Collection. Missing Values May Affect Model Performance Negatively. Techniques Such As Imputation Can Handle Missing Data. Proper Treatment Prevents Biased Results. Understanding Missing Patterns Is Important. Effective Handling Improves Data Quality And Accuracy.
54. Write A Program To Implement K-Means Clustering
Ans:
K-Means Is An Unsupervised Learning Algorithm Used For Clustering. It Groups Similar Data Points Together. The Algorithm Creates Clusters Based On Distance Measures.
- from sklearn.cluster import KMeans
- model = KMeans(n_clusters=2, random_state=0)
- model.fit([[1],[2],[8],[9]])
- print(model.labels_)
55. What Is Standardization?
Ans:
Standardization Transforms Data To Have A Mean Of Zero. It Also Adjusts Values To A Standard Deviation Of One. This Technique Helps Compare Different Features Fairly. Many Machine Learning Algorithms Benefit From Standardization. It Improves Optimization Efficiency During Training. Standardized Data Reduces Scale-Related Issues. It Is Commonly Applied In Predictive Modeling.
56. What Is One-Hot Encoding?
Ans:
- One-Hot Encoding Converts Categorical Data Into Binary Columns. Each Category Receives Its Own Unique Column.
- This Prevents Algorithms From Assuming Numeric Relationships. It Is Widely Used In Classification Problems.
- One-Hot Encoding Improves Data Representation. Many Machine Learning Libraries Support This Technique. It Helps Models Process Categorical Variables Correctly.
57. What Is Label Encoding?
Ans:
Label Encoding Assigns Numerical Values To Categories. Each Unique Category Receives A Distinct Number. It Simplifies Data Processing For Algorithms. The Technique Is Easy To Implement And Efficient. Care Must Be Taken To Avoid Misleading Relationships. Label Encoding Is Suitable For Certain Models. It Is Commonly Used In Data Preparation.
58. What Is Feature Scaling?
Ans:
Feature Scaling Adjusts Variables To Similar Numerical Ranges. It Prevents Features With Large Values From Dominating. Scaling Improves Training Speed And Stability. Algorithms Like KNN And SVM Benefit Greatly. Standardization And Normalization Are Common Methods. Proper Scaling Enhances Model Accuracy. It Is An Important Preprocessing Technique.
59. What Is Outlier Detection?
Ans:
Outlier Detection Identifies Unusual Data Points In A Dataset. These Values Differ Significantly From Normal Observations. Outliers Can Affect Model Accuracy And Predictions. Statistical And Machine Learning Methods Detect Them. Handling Outliers Improves Data Quality. Sometimes Outliers Provide Valuable Business Insights. Proper Analysis Is Necessary Before Removal.
60. What Is Data Imbalance?
Ans:
Data Imbalance Occurs When Classes Have Unequal Sample Sizes. One Class May Contain Significantly More Records. This Can Lead To Biased Predictions. Models Often Favor The Majority Class. Techniques Such As Resampling Help Address Imbalance. Evaluation Metrics Must Be Chosen Carefully. Balanced Data Improves Classification Performance.
61. What Is ROC Curve?
Ans:
ROC Curve Stands For Receiver Operating Characteristic Curve. It Evaluates Classification Performance Across Thresholds. The Curve Compares True Positive And False Positive Rates. Better Models Produce Curves Closer To The Top Left. ROC Analysis Helps Compare Different Models. It Provides A Visual Performance Representation. The Technique Is Widely Used In Classification Tasks.
62. Difference Between Precision And Recall
Ans:
| Aspect | Precision | Recall |
|---|---|---|
| Definition | Measures How Many Predicted Positive Cases Are Actually Positive | Measures How Many Actual Positive Cases Are Correctly Identified. |
| Formula | Precision = TP / (TP + FP) | SRecall = TP / (TP + FN) |
| Focus | Focuses On The Accuracy Of Positive Predictions. | Focuses On Finding All Positive Instances. |
| Important Metric | False Positives. | False Negatives. |
63. What Is Hyperparameter Tuning?
Ans:
- Hyperparameter Tuning Optimizes Model Settings Before Training. Parameters Such As Learning Rate Can Be Adjusted.
- Proper Tuning Improves Accuracy And Efficiency. Different Combinations Produce Different Results.
- Automated Search Techniques Are Often Used. Tuning Helps Achieve Better Generalization. It Is An Important Step In Model Development.
64. What Is Grid Search?
Ans:
Grid Search Is A Hyperparameter Optimization Technique. It Tests All Possible Parameter Combinations. The Best Configuration Is Selected Based On Performance. Grid Search Is Systematic And Easy To Understand. It Can Be Computationally Expensive. The Method Works Well For Small Parameter Spaces. It Helps Improve Model Accuracy.Grid Search Is Widely Used To Fine-Tune Machine Learning Models For Better Results.
65. What Is Random Search?
Ans:
Random Search Selects Hyperparameter Combinations Randomly. It Explores The Search Space More Efficiently. Fewer Evaluations May Produce Good Results. Random Search Often Requires Less Computation. It Is Useful For Large Parameter Spaces. The Technique Can Outperform Grid Search In Some Cases. It Is Popular In Practical Machine Learning.
66. What Is Natural Language Processing?
Ans:
Natural Language Processing Enables Computers To Understand Human Language. It Combines Linguistics And Artificial Intelligence. NLP Is Used In Chatbots And Translation Systems. Text Classification Is A Common Application. The Technology Extracts Meaning From Text Data. NLP Continues To Advance Rapidly. It Powers Many Modern AI Solutions.

67. What Is Tokenization?
Ans:
okenization Splits Text Into Smaller Meaningful Units. These Units May Be Words Or Sentences. It Is A Fundamental NLP Preprocessing Step. Tokens Help Models Analyze Language Efficiently. Different Languages Require Different Tokenization Methods. Accurate Tokenization Improves NLP Performance. The Process Supports Various Text Analytics Tasks.
68. Write A Program To Split Dataset Into Training And Testing Sets
Ans:
Data Splitting Is An Important Step In Machine Learning. It Separates Data Into Training And Testing Portions. Training Data Is Used To Build The Model.
- from sklearn.model_selection import train_test_split
- X_train,X_test = train_test_split([1,2,3,4,5], test_size=0.2)
- print(X_train, X_test)
69. What Is Lemmatization?
Ans:
- Lemmatization Converts Words Into Their Dictionary Forms. It Uses Linguistic Knowledge To Determine Base Words.
- The Results Are Usually More Meaningful Than Stemming. Lemmatization Improves Text Understanding.
- It Is Commonly Used In NLP Applications. The Process Preserves Language Structure Better. Accurate Lemmatization Enhances Model Performance.
70. What Is Sentiment Analysis?
Ans:
Sentiment Analysis Determines Emotional Tone In Text Data. It Identifies Positive, Negative, Or Neutral Opinions. Businesses Use It To Analyze Customer Feedback. Social Media Monitoring Often Relies On Sentiment Analysis. Machine Learning Models Detect Language Patterns. The Technique Supports Better Decision Making. It Is A Popular NLP Application.
71. What Is Computer Vision?
Ans:
Computer Vision Enables Machines To Interpret Visual Information. It Processes Images And Videos Automatically. Applications Include Facial Recognition And Medical Imaging. Deep Learning Has Improved Computer Vision Significantly. Large Datasets Are Often Required For Training. The Technology Supports Automation Across Industries. Computer Vision Is A Key AI Domain.
72. What Is Image Classification?
Ans:
Image Classification Assigns Labels To Images Based On Content. Models Learn Patterns From Labeled Visual Data. Applications Include Object Recognition And Medical Diagnosis. Deep Learning Techniques Achieve High Accuracy. Large Training Datasets Improve Results. Classification Helps Automate Visual Analysis Tasks. It Is Widely Used In Industry Applications.
73. What Is Object Detection?
Ans:
- Object Detection Identifies And Locates Objects Within Images. It Combines Classification And Localization Tasks.
- Bounding Boxes Mark Object Positions. Applications Include Autonomous Vehicles And Surveillance Systems.
- Deep Learning Models Perform Detection Efficiently. Accuracy Depends On Data Quality And Training. Object Detection Enhances Visual Intelligence Systems.
.
74. What Is CNN?
Ans:
Convolutional Neural Networks Are Specialized Deep Learning Models. They Are Designed For Image Processing Tasks. CNNs Automatically Learn Visual Features. Convolution Layers Extract Important Patterns. Pooling Layers Reduce Computational Complexity. CNNs Achieve Excellent Results In Computer Vision. They Are Widely Used In Industry Applications.
75. What Is RNN?
Ans:
Recurrent Neural Networks Process Sequential Data Effectively. They Maintain Information From Previous Inputs. RNNs Are Useful For Language And Time-Series Tasks. The Architecture Handles Ordered Data Naturally. Traditional RNNs Struggle With Long Dependencies. Improved Variants Address This Limitation. RNNs Are Important In Sequence Modeling. RNNs Enable Machines To Learn Patterns From Sequential And Temporal Data.
76. What Is LSTM?
Ans:
LSTM Stands For Long Short-Term Memory Network. It Is An Advanced Form Of RNN. LSTM Handles Long-Term Dependencies Effectively. Memory Cells Store Important Information Over Time. It Is Commonly Used In Language Processing. LSTMs Improve Prediction Accuracy For Sequential Data. They Are Widely Applied In Deep Learning. LSTMs Are Highly Effective For Tasks Requiring Long-Term Context Retention.
77. What Is Time Series Analysis?
Ans:
Time Series Analysis Examines Data Collected Over Time. It Identifies Trends And Seasonal Patterns. Forecasting Is A Common Objective. Historical Data Supports Future Predictions. Financial And Business Applications Frequently Use It. Specialized Models Handle Time Dependencies. Time Series Analysis Supports Strategic Planning.
78. What Is Forecasting?
Ans:
Forecasting Predicts Future Outcomes Using Historical Data. Machine Learning Models Analyze Past Patterns. Businesses Use Forecasting For Demand Planning. Accurate Forecasts Improve Decision Making. Various Statistical And AI Techniques Are Available. Forecasting Supports Resource Optimization. It Is Important Across Many Industries.
79. What Is A Recommendation System?
Ans:
- Recommendation Systems Suggest Relevant Products Or Content. They Analyze User Preferences And Behavior.
- Online Shopping Platforms Use Recommendations Extensively. Personalized Suggestions Improve User Experience.
- Machine Learning Powers Modern Recommendation Engines. The Systems Increase Engagement And Revenue. They Are Widely Used Across Digital Platforms.
80. What Is Collaborative Filtering?
Ans:
- Collaborative Filtering Recommends Items Based On Similar Users. It Analyzes Past Interactions And Preferences.
- The Method Does Not Require Product Details. Recommendations Improve As More Data Becomes Available.
- Collaborative Filtering Is Common In Streaming Services. It Supports Personalized User Experiences. The Technique Is Popular In Recommendation Systems.
81. Write A Program To Calculate Accuracy Score
Ans:
Accuracy Measures The Percentage Of Correct Predictions Made By A Model. It Is One Of The Most Common Evaluation Metrics. Higher Accuracy Indicates Better Performance.
- from sklearn.metrics import accuracy_score
- y_true = [1,0,1]
- y_pred = [1,0,0]
- print(accuracy_score(y_true, y_pred))
82. What Is Reward Function?
Ans:
A Reward Function Defines Feedback For Agent Actions. Positive Rewards Encourage Desired Behavior. Negative Rewards Discourage Incorrect Actions. The Function Guides Learning Progress. Reward Design Influences Performance Significantly. Effective Rewards Lead To Better Strategies. It Is A Core Component Of Reinforcement Learning.
83. What Is Q-Learning?
Ans:
Q-Learning Is A Popular Reinforcement Learning Algorithm. It Learns The Value Of Actions In States. The Algorithm Builds A Q-Table Over Time. No Model Of The Environment Is Required. Q-Learning Supports Optimal Decision Making. It Learns Through Repeated Interactions. The Technique Is Widely Studied In AI.
84. What Is Model Deployment?
Ans:
Model Deployment Makes Trained Models Available For Use. Predictions Can Be Generated In Real Time. Deployment Connects Models To Applications. Monitoring Continues After Deployment. Reliable Deployment Ensures Business Value. Different Platforms Support Model Hosting. It Is The Final Stage Of Machine Learning Projects.Proper Deployment Ensures Models Deliver Consistent And Scalable Performance In Production Environments.
85. What Is MLOps?
Ans:
MLOps Combines Machine Learning With Operational Practices. It Automates Model Development And Deployment Processes. Collaboration Between Teams Improves Efficiency. MLOps Supports Continuous Monitoring And Updates. It Enhances Scalability And Reliability. The Approach Reduces Deployment Challenges. MLOps Is Important In Enterprise AI Systems.
86. What Is Model Monitoring?
Ans:
Model Monitoring Tracks Performance After Deployment. It Detects Accuracy Drops And Unexpected Behavior. Monitoring Ensures Reliable Predictions Over Time. Alerts Can Identify Potential Issues Quickly. Data Changes May Affect Performance. Continuous Evaluation Supports Maintenance Efforts. Monitoring Is Essential For Production Systems.
87. What Is Concept Drift?
Ans:
Concept Drift Occurs When Data Patterns Change Over Time. Models Trained On Old Data May Become Less Accurate. Continuous Monitoring Helps Detect Drift. Retraining Is Often Required. Business Environments Frequently Experience Changes. Drift Can Impact Prediction Reliability. Managing It Maintains Model Effectiveness.
88. What Is Explainable AI?
Ans:
Explainable AI Helps Users Understand Model Decisions. It Improves Transparency And Trust In AI Systems. Organizations Need Explanations For Critical Decisions. Explainability Supports Regulatory Compliance. Users Gain Insights Into Prediction Factors. It Reduces Concerns About Black-Box Models. Explainable AI Is Growing In Importance.
89. What Is SHAP?
Ans:
Explainable AI Helps Users Understand Model Decisions. It Improves Transparency And Trust In AI Systems. Organizations Need Explanations For Critical Decisions. Explainability Supports Regulatory Compliance. Users Gain Insights Into Prediction Factors. It Reduces Concerns About Black-Box Models. Explainable AI Is Growing In Importance.
90. What Is Data Augmentation?
Ans:
Data Augmentation Creates Additional Training Examples Artificially. It Increases Dataset Diversity And Size. Techniques Include Rotation And Flipping For Images. Augmentation Improves Model Generalization. It Reduces Overfitting Risks. More Data Often Leads To Better Performance. The Method Is Common In Deep Learning. Data Augmentation Enhances Model Robustness Without Collecting New Data.
91. What Is Generative AI?
Ans:
Generative AI Creates New Content Such As Text And Images. It Learns Patterns From Large Datasets. Modern Models Produce Human-Like Outputs. Applications Include Content Creation And Design. Generative AI Enhances Productivity. Businesses Use It For Innovation And Automation. It Is A Rapidly Growing AI Field. Generative AI Is Transforming The Way People Create And Interact With Digital Content.
92. What Is ChatGPT?
Ans:
ChatGPT Is A Conversational AI Model Designed For Human Interaction. It Understands And Generates Natural Language. The Model Assists With Questions And Content Creation. It Uses Advanced Language Learning Techniques. Businesses Employ It For Customer Support. ChatGPT Improves Productivity And Accessibility. It Demonstrates The Power Of Generative AI.
93. What Is A Large Language Model?
Ans:
A Large Language Model Is Trained On Massive Text Datasets. It Learns Grammar, Context, And Knowledge Patterns. These Models Generate Human-Like Responses. Applications Include Chatbots And Content Generation. Large Models Require Significant Computing Resources. They Support Various Natural Language Tasks. LLMs Are Transforming AI Applications Worldwide.
94. What Is Transformer Architecture?
Ans:
- Transformer Architecture Uses Attention Mechanisms For Learning. It Processes Data More Efficiently Than Traditional RNNs.
- Transformers Handle Long-Range Dependencies Well. They Form The Foundation Of Modern Language Models.
- Training Can Be Performed In Parallel. The Architecture Delivers Excellent Performance. It Revolutionized Natural Language Processing.
95. What Is Attention Mechanism?
Ans:
Attention Mechanism Helps Models Focus On Relevant Information. It Assigns Importance To Different Inputs. Context Understanding Improves Significantly. Attention Is Central To Transformer Models. The Technique Enhances Language Processing Accuracy. It Supports Better Learning Efficiency. Modern AI Systems Depend Heavily On Attention.
96. What Is Fine-Tuning?
Ans:
Fine-Tuning Adapts A Pretrained Model To Specific Tasks. Existing Knowledge Is Reused Efficiently. Less Training Data Is Often Required. Fine-Tuning Improves Domain-Specific Performance. It Saves Time And Computational Resources. The Process Is Common In NLP And Vision Applications. Fine-Tuning Accelerates AI Development. Fine-Tuning Enables Models To Deliver Better Results For Specialized Use Cases.
97. What Is Prompt Engineering?
Ans:
Prompt Engineering Designs Effective Inputs For AI Models. Well-Structured Prompts Improve Response Quality. It Helps Achieve Desired Outputs Consistently. Prompt Design Is Important For Generative AI. Different Instructions Produce Different Results. Experimentation Often Improves Performance. Prompt Engineering Is A Valuable AI Skill. Effective Prompt Engineering Maximizes The Accuracy And Relevance Of AI Responses..
98. What Is AI Ethics?
Ans:
AI Ethics Focuses On Responsible AI Development And Use. It Addresses Fairness And Transparency Concerns. Privacy Protection Is A Key Consideration. Ethical AI Reduces Risks And Bias. Organizations Must Follow Responsible Practices. Regulations Often Influence AI Governance. Ethics Ensures Trustworthy Artificial Intelligence Systems. AI Ethics Promotes Safe, Fair, And Accountable Use Of Artificial Intelligence Technologies.
99. What Are Applications Of Machine Learning?
Ans:
Machine Learning Has Applications Across Multiple Industries. Healthcare Uses It For Disease Prediction. Finance Employs It For Fraud Detection. Retail Benefits From Recommendation Systems. Manufacturing Uses Predictive Maintenance Solutions. Transportation Supports Route Optimization. Machine Learning Continues To Transform Businesses Worldwide.
100. Why Does Want To Work On Machine Learning Projects At Infosys?
Ans:
- Machine Learning Projects At Infosys Offer Valuable Learning Opportunities. They Involve Solving Real Business Challenges Using AI.
- Exposure To Advanced Technologies Enhances Professional Growth. Collaboration With Skilled Teams Improves Knowledge.
- Infosys Encourages Innovation And Continuous Learning. The Experience Helps Build Strong Technical Expertise. Contributing To Digital Transformation Makes The Role Highly Rewarding.
LMS
