Genpact Data Analyst Interviews Assess A Candidate’s Knowledge Of Data Analysis, SQL, Excel, Statistics, Data Visualization, And Business Intelligence Tools. Freshers May Be Asked Questions Related To Data Cleaning, Data Transformation, Database Concepts, Power BI, Python, And Analytical Problem-Solving. The Interview Process Can Also Include Practical Scenarios That Test The Ability To Interpret Data, Identify Trends, Create Reports, And Communicate Business Insights Clearly. Preparing Common Technical, Analytical, And Scenario-Based Questions Helps Candidates Build Confidence And Demonstrate Strong Data Analyst Skills. This Guide Covers Important Genpact Data Analyst Interview Questions And Answers To Support Effective Interview Preparation And Career Success.
1. What Is Data Analysis?
Ans:
Data Analysis Is The Process Of Examining, Cleaning, Transforming, And Interpreting Data To Find Useful Insights. It Helps Organizations Understand Trends, Patterns, Relationships, And Business Performance. Data Analysts Use Tools Such As Excel, SQL, Python, And Power BI To Analyze Information. The Results Help Businesses Make Better Data-Driven Decisions.
2. What Does A Data Analyst Do?
Ans:
A Data Analyst Collects, Cleans, Processes, And Analyzes Data To Support Business Decisions. The Role Often Includes Creating Reports, Dashboards, Visualizations, And Performance Metrics. Data Analysts Work With Business Teams To Understand Requirements And Identify Useful Insights. They Also Communicate Findings Clearly To Technical And Non-Technical Stakeholders.
3. What Is The Difference Between Data Analysis And Data Science?
Ans:
| Aspect | Data Analysis | Data Science |
|---|---|---|
| Focus | Focuses On Examining Existing Data To Find Insights, Trends, And Patterns. | Focuses On Extracting Insights And Building Predictive Or Advanced Data-Driven Solutions. |
| Techniques | Commonly Uses SQL, Excel, Statistics, And Data Visualization. | Uses Statistics, Machine Learning, Deep Learning, Python, And Advanced Analytics. |
| Purpose | Helps Understand Past And Current Business Performance. | Helps Predict Future Outcomes And Solve Complex Business Problems. |
| Output | Reports, Dashboards, KPIs, And Business Insights. | Predictive Models, Machine Learning Solutions, And Advanced Analytical Systems. |
4. What Is Structured Data?
Ans:
- Structured Data Is Information Organized In A Clearly Defined Format Such As Rows And Columns.
- Relational Databases And Spreadsheet Tables Are Common Examples Of Structured Data. Each Column Usually Represents A Specific Attribute, While Each Row Represents A Record.
- Structured Data Is Easy To Query, Filter, Sort, And Analyze Using SQL Or Other Tools.
5. What Is Unstructured Data?
Ans:
Unstructured Data Does Not Follow A Fixed Tabular Structure And Can Include Text, Images, Audio, Video, And Documents. Examples Include Customer Emails, Social Media Content, Call Recordings, And Business Documents. Special Processing Techniques Are Often Required To Extract Useful Information From Such Data. Natural Language Processing And Computer Vision Can Help Analyze Certain Types Of Unstructured Data
6. What Is Data Cleaning?
Ans:
Data Cleaning Is The Process Of Identifying And Correcting Errors Or Inconsistencies In A Dataset. It Can Include Handling Missing Values, Removing Duplicates, Correcting Incorrect Formats, And Managing Outliers. Clean Data Improves The Reliability Of Analysis And Business Reports. Data Cleaning Is Usually An Important Step Before Performing Statistical Or Analytical Operations.
7. What Are Missing Values?
Ans:
Missing Values Occur When Information For One Or More Fields Is Not Available In A Dataset. They Can Result From Data Entry Problems, System Failures, Optional Fields, Or Integration Issues. Analysts Can Handle Them By Removing Records, Replacing Values, Or Using Appropriate Imputation Methods. The Correct Approach Depends On The Amount And Importance Of The Missing Data.
8. What Is Data Validation?
Ans:
Data Validation Is The Process Of Checking Whether Data Meets Defined Accuracy, Format, And Business Rules. It Can Verify Data Types, Ranges, Required Fields, Relationships, And Other Conditions. Validation Helps Prevent Incorrect Information From Entering Reports Or Analytical Systems. It Is Important For Maintaining Data Quality And Reliable Business Decisions..
9. What Is Data Transformation?
Ans:
- Data Transformation Converts Data From One Structure, Format, Or Representation Into Another Useful Form.
- It Can Include Aggregation, Filtering, Formatting, Encoding, Scaling, And Creating New Calculated Fields.
- Transformation Makes Data More Suitable For Analysis, Reporting, Or Modeling. SQL, Excel, Python, And ETL Tools Are Commonly Used For Data Transformation.
10. What Is Exploratory Data Analysis?
Ans:
- Exploratory Data Analysis Is The Process Of Investigating A Dataset To Understand Its Main Characteristics And Patterns.
- Analysts Use Summary Statistics, Visualizations, Distributions, And Relationships Between Variables During EDA.
- It Helps Identify Missing Values, Outliers, Trends, And Unexpected Patterns. EDA Provides A Strong Foundation For Further Analysis And Decision-Making.
11. What Is SQL?
Ans:
SQL Stands For Structured Query Language And Is Used To Manage And Analyze Data Stored In Relational Databases. Analysts Use SQL To Retrieve, Filter, Join, Aggregate, Insert, Update, And Delete Data. Common SQL Commands Include SELECT, WHERE, GROUP BY, ORDER BY, JOIN, And HAVING. SQL Is One Of The Most Important Skills For A Data Analyst.
12. What Is A Primary Key?
Ans:
A Primary Key Is A Column Or Combination Of Columns That Uniquely Identifies Each Record In A Database Table. A Primary Key Cannot Normally Contain Duplicate Values And Is Used To Maintain Record Uniqueness. It Helps Establish Relationships Between Tables Through Related Keys. Examples Include Customer ID, Employee ID, Or Product ID.
13. What Is A Foreign Key?
Ans:
A Foreign Key Is A Column That References A Primary Key Or Unique Key In Another Table. It Helps Establish Relationships Between Tables In A Relational Database. Foreign Keys Improve Data Consistency By Connecting Related Records Across Tables. For Example, An Order Table Can Use Customer ID As A Foreign Key To Reference A Customer Table..
14. What Is A SQL JOIN?
Ans:
A SQL JOIN Combines Data From Two Or More Tables Based On A Related Column. Common JOIN Types Include INNER JOIN, LEFT JOIN, RIGHT JOIN, And FULL OUTER JOIN. JOINs Allow Analysts To Combine Information Stored Across Different Database Tables. They Are Frequently Used To Create Complete Datasets For Business Analysis And Reporting.
15. What Is The Difference Between INNER JOIN And LEFT JOIN?
Ans:
An INNER JOIN Returns Only Records That Have Matching Values In Both Joined Tables. A LEFT JOIN Returns All Records From The Left Table And Matching Records From The Right Table. When No Match Exists, The Right-Side Columns Usually Contain NULL Values. The Choice Depends On Whether Unmatched Records Need To Be Included In The Analysis.
16. What Is GROUP BY In SQL?
Ans:
GROUP BY Is Used To Organize Rows With Similar Values Into Groups For Aggregation. It Is Commonly Used With Functions Such As COUNT, SUM, AVG, MIN, And MAX. For Example, Sales Data Can Be Grouped By Region To Calculate Total Sales For Each Region. GROUP BY Is Essential For Creating Business Summaries And Performance Reports.
17. What Is HAVING In SQL?
Ans:
HAVING Is Used To Filter Groups Created By The GROUP BY Clause. Unlike WHERE, It Is Commonly Applied After Aggregation Has Been Performed. For Example, HAVING Can Be Used To Find Departments With Total Sales Above A Specific Amount. It Is Useful When Filtering Based On Aggregate Results.
18. What Is The Difference Between WHERE And HAVING
Ans:
- WHERE Filters Individual Rows Before Grouping And Aggregation Are Performed. HAVING Filters Groups After Aggregate Functions Have Been Applied.
- WHERE Is Commonly Used For Conditions On Individual Records, While HAVING Is Used For Aggregate Conditions.
- Understanding This Difference Helps Analysts Write Accurate SQL Queries.
19. What Is A Subquery?
Ans:
A Subquery Is A Query Written Inside Another SQL Query. It Can Be Used To Calculate Intermediate Results Or Filter Data Based On Another Query. Subqueries Can Appear In SELECT, FROM, Or WHERE Clauses Depending On The Requirement. They Are Useful For Solving Complex Data Retrieval And Analysis Problems.
20. What Is A CTE?
Ans:
A CTE Stands For Common Table Expression And Provides A Temporary Named Result Set Within A SQL Query. It Is Created Using The WITH Keyword And Can Make Complex Queries Easier To Read. CTEs Are Useful For Breaking Large Queries Into Smaller Logical Sections. They Can Also Support Recursive Queries In Appropriate Database Systems.
21. What Is A Window Function?
Ans:
AA Window Function Performs Calculations Across Related Rows Without Combining Them Into A Single Result Row. Common Window Functions Include ROW_NUMBER, RANK, DENSE_RANK, LAG, LEAD, And SUM. They Are Useful For Ranking, Running Totals, Comparisons, And Trend Analysis. Window Functions Are Frequently Used In Advanced Data Analyst SQL Interviews.
22. What Is Normalization?
Ans:
- Normalization Is The Process Of Organizing Database Data To Reduce Redundancy And Improve Data Consistency.
- It Divides Data Into Related Tables Based On Logical Relationships. Normalization Helps Avoid Duplicate Information And Update Anomalies.
- Relational Database Design Commonly Uses Normalization To Maintain Structured And Reliable Data.
23. What Is Denormalization?
Ans:
Denormalization Is The Process Of Intentionally Combining Or Duplicating Data To Improve Query Performance Or Simplify Reporting. It Can Reduce The Number Of JOIN Operations Required During Data Retrieval. However, It May Increase Storage Requirements And Data Maintenance Complexity. Denormalization Is Often Used In Analytical Systems Where Read Performance Is Important.
24. What Is Excel Used For In Data Analysis?
Ans:
- Excel Is Widely Used For Data Cleaning, Calculations, Reporting, Visualization, And Quick Business Analysis.
- Analysts Can Use Formulas, Pivot Tables, Charts, Filters, And Conditional Formatting To Understand Data.
- Functions Such As VLOOKUP, XLOOKUP, IF, SUMIF, And COUNTIF Are Commonly Used. Excel Is Particularly Useful For Small And Medium-Sized Analytical Tasks.
25. What Is A Pivot Table?
Ans:
A Pivot Table Is An Excel Feature Used To Summarize And Analyze Large Amounts Of Data. It Can Group Records And Calculate Values Such As Sum, Count, Average, Minimum, And Maximum. Analysts Can Quickly Change Rows, Columns, Filters, And Values To Explore Different Views. Pivot Tables Are Useful For Creating Business Summaries Without Writing Complex Formulas
26. What Is VLOOKUP?
Ans:
VLOOKUP Is An Excel Function Used To Search For A Value In The First Column Of A Table And Return Related Information From Another Column. It Is Commonly Used To Combine Information From Different Tables Or Worksheets. VLOOKUP Requires A Lookup Value, Table Range, Column Number, And Matching Option. XLOOKUP Is A More Flexible Modern Alternative In Supported Excel Versions.
27. What Is XLOOKUP?
Ans:
XLOOKUP Is An Excel Function Used To Search For A Value And Return A Corresponding Result From Another Range. It Provides More Flexibility Than Traditional VLOOKUP Because The Lookup And Return Ranges Can Be Independently Specified. XLOOKUP Can Search In Different Directions And Handle Missing Matches More Easily. It Is Useful For Data Matching And Record Retrieval.
28. What Is Conditional Formatting?
Ans:
Conditional Formatting Automatically Changes The Appearance Of Cells Based On Defined Conditions. Analysts Can Use It To Highlight High Values, Low Values, Duplicates, Errors, Or Specific Business Conditions. It Helps Users Identify Important Patterns Quickly In Large Worksheets. Conditional Formatting Is Commonly Used In Operational Reports And Dashboards.
29. What Is Power BI?
Ans:
Power BI Is A Business Intelligence And Data Visualization Platform Used To Connect, Transform, Analyze, And Present Data. It Allows Analysts To Create Interactive Reports, Dashboards, Charts, And Business Metrics. Power BI Can Connect To Databases, Excel Files, Cloud Sources, And Other Data Systems. It Helps Organizations Monitor Performance And Make Data-Driven Decisions.
30. What Is A Power BI Dashboard?
Ans:
A Power BI Dashboard Is A Visual Interface That Displays Important Business Metrics And Insights In An Easily Understandable Format. It Can Contain Charts, Cards, Graphs, Maps, Tables, And Interactive Filters. Dashboards Help Stakeholders Monitor Key Performance Indicators And Identify Trends. A Good Dashboard Should Be Clear, Relevant, Accurate, And Focused On Business Objectives..
31. What Is DAX?
Ans:
- DAX Stands For Data Analysis Expressions And Is A Formula Language Used In Power BI And Related Microsoft Data Tools.
- It Can Create Measures, Calculated Columns, And Advanced Analytical Calculations. DAX Supports Functions For Aggregation, Filtering, Time Intelligence, And Logical Operations.
- Understanding DAX Helps Analysts Build More Powerful And Interactive Power BI Reports.
32. What Is ETL?
Ans:
ETL Stands For Extract, Transform, And Load And Is A Process Used To Move Data From Source Systems Into A Target Data Store. Extraction Collects Data From Sources, Transformation Cleans And Converts It, And Loading Places It Into The Target System. ETL Helps Prepare Data For Reporting, Analytics, And Business Intelligence. Data Analysts May Work With ETL Pipelines To Ensure Reliable Analytical Data.
33. What Is Data Warehousing?
Ans:
A Data Warehouse Is A Centralized System Designed To Store Large Amounts Of Data For Reporting And Analysis. It Usually Combines Information From Multiple Operational And External Sources. Data Warehouses Support Historical Analysis, Aggregation, Business Intelligence, And Decision-Making. Analysts Use SQL And BI Tools To Query Data Stored In Warehouses.
34. What Is The Difference Between OLTP And OLAP?
Ans:
| Aspect | OLTP | OLAP |
|---|---|---|
| Purpose | Used For Managing And Processing Day-To-Day Transactions. | Used For Data Analysis, Reporting, And Decision-Making. |
| Data | Usually Stores Current And Detailed Transaction Data. | Usually Stores Historical And Aggregated Data. |
| Queries | Handles Frequent Insert, Update, And Delete Operations. | Handles Complex Queries, Aggregations, And Analytical Operations.s |
| Example | Banking Transactions, Order Processing, And Online Payment | Sales Analysis, Business Reports, Forecasting, And Trend Analysis. |
35. What Is Descriptive Analytics?
Ans:
Descriptive Analytics Focuses On Understanding What Has Already Happened In A Dataset Or Business Process. It Uses Historical Data, Reports, Dashboards, And Summary Statistics To Describe Performance. Examples Include Monthly Sales Reports, Customer Counts, And Revenue Trends. It Provides A Foundation For More Advanced Diagnostic And Predictive Analysis.
36. What Is Diagnostic Analytics?
Ans:
Diagnostic Analytics Attempts To Understand Why A Particular Event Or Result Occurred. Analysts Examine Relationships, Trends, Segments, And Contributing Factors To Identify Possible Causes. Techniques Can Include Drill-Down Analysis, Correlation Analysis, Comparisons, And Root Cause Investigation. It Helps Businesses Move From Knowing What Happened To Understanding Why It Happened.
37. What Is Predictive Analytics?
Ans:
Predictive Analytics Uses Historical Data, Statistical Methods, And Machine Learning To Estimate Future Outcomes. Examples Include Sales Forecasting, Customer Churn Prediction, And Demand Forecasting. Predictions Are Based On Patterns Found In Available Data And Always Contain Some Uncertainty. Data Analysts May Support Predictive Projects By Preparing Data And Evaluating Results.
38. What Is Prescriptive Analytics?
Ans:
Prescriptive Analytics Goes Beyond Prediction By Suggesting Actions That May Lead To Better Outcomes. It Can Use Optimization, Simulation, Rules, And Predictive Models To Recommend Decisions. For Example, It Can Suggest Inventory Levels Or Resource Allocation Strategies. Prescriptive Analytics Helps Organizations Decide What Actions Should Be Taken Based On Data.
39. What Is Mean?
Ans:
- Mean Is A Measure Of Central Tendency Calculated By Adding All Values And Dividing The Total By The Number Of Values.
- It Provides A Simple Representation Of The Average Value In A Dataset. Mean Can Be Strongly Influenced By Extreme Values Or Outliers.
- It Is Commonly Used In Statistical Analysis And Business Reporting.
40. What Is Median?
Ans:
- Median Is The Middle Value When Data Points Are Arranged In Ascending Or Descending Order.
- If The Dataset Contains An Even Number Of Values, The Median Is Usually The Average Of The Two Middle Values.
- Median Is Less Sensitive To Extreme Values Than Mean. It Is Useful For Analyzing Skewed Data Such As Income Or Transaction Values.
41. What Is Mode?
Ans:
Mode Is The Value That Appears Most Frequently In A Dataset. A Dataset Can Have One Mode, Multiple Modes, Or No Mode If Values Occur With Equal Frequency. Mode Can Be Used For Both Numerical And Categorical Data. It Is Particularly Useful For Identifying The Most Common Category Or Value.
42. What Is Standard Deviation?
Ans:
Standard Deviation Measures How Much Individual Data Values Differ From The Mean. A Low Standard Deviation Indicates That Values Are Relatively Close To The Average. A High Standard Deviation Indicates Greater Variation In The Dataset. It Is Widely Used To Understand Data Spread And Statistical Risk.
43. What Is Variance?
Ans:
Variance Measures The Average Squared Difference Between Data Values And Their Mean. It Provides An Indication Of How Widely Values Are Distributed Around The Average. A Larger Variance Indicates Greater Data Spread, While A Smaller Variance Indicates Less Spread. Standard Deviation Is The Square Root Of Variance And Is Easier To Interpret In The Original Units.
44. What Is Correlation?
Ans:
- Correlation Measures The Strength And Direction Of The Relationship Between Two Variables.
- Its Value Generally Ranges From -1 To +1, Where Values Near +1 Indicate Strong Positive Association And Values Near -1 Indicate Strong Negative Association.
- A Value Near Zero Indicates Little Or No Linear Relationship. Correlation Does Not Automatically Mean That One Variable Causes Another.
45. What Is Regression Analysis?
Ans:
Regression Analysis Is A Statistical Technique Used To Study Relationships Between A Dependent Variable And One Or More Independent Variables. It Can Be Used To Understand Relationships And Make Predictions About Numerical Outcomes. Linear Regression Is A Common Example Used In Data Analysis. Regression Results Can Help Identify Important Factors Affecting A Business Metric.
46. What Is An Outlier?
Ans:
An Outlier Is A Data Point That Is Significantly Different From Most Other Observations In A Dataset. Outliers Can Result From Data Entry Errors, Measurement Problems, Fraud, Or Genuine Unusual Events. Analysts Can Detect Them Using Statistical Methods, Visualizations, Or Domain-Specific Rules. Outliers Should Be Investigated Before Deciding Whether To Remove Or Retain Them..
47. What Is Data Visualization?
Ans:
Data Visualization Is The Process Of Representing Data Using Charts, Graphs, Maps, And Other Visual Elements. It Helps Analysts And Business Users Understand Trends, Comparisons, Relationships, And Patterns More Easily. Common Visualizations Include Bar Charts, Line Charts, Pie Charts, Scatter Plots, And Heatmaps. Effective Visualization Should Clearly Communicate The Most Important Business Insight.
48. What Is A KPI?
Ans:
KPI Stands For Key Performance Indicator And Represents A Metric Used To Measure Progress Toward A Business Objective. Examples Include Revenue Growth, Customer Retention, Conversion Rate, And Average Resolution Time. KPIs Should Be Relevant, Measurable, And Connected To Business Goals. Data Analysts Often Build Reports And Dashboards To Monitor KPIs.
49. What Is A Dashboard?
Ans:
The Second Highest Salary Can Be Found Using DENSE_RANK(), A Dashboard Is A Visual Collection Of Important Metrics And Charts Designed To Provide A Quick View Of Business Performance. It Can Combine Data From Multiple Sources And Present Information Through Interactive Filters And Visualizations. Dashboards Help Users Monitor Trends And Identify Problems Quickly. Effective Dashboards Focus On Relevant Metrics Rather Than Presenting Excessive Information.
50. How Does Choose The Right Chart?
Ans:
The Right Chart Depends On The Type Of Data And The Business Question Being Answered. Bar Charts Are Useful For Comparisons, Line Charts For Trends, And Scatter Plots For Relationships Between Numerical Variables. Maps Can Be Useful For Geographic Analysis, While Tables Can Show Detailed Values. The Main Goal Should Be To Present The Insight Clearly And Avoid Misleading Visuals.
51. What Is Python Used For In Data Analysis?
Ans:
Python Is Widely Used For Data Cleaning, Statistical Analysis, Automation, Visualization, And Machine Learning. Libraries Such As Pandas, NumPy, Matplotlib, And Scikit-Learn Provide Tools For Different Analytical Tasks. Python Can Handle Repetitive Data Processing More Efficiently Than Manual Spreadsheet Operations. It Is Particularly Useful For Large Or Complex Analytical Workflows.
52. What Is Pandas?
Ans:
Pandas Is A Python Library Used For Data Manipulation And Analysis. It Provides Data Structures Such As DataFrame And Series For Working With Tabular Data. Analysts Can Use Pandas To Filter Records, Handle Missing Values, Group Data, Merge Tables, And Transform Columns. Pandas Is One Of The Most Common Python Libraries Used By Data Analysts.
53. What Is NumPy?
Ans:
Denormalization Is The Intentional Combination Or Duplication Of NumPy Is A Python Library Designed For Numerical Computing And Efficient Array Operations. It Provides Multidimensional Arrays And Mathematical Functions For Processing Numerical Data. NumPy Forms A Foundation For Many Other Python Data Science Libraries. It Is Commonly Used For Mathematical Calculations, Statistical Operations, And Data Processing.
54. What Is Matplotlib?
Ans:
Matplotlib Is A Python Visualization Library Used To Create Charts And Graphs From Data. It Supports Common Visualizations Such As Line Charts, Bar Charts, Scatter Plots, Histograms, And Box Plots. Analysts Can Customize Titles, Labels, Axes, And Other Chart Elements. Matplotlib Is Useful For Exploratory Analysis And Presenting Analytical Results.
55. What Is A Histogram?
Ans:
A Histogram Is A Chart Used To Show The Distribution Of Numerical Data Across Different Ranges. It Divides Values Into Intervals Called Bins And Displays The Frequency Within Each Bin. Histograms Help Analysts Understand Distribution Shape, Spread, Skewness, And Potential Outliers. They Are Commonly Used During Exploratory Data Analysis.
56. What Is A Scatter Plot?
Ans:
- A Scatter Plot Displays Individual Data Points Based On Two Numerical Variables. It Helps Analysts Identify Relationships, Trends, Clusters, And Potential Outliers Between Variables.
- A Positive Pattern May Indicate That Both Variables Increase Together, While A Negative Pattern Shows The Opposite.
- Scatter Plots Are Frequently Used For Correlation And Regression Analysis.
57. What Is A Box Plot?
Ans:
A Box Plot Summarizes The Distribution Of Numerical Data Using Quartiles And Other Statistical Measures. It Can Show The Median, Spread, And Potential Outliers In A Compact Format. Box Plots Are Useful For Comparing Distributions Across Multiple Groups. They Are Particularly Helpful When Analysts Need To Identify Differences In Data Spread.
58. What Is Sampling?
Ans:
Sampling Is The Process Of Selecting A Smaller Group Of Observations From A Larger Population. It Allows Analysts To Study Data More Efficiently When Examining The Entire Population Is Expensive Or Impractical. Common Methods Include Random, Stratified, Systematic, And Cluster Sampling. A Good Sample Should Represent The Population As Closely As Possible.
59. What Is A/B Testing?
Ans:
A/B Testing Is An Experiment That Compares Two Versions Of A Product, Feature, Page, Or Process. Users Are Typically Divided Into Groups And Exposed To Different Versions. Their Outcomes Are Measured Using A Defined Metric Such As Conversion Rate Or Engagement. Statistical Analysis Helps Determine Whether The Observed Difference Is Meaningful.
60. What Is Hypothesis Testing?
Ans:
Hypothesis Testing Is A Statistical Method Used To Determine Whether Evidence Supports A Specific Claim About A Population. It Usually Involves A Null Hypothesis And An Alternative Hypothesis. Analysts Calculate A Test Statistic And Often Use A P-Value To Evaluate The Evidence. Hypothesis Testing Helps Support Data-Driven Conclusions About Differences Or Relationships.
61. What Is A P-Value?
Ans:
A P-Value Indicates How Consistent The Observed Data Is With The Null Hypothesis Under A Chosen Statistical Model. A Small P-Value Can Provide Evidence Against The Null Hypothesis. However, It Does Not Directly Measure The Size Or Practical Importance Of An Effect. Analysts Should Consider Effect Size, Confidence Intervals, And Business Context Along With The P-Value.
62. What Is A Confidence Interval?
Ans:
A Confidence Interval Provides A Range Of Plausible Values For A Population Parameter Based On Sample Data And A Chosen Confidence Level. It Helps Express The Uncertainty Associated With An Estimate. Wider Intervals Generally Indicate Greater Uncertainty Than Narrower Intervals. Confidence Intervals Are Useful For Interpreting Means, Proportions, Model Parameters, And Business Metrics.
63. What Is Data Quality?
Ans:
- Data Quality Refers To How Accurate, Complete, Consistent, Timely, And Reliable Data Is For Its Intended Purpose.
- Poor Data Quality Can Lead To Incorrect Reports, Wrong Decisions, And Operational Problems
- . Analysts Can Improve Quality Through Validation, Cleaning, Standardization, And Monitoring. Data Quality Should Be Evaluated According To Business Requirements.
64. What Is Data Deduplication?
Ans:
Data Deduplication Is The Process Of Identifying And Removing Duplicate Records From A Dataset. Duplicate Data Can Occur Due To Repeated Entries, System Integration Problems, Or Data Migration Issues. Analysts Can Use Unique Identifiers, Matching Rules, And SQL Queries To Detect Duplicates. Removing Unwanted Duplicates Helps Improve Accuracy In Reports And Analysis.
65. How Does Handle Duplicate Records In SQL?
Ans:
Duplicate Records Can Be Identified Using GROUP BY, COUNT, Window Functions, Or Unique Business Keys. Analysts Can Use ROW_NUMBER To Assign A Sequence To Duplicate Records And Retain Only The Required Row. Before Deleting Data, The Business Rules For Identifying A Valid Record Should Be Confirmed. This Prevents Accidental Removal Of Legitimate Repeated Transactions.
66. What Is A CASE Statement In SQL?
Ans:
A CASE Statement Allows Analysts To Apply Conditional Logic Inside A SQL Query. It Can Categorize Records, create calculated fields, Or assign different values based on specified conditions. For Example, Sales Amounts Can Be Classified Into Low, Medium, And High Categories. CASE Statements Are Commonly Used In Reports, Transformations, And Business Rule Implementation.
67. What Is NULL In SQL?
Ans:
NULL Represents A Missing, Unknown, Or Undefined Value In A Database. It Is Different From Zero, An Empty String, Or A Blank Value. SQL Provides Functions And Operators Such As IS NULL, IS NOT NULL, COALESCE, And NULLIF For Handling NULL Values. Analysts Must Handle NULL Carefully Because It Can Affect Calculations And Query Results.
68. What Is COALESCE In SQL?
Ans:
COALESCE Returns The First Non-NULL Value From A List Of Expressions. It Is Commonly Used To Replace Missing Values With A Default Or Alternative Value. For Example, An Analyst Can Display Zero When A Numerical Field Is NULL. COALESCE Is Useful For Data Cleaning, Reporting, And Building Reliable Calculated Fields.
69. What Is The Difference Between UNION And UNION ALL??
Ans:
UNION Combines Results From Multiple Queries And Removes Duplicate Rows From The Final Result. UNION ALL Combines Results Without Removing Duplicates, Which Can Make It Faster. Both Queries Generally Require Compatible Column Counts And Data Types. The Choice Depends On Whether Duplicate Records Should Be Preserved Or eliminated.
70. What Is The Difference Between DELETE, TRUNCATE, And DROP
Ans:
DELETE Removes Selected Rows From A Table And Can Usually Use A WHERE Condition. TRUNCATE Removes All Rows From A Table More Directly And Is Generally Used When The Entire Table Data Must Be Cleared. DROP Removes The Table Structure Along With Its Data. These Commands Have Different Effects And Should Be Used Carefully In Production Systems.
71. What Is A Database Index?
Ans:
A Database Index Is A Data Structure That Helps The Database Find Rows More Efficiently. Indexes Can Improve Query Performance When Columns Are Frequently Used For Searching, Joining, Or Sorting. However, Indexes Require Additional Storage And Can Increase The Cost Of Insert Or Update Operations. Proper Index Design Requires Understanding Query Patterns And Data Characteristics.
72. How Does Optimize A Slow SQL Query?
Ans:
A Slow SQL Query Can Be Improved By Reviewing The Execution Plan And Identifying Expensive Operations. Appropriate Indexes, Better JOIN Conditions, Efficient Filtering, And Reduced Unnecessary Columns Can Improve Performance. Analysts Should Avoid Processing More Data Than Required And Review Whether Aggregations Can Be Optimized. Query Optimization Should Always Be Validated Using Actual Performance Measurements.
73. What Is An SQL Execution Plan?
Ans:
An SQL Execution Plan Shows How A Database Engine Intends To Execute A Query. It Can Display Operations Such As Table Scans, Index Usage, Joins, Sorting, And Aggregation. Analysts And Database Professionals Use Execution Plans To Identify Performance Bottlenecks. Understanding The Plan Helps Determine Whether Query Rewriting Or Index Improvements Are Needed.
74. What Is A Star Schema?
Ans:
- A Star Schema Is A Data Warehouse Design That Contains A Central Fact Table Connected To Multiple Dimension Tables.
- The Fact Table Usually Stores Measurable Business Events Such As Sales Or Transactions.
- Dimension Tables Provide Descriptive Information Such As Customer, Product, Location, Or Date. This Structure Is Common In Business Intelligence And Analytical Reporting.
75. What Is A Fact Table?
Ans:
A Fact Table Stores Quantitative Business Events Or Measurements In A Data Warehouse. It Often Contains Foreign Keys That Connect The Facts To Related Dimension Tables. Examples Of Facts Include Sales Amount, Quantity, Cost, Revenue, And Transaction Count. Fact Tables Are Used For Aggregation And Analysis In Business Intelligence Systems.
76. What Is A Dimension Table?
Ans:
A Dimension Table Stores Descriptive Information Used To Provide Context To Business Facts. Examples Include Customer, Product, Employee, Location, And Date Dimensions. Dimension Attributes Help Analysts Filter, Group, And Analyze Measures From Fact Tables. They Are An Important Component Of Star And Snowflake Data Warehouse Schemas.
77. What Is Time Series Analysis?
Ans:
Time Series Analysis Examines Data Collected Over Time To Identify Trends, Seasonality, Cycles, And Other Patterns. Examples Include Daily Sales, Monthly Revenue, Website Traffic, And Customer Transactions. Analysts Can Use Time-Based Aggregations, Moving Averages, And Forecasting Techniques To Study Such Data. Time Series Analysis Helps Organizations Understand Historical Performance And Plan For The Future.
78. What Is Forecasting?
Ans:
Forecasting Is The Process Of Estimating Future Values Based On Historical Patterns And Relevant Information. It Can Be Applied To Sales, Demand, Revenue, Inventory, Website Traffic, And Other Business Metrics. Forecasting Methods Range From Simple Moving Averages To Statistical And Machine Learning Models. The Accuracy Of A Forecast Depends On Data Quality, Model Choice, And Changing Business Conditions..
79. What Is Customer Segmentation?
Ans:
Oracle Can Store And Provide Access To Large Volumes Of Customer Segmentation Is The Process Of Dividing Customers Into Groups Based On Shared Characteristics Or Behaviors. Segments Can Be Created Using Demographics, Purchase Behavior, Engagement, Geography, Or Customer Value. Analysts Can Use SQL, Statistics, And Clustering Techniques To identify Meaningful Groups. Segmentation Helps Businesses Create More Targeted Marketing And Customer Strategies.
80. What Is Customer Churn
Ans:
- Customer Churn Refers To Customers Stopping Their Use Of A Product Or Service During A Defined Period.
- Analysts Study Churn Rates And Customer Behavior To Identify Factors Associated With Customer Loss.
- Common Factors May Include Price, Poor Service, Low Engagement, Or Competitive Offers. Churn Analysis Helps Organizations Develop Strategies To Improve Customer Retention.
81. What Is Conversion Rate?
Ans:
- Conversion Rate Measures The Percentage Of Users Or Visitors Who Complete A Desired Action.
- The Action Can Include Purchasing A Product, Registering For A Service, Submitting A Form, Or Completing Another Business Goal.
- It Is Usually Calculated By Dividing Conversions By The Relevant Total Number Of Users Or Visits
82. How Does Identify Trends In Data?
Ans:
Trends Can Be Identified By Examining Data Over Time Using Line Charts, Aggregations, Moving Averages, And Statistical Techniques. Analysts Compare Current Values With Historical Periods To Detect Growth, Decline, Or Seasonal Patterns. Segmenting Data By Relevant Dimensions Can Reveal Trends That Are Hidden In Overall Results. Business Context Should Always Be Considered Before Drawing Conclusions From A Trend.
83. How Does Handle Conflicting Data From Different Sources
Ans:
Conflicting Data Should First Be Investigated To Understand Differences In Definitions, Timing, Formats, Or Source Systems. Analysts Can Compare Data Quality Rules, Record Counts, Business Logic, And Source Reliability. A Standardized Definition And Trusted Source Should Be Established With Relevant Stakeholders. The Final Data Should Be Validated Before It Is Used In Reports Or Business Decisions
84. How Does Ensure Accuracy In A Report?
Ans:
Report Accuracy Can Be Improved By Validating Source Data, Business Rules, Calculations, Filters, And Aggregations. Analysts Should Reconcile Important Metrics With Trusted Sources And Test Results Across Different Scenarios. Peer Review And Automated Data Quality Checks Can Help Identify Errors. Reports Should Also Clearly Define Metrics So Users Understand What Each Number Represents.
85. How Does Explain Data Insights To Non-Technical Stakeholders?
Ans:
Data Insights Should Be Explained Using Simple Language And A Clear Connection To The Business Objective. Instead Of Focusing Only On Technical Methods, The Analyst Should Explain What Happened, Why It Matters, And What Action May Be Considered. Charts And Examples Can Make Complex Findings Easier To Understand. The Explanation Should Be Concise, Evidence-Based, And Relevant To The Audience.
86. How Would Handle A Last-Minute Data Request?
Ans:
A Last-Minute Request Should First Be Clarified To Understand The Exact Business Question, Required Data, Deadline, And Expected Output. The Analyst Should Prioritize The Most Important Requirements And Check Whether The Necessary Data Is Available. Quick Validation Is Important Even When The Deadline Is Tight To Avoid Delivering Incorrect Results. If Full Analysis Is Not Possible, Clear Communication About Scope And Limitations Is Essential.
87. How Would Investigate An Unexpected Drop In Sales?
Ans:
The Sales Data Would First Be Validated To Confirm That The Drop Is Not Caused By Missing Records Or Reporting Errors. The Decline Would Then Be Compared Across Time Periods, Products, Regions, Channels, And Customer Segments. Related Factors Such As Pricing, Promotions, Inventory, Website Performance, And Customer Behavior Would Also Be Examined. The Analysis Would Focus On Identifying The Most Likely Business Drivers Behind The Decline.
88. How Would Approach A Data Analyst Project At Genpact?
Ans:
The Process Begins By Understanding The Business Problem, Stakeholder Requirements, Available Data, And Expected Outcome. Next, The Data Is Cleaned And Validated Before Performing Exploratory Analysis And Identifying Relevant Patterns. SQL, Excel, Python, Or Power BI Can Be Used Depending On The Nature Of The Task And Required Deliverable. Finally, Clear Insights Are Presented, Results Are Validated, And The Findings Are Explained To Support Effective Business Decisions.
89. Why Should Genpact Hire As A Data Analyst?
Ans:
A Strong Candidate For A Data Analyst Role Should Demonstrate Analytical Thinking, Problem-Solving Ability, And Practical Knowledge Of Data Tools. Skills In SQL, Excel, Python, Data Visualization, And Statistics Can Help Handle Different Analytical Requirements. The Ability To Validate Data And Communicate Insights Clearly Is Also Important For Business-Focused Work. A Willingness To Learn New Technologies And Understand Business Processes Can Further Support Success In The Role.
90. How Should A Fresher Prepare For A Genpact Data Analyst Interview?
Ans:
A Fresher Should Revise SQL, Excel, Statistics, Data Cleaning, Python, Power BI, Data Visualization, And Basic Business Analytics Concepts. Practical Practice Should Include SQL Queries, Excel Functions, Dashboard Creation, Data Interpretation, And Real-Time Business Scenarios. Candidates Should Also Prepare To Explain Academic Projects, Datasets, Analytical Methods, Challenges, And Results Clearly.
LMS

