Tuesday, March 2, 2021

Inside the mind of a Product Manager

Recently read some articles, watched some YouTube videos on the roles and responsibilities of a Product Manager. I found it very fascinating to know and understand the mindset of a Product Manager. In a layman's term - he/she is a glorified Business Analyst with a Developer skillset with marketing acumen operating within the realms of legal framework by providing the maximum utility of a Product - or to put it even exalted level - CEO of a product.

A Product Manager is responsible for owning the product, have a vision of that product, and think of creative ways to work with developers, the marketing team, and any other peripheral teams (legal, security, design, communication, etc.) to help shape that vision to fruition.
Also, I feel that a Product Manager needs to be in touch with ground realities time and again. Working closely with a Product cab gives them tunnel vision and they feel very protective towards their creation.

As a Consultant I can connect in many ways a Product Manager navigates different hoops of obstacles and work for the greater good.

I have worked with many different teams, different team dynamics coupled with different technologies and domains in play. I have my eyes set on the prize - the delivery of the project which makes one immune to internal politics or any disruptions. Resolve to get things done, grit to endure hierarchical nomenclature, and navigate the organizational terrain. 

The focus should be on the delivery (launch) as well as sailing (land) of a Project.

Validation is one of the key indicators if processes are flowing smoothly and not having any bottlenecks or data integrity compromised. Having checkpoints in form of Reports to verify - data coming in - data processed - output result. In short, the landing of a space shuttle is equally important if not more than the launching of a space shuttle.





Supervised Learning : KNN Algorithm

K-Nearest Neighbors (KNN) Algorithm*

Motivation

The k-nearest neighbor algorithm, commonly known as the KNN algorithm, is a simple yet effective classification and regression supervised machine learning algorithm.

Classification is more intuitive to human learning. Since we were kids, our parents made sure we recognize objects , letters , people correctly. This method is formally known as supervised learning. In terms of notation, supervised learning is a function that maps an input to an output based on examples, input-output pairs

As a parent myself, when I tell my then 2-year old that a round object that bounces is a ball. ‘Round’ and ‘Bounce’ are features (input) and ‘Ball’ is variable (output)
This entire process can be called model training. Similarly, when the independent variable/output is a constant value, it is known as a regression problem.

What is the K-Nearest Neighbors (KNN) Algorithm?

The KNN algorithm is a classical machine learning algorithm that focuses on the distance from new unclassified/unlabeled data points to existing classified/labeled data points.

Figure 1: Colors of the three flower categories in our example dataset, red, yellow and green.

We already have labeled data points from our dataset. These data points belong to three categories represented by red, green, and yellow colors as shown above

Along comes a new ‘alien’ data point represented by the black cross mark and we need to determine which category this new data point belongs to from the three colors.

First, we assign ‘k’ a random value. The k-value tells the no. of nearest points to look for from the new unlabelled data point. Next, we calculate the distance from the unlabelled data point to every data point on the graph and select the top 5 shortest distances.

Figure 2: Measuring the distance between the categories from our dataset.
Amongst the top 5 nearest data points, It is evident that the new data point belongs to category red as most of its nearest neighbors are from category red.

Figure 3: Seeing how the new data point belongs to the category red from our example dataset.

Similarly, in a regression problem, the aim is to predict a new data point’s value. Based on a x-y graph, there are 2 features for each data point. The x-axis represents feature-1, and the y-axis represents feature-2.

We introduce a new data point for which only feature-1 value is known, and we need to predict feature-2 value. We start with a k-value of 5 and get the 5 nearest neighboring points from the new data point. The predicted value of feature-2 for the new data point is the mean of the feature-2 of 5 nearest neighbors.

Figure 4: Our example of introducing a new data point to illustrate the k-nearest neighbor (KNN) algorithm.

When to use KNN?

1.      KNN algorithm for applications that require high accuracy

2.      When the dataset is small.

3.      When data labeled correctly, and the predicted value will be among the given labels.

4.      KNN is used to solve regression, classification, or search problems

Pros and Cons of Using KNN

Pros

·        It is straightforward to implement since it requires only 2 param: k-value and distance function.

·        A good value of k will make the algorithm robust to noise.

·        It learns a nonlinear decision boundary.

·        There are almost no assumptions on the given data.

·        It is a non-parametric approach. No model fitting/training is required.

Cons

·        Inefficient for large datasets since distance needs to be calculated.

·        The model is susceptible to outliers.

·        It cannot handle imbalanced data. It needs to be handled explicitly.

·        If dataset requires large K value, it will increase the computation of the algorithm.

Math Behind KNN

T algorithm calculates the distance from the new data point to each existing data point. There are 2 common methods.

Figure 5: A representation of the Minkowski distance.

Minkowski Distance:

a) The Minkowski Distance is a generalized distance function in a normed vector space.

 

b) In cases,

1.      When p=1 — this is the Manhattan Distance

2.      When p=2 — this is the Euclidean Distance

Euclidean distance

the Euclidean distance between two points in the Euclidean space is the line segment’s length between the two points. We use the following formula:Figure 7: Euclidean distance formula.

Manhattan distance

The Manhattan distance calculation is similar to the calculation of Euclidean distance, with the difference being that we take absolute value instead of taking the square root of the sum of squared difference.

Fundamentally, the Euclidean distance represents “flying from one point to another,” and the Manhattan distance is “traveling from one point to another point” in a city following the pathway or the road.Figure 8: Manhattan distance formula. 

The common method is the Euclidean distance formula. One of the reasons being, Euclidean distance can calculate in any dimension, whereas Manhattan finds the elements on a vertical or a horizontal plane.

How to Choose the Right Value for K?

There is no standard statistical method to compute the most optimal K value. We want to choose a K value that will reduce errors. As we increase K, our predictions become more stable due to averaging or majority voting but then the error rate will increase as it will underfit the model. A small K yields a low bias and high variance (higher complexity), while a large K yields a high bias and a low variance. There are few different methods to try: Domain knowledge, Cross-Validation or Square Root

Implementation of KNN in Python

Using Iris dataset of Iris flowers of three related species: Setosa, Versicolor, Virginica. The observed features are sepal length, sepal width, petal length, and petal width.

import numpy as npimport
pandas as pdimport
matplotlib.pyplot as pltimport
seaborn as sns
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KneighborsClassifier
from sklearn.model_selection import cross_val_score
from sklearn.datasets import load_iris
iris = load_iris()

The table below shows the first 5 rows of data in a Pandas DataFrame. The target values are 0.0, 1.0, and 2.0, representing Setosa, Versicolor, and Virginica, respectively.


Since KNN is sensitive to outliers and imbalanced data. Plotting a count plot for the target variable, there are 50 samples of each flower type. Thankfully, the data is perfectly balanced.

sns.countplot(x=’target’, data=iris



for feature in [‘sepal length (cm)’, ‘sepal width (cm)’, ‘petal length (cm)’, ‘petal width (cm)’]:
     sns.boxplot(x=’target’, y=feature, data=iris)
     plt.show()

 


Next, we split the data into training and testing sets to measure how accurate the model is. The model will be trained on the training set, which is randomly selected 60:40 of the original data. Before splitting it into training and testing sets, it is essential to separate the feature/dependent and target/independent variable.

X = iris.drop([‘target’], axis=1)
y = iris[‘target’]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.4, random_state=0)

Building the initial model with a k-value of 1, meaning only 1 nearest neighbor will be considered for classifying a new data point. Internally, the distance from the new data point to all the data points will be calculated and then sorted ascending, along with their respective classes. Since the k-value is 1, the class (target value) of the first instance from the sorted array will determine the new data class. As we can see, we obtain a decent accuracy score of 91.6%. However, the optimal k-value needs to be selected.

knn = KNeighborsClassifier(n_neighbors=1)
knn.fit(X_train, y_train)
print(knn.score(X_test, y_test))
Output: 0.9166666666666666

To find the optimal k-value by cross-validation method, we calculate the k-values’ accuracy ranging from 1 to 26 in this case and choosing the optimal k-value.

k_range = list(range(1,26))
scores = []
for k in k_range:
    knn = KNeighborsClassifier(n_neighbors=k)
    knn.fit(X_train, y_train)
    y_pred = knn.predict(X_test)
    scores.append(metrics.accuracy_score(y_test, y_pred))
    plt.plot(k_range, scores)
    plt.xlabel(‘Value of k’)
    plt.ylabel(‘Accuracy Score’)
    plt.title(‘Accuracy Scores for different values of k’)
    plt.show()

 


Introducing a new unlabeled data point whose class we need to predict is the flower type, which belongs to a category. We will build the model with a k-value of 11.

knn = KNeighborsClassifier(n_neighbors=11)
knn.fit(iris.drop([‘target’], axis=1), iris[‘target’])
X_new = np.array([[1, 2.9, 10, 0.2]])
prediction = knn.predict(X_new)
print(prediction)
if prediction[0] == 0.0:
  print(‘Setosa’)
elif prediction[0] == 1.0:
  print(‘Versicolor’)
else: print(‘Virginica’)

Output: [2.]Virginica

KNN Applications

From forecasting epidemics and the economy to information retrieval ,recommender systems, data compression, and healthcare, the k-nearest neighbors (KNN) algorithm has become fundamental in such applications. KNN is primarily used during regression and classification tasks.

Conclusion

KNN is a highly effective and easy-to-implement supervised machine learning algorithm that can be used for classification and regression problems. The model functions by calculating distances of a selected number of examples, K, nearest to the predicting point.

For a classification problem, the label becomes the majority vote of the nearest K points. For a regression problem, the label becomes the average of the nearest K points.

 

* = Source Credit : TowardsAI.net


Tuesday, December 15, 2020

Chapli Kabab

Chapli Kabab is my favourite Kabab after Bihari Kabab. I had Chapli Kabab recently at my friends place - Sohail Bhai's. I was dropping him something at his door and he being his usual hospitable and gracious - offered me something he had made - I went after dinner but it was Sohail bhai's creation - couldn't say no to it (and I am bit of a glutton anyways lol). He shared the recipe with me. This past weekend when my parents were vistitng me I got excited and decided to excite them by attempting to make Chapli Kabab.



Recipe Process:

- 1/2 kg of minced beef (25% fat)
- 2 tblsp of oil and whisk 2 eggs - fry both sides
- 3 medium sized red onions - finely chopped and remove excess water
- 10-12 desi green chillis - finely chopped
- Mix the green chillis and onions in the minced beef (qeema)
- + 1 tblspn of ginger + garlic paste
- + 2 cup of freshly chopped coriander
- + 1.5 tblsp of crushed coriander seeds
- + 1 tblspn red chilli flakes
- + 1 tsp cumin powder
- + 1 tsp garam masala
- + 1.5 tsp Ajwain crushed (carom seeds)
- + 1.5 tblsp anardana crushed
- + 0.5 tsp black pepper powder
- Mix everything
- + 1 cup of maize flour - mix it
- +1 raw egg - mix it
- + 2 chopped de-seeded tomatoes
- Refrigerate it for at least 30 mins.

- Heat the oil and fry each side 7-8 mins.

My parents and family really enjoyed it and I surprised myself (again) . It was lot of work but it paid off. Alhamdulilah!






Thursday, October 22, 2020

Data Science : Chapter 2

Wanted the title to be "Data Science: Reloaded" or "Data Science 2.0' but "Chapter 2" gives it a growth-mind-set connotation instead of a dramatic (read cheesy lol) title 

I was introduced to this 6-month Data Science certificate program by my ex-colleague Noman Bhai. It was an instant 'yes' as back of my mind I always wanted to reinforce my past learning about Data Science. I completed my Masters in Management Analytics from Queen's Univ and it has been 5 years since then.

This program uses Python which was another driver in plunging into this certification. I am halfway done (almost end of Module 3) and must say it has been a great journey. Learning a lot from Zoom class sessions, met some dynamic members from cohort from different backgrounds. It has given me lot of exposure to the statistical modeling and make me appreciate it even more with the introduction of Python. 

Young and enthusiast faculty. The passion and desire in explaining Data Science and stats concept is evident. Most of them are pursuing Masters or PhDs; delivering gems in professional and respectful manner.  

Presentations complimented with lab sessions to reinforce the concepts. TA always accessible for any questions. Each module has sporadic quizzes. Every module has written (virtual) and practical (take-away) exam - honestly not my strong forte studying for exams (anymore lol).

Nevertheless, I have something to look forward to on weekends. It is painful to get up early on weekends as weekends are the only 'no alarm' sleeps - but not anymore. Hey - I ain't complaining. I did mention I am enjoying them - learning is fun (sometimes) :)

The biggest 'surprise' (as so may call it) was my introduction to Konain Qurban. In hindsight, embarking on this certification seems like an excuse as destiny wanted me to meet him. Younger than me but an enormous wealth of experience, information and wisdom (had to say that part lol).  He is a trail blazer - he is the founder of this institute - Frontier Institute of Technology. A visionary. If a title exists, he is clearly the front runner for "30 under 30" in Pakistan !

On a lighter side, he also introduced me and family to the best Hyderabadi restaurant in Canada! Will go again when you are back in December !

Almost forgot, we also must do a Capstone Project. I have been assigned a group - looking forward to this sub-journey.

On a serious note, I should start my Assignment for my Module 3. One thing to add,  from work front - it is so damn busy as Covid has shamelessly erased the work-home boundaries, end up working much more and feeling burnt-down as well. I can write more but will hold that thought for later post (plus need to stop procrastinating and get on with the assignment! ) 

 

*Edits. Captured my social media 'fame' lol :


 


Tuesday, October 6, 2020

Strong-Link vs Weak-Link + High-Treshold vs Low-Treshold

Interesting concepts I came across in a podcast called Revisionist by Malcom Gladwell  that I wanted to pen down through this post and give examples to explain the terms.

Strong-Link and Weak-Link

In soccer, your best player is not good if the weakest player is also good enough to make a strong-link. 

In basketball, your best player does not necessarily needs to have all good players. You can have a weak-link and yet your star player can shine and lead the team to victory.


Low-Threshold and High-Threshold

Low Threshold - some people don't 'absorb' when communicating with their partners and  'steam off' when they have something to complain about

High-Threshold - some people just build up all the 'not rights' in the relationships and keep a mental note - and when they 'erupt' on an incident - they bring out the complaints dating back from wedding/honeymoon days. Meaning they have high-threshold in the relationship.


Thursday, October 1, 2020

Stuff it, leave no Room in that 'MushRoom' !

Kids love it.My wife loves it. Our friends love it. Its an instant hit ! It falls in appetizer category but special enough for my family to sit together and savor the taste and enjoy the company and the whole experience. Hey as they say - family that eats together , stays together (IA)

I have made this recipe couple of times and now I think I can moralize the experience by this blog - also serves as the reference for the recipe. Shoutout to Emad for the inspiration as I viewed the recipe for the first time watching his video

Process:

1) De-throne the mushroom from its stalk or should I say uproot the stalk from its head and place the heads on the baking pan. (Chop the stalks and put it aside for later)
2) Chop 1 onion
3) Chop 1/2 bell pepper (which ever color is available honestly, but I usually pick the yellow one. Hello, why not?
4) Heat the pan with Olive/Avocado oil with fresh ginger (1/4) and garlic (2 cloves)
5) Now throw in the chopped onions, pepper and mushrooms (yes in that order with 5-6 mins interval)
6) I used Kirkland spice mix - 2 tablespoon. Or you can add salt, pepper, crushed red powder, chili flakes
7) * Optional : Add a pinch of turmeric powder for satisfaction purposes (health is wealth - duh ! )

Once the above spices are fried (make sure you don't burn them or overcook them). Follow these steps:

8) In a mixing bowl, put the mix and add 2 tablespoon of cream cheese (Kiri brand preferably)
9) 'wing' the creamy spice mix into the holes of the mushroom heads. Over-fill it - 'concave' it !
10) Place the baking tray in the oven for 20-25 mins on 400F.

Easy recipe and loads of goodness (and brownie points ;) )












Wednesday, May 13, 2020

Versa - Viva La Vida !


12 years, yes - I owned Nissan Versa for 12 years. That is way above average (8 years) that Canadians hold on to their car for.

It was my first ever car. I bought it on May 2008 - few months before my son was born that year. I had bought it from an auction for $12k. It was a 2007 make and had 28,000km mileage. By the time it left my garage forever it was clicked around 276,000 km!

I can vouch for the car as the most economical car and required bare-minimum maintenance. Just like the car, I grew old with it. Ironically, I appreciated in value and Versa depreciated (pun intended).

I have lots of found memories with (and in) the car. During my career, it drove me to all my workplace in different cities - Markham, Mississauga, Thornhill, Cambridge, Waterloo and lastly- Vaughan. I have driven to Chicago twice in that car (both times to witness nuptial unions).

It was a bittersweet feeling - on one hand I was getting my first luxury car and on the other I was letting go my first car. Had I not been idly browsing for cars and getting snitched by the car dealer I would not had let go of you Versa. Yes Versa, I want you to know that I will always remember you and letting you go was not easy. I traded you with mere $500 - I know no one can put dollar value to you - you are priceless.

Doing a transaction in the holy and blessful month of Ramadan, I hope my new car brings joy and happiness for me and my family IA - Ameen.

P.S - I did not test drive Versa(2008) nor its replacement(2020) ~ somethings never change