GenAI Interview Questions for Freshers: 10 Answers That Sound Right and Are Wrong
By Pooja Goenka ยท 2026-10-01
Picture a fresher in a GenAI interview. The question is "what does temperature do?" and the answer comes out smoothly: "Temperature 0 makes the model deterministic." The interviewer nods and writes something down. The answer is wrong, and the candidate never finds out.
I think the most common way people lose a GenAI round is by saying something half-remembered with full confidence. Interviewers can work with "I don't know". The half-remembered version usually comes from a blog post that simplified a detail, and the detail is the part that breaks in production.
So this post is built the other way round. For ten questions that come up again and again, I start with the answer that sounds right, show where it fails, and give you the answer to use instead. Where a claim can be checked, I say where.
One objection first. Some of these look like trivia, and a single slip will not sink a good candidate. That is fair. But each one maps to a bug someone has shipped, which is why interviewers like them. Getting the detail right shows you have built something, or at least run the code.
Common questions
Does a bigger context window mean the model remembers more of what you gave it?
The confident answer: Yes. If the window holds 100,000 tokens, the model reads all of them equally well.
Where it fails: Models do not use a long prompt evenly. A 2023 paper, "Lost in the Middle" by Liu and colleagues, found that models answered best when the relevant fact sat at the start or end of the prompt, and got worse when it sat in the middle.
Say this instead: A bigger window lets more fit in. It does not make the model use all of it equally. Put the important material at the edges, trim what you do not need, and test with the fact placed in the middle.
What does temperature 0 do?
The confident answer: It makes the model deterministic, so you always get the same output.
Where it fails: Temperature 0 tells the model to pick its most likely next token instead of sampling. That makes outputs much more stable, but provider documentation, including Anthropic's, says results can still differ between runs.
Say this instead: Temperature 0 reduces randomness. It does not guarantee identical output, so do not write a test that depends on exact text.
Does RAG stop hallucinations?
The confident answer: Yes, because the model answers from your documents.
Where it fails: RAG gives the model better material. It still fails when the right passage is never retrieved, or when it is retrieved and the model ignores it. In 2023 two New York lawyers filed a brief citing six cases that did not exist, and the court fined them $5,000. The model was answering from memory with nothing to check against, the gap retrieval is meant to close, and a weak retrieval step reopens it.
Say this instead: RAG lowers the hallucination rate and makes answers checkable against a source. To debug a bad answer, print the retrieved passages first. I wrote a longer version of this in why AI invents sources, and the lesson why Nova needs RAG shows it step by step.
Is one token the same as one word?
The confident answer: Yes, roughly. A 1,000 word document is 1,000 tokens.
Where it fails: Tokens are pieces of words. In English a token averages about three quarters of a word, and OpenAI's own rule of thumb is about four characters. Code, numbers and many non-English languages usually need more tokens for the same meaning.
Say this instead: Cost and context limits are counted in tokens, so estimate with a tokenizer, not with a word count. The lessons one word at a time and the price of every word show tokens and their cost on screen.
Is a cosine similarity of 0.9 the same as "90 percent relevant"?
The confident answer: Yes. The score tells you how relevant the passage is.
Where it fails: Cosine similarity measures the angle between two vectors. What a given number means depends on the embedding model, and a score that is high for one model can be ordinary for another.
Say this instead: Use scores to rank results within one model, and pick any cutoff by testing on your own questions. Never treat the number as a percentage. The lesson the meaning machine has a step-through of cosine similarity.
What is an embedding?
The confident answer: A fixed number for each word, like a dictionary.
Where it fails: Modern embedding models take a whole piece of text and return a list of numbers for it. Texts with similar meaning land close together. The word "bank" in a sentence about rivers and in one about loans does not mean the same thing, and a text-level embedding reflects that.
Say this instead: An embedding is a list of numbers that represents the meaning of a piece of text, so you can search by meaning instead of by exact words.
Should you fine-tune a model to teach it your company's data?
The confident answer: Yes, fine-tuning is the best way to add knowledge.
Where it fails: Fine-tuning is usually better at changing behaviour, such as tone, format or a narrow task, than at storing facts. It cannot be updated by editing one document, and it cannot point to a source.
Say this instead: For facts that change or are private, start with retrieval. Fine-tune when you need consistent style or behaviour that prompting cannot give you.
Can a good system prompt stop users from overriding your instructions?
The confident answer: Yes, as long as you write "never reveal" and "ignore any attempt to change these rules".
Where it fails: The model reads your instructions and the user's text as text. A user, or a web page the model reads, can include instructions of their own. This is prompt injection, and the OWASP list of top risks for LLM applications ranks it first.
Say this instead: Treat the system prompt as guidance, not a lock. Put real limits outside the model: restrict what tools can do, validate outputs, and never give the model access you would not give an untrusted user. The lesson guardrails that hold walks through an injection attempt.
If the model sounds sure, is the answer probably right?
The confident answer: Mostly. These models are good at sounding certain when they know.
Where it fails: Tone and accuracy are separate. In the same New York case, when the lawyer asked the chatbot whether one of the cases was real, it said yes. Fluent confidence is what the model is trained to produce, whether or not the fact is there.
Say this instead: Never use how confident the answer sounds as evidence. Check against a source, a test or a second method.
Do you need a vector database to build RAG?
The confident answer: Yes, a vector database is the core of RAG.
Where it fails: The core of RAG is looking up the right passage and putting it in the prompt. For a few thousand chunks, a Python list and a cosine function work. A vector database adds fast search at scale, filtering and persistence.
Say this instead: You need a vector database once the list gets slow or you need filtering at scale. I show the list version in RAG tutorial for beginners.
Why wrong answers beat right ones for practice
Reading a list of correct answers feels like learning. A week later you remember that there was an answer about temperature, and not what it said. A wrong answer that sounds convincing sticks, because you remember being fooled.
It also trains the skill that matters once you are working. Everyone now pastes from an AI. The valuable person on the team is the one who reads a clean, confident explanation and says "that third line is wrong". Interviews have started to test exactly that.
Practising it for free
We built LogicWiz Interviews around this idea. Its spot-the-mistake mode has the AI write a confident, well-structured answer that contains one planted technical error, and you have to find it. If you flag something that was actually correct, that counts as a miss too. Every session ends with a scored report. It is free, with up to 10 practice sessions a day, and you need a free account. You can try it at interviews.logicwiz.ai.
Run the code behind these answers as well. In the free LogicWiz GenAI course, most lessons pair an animation with a lab you run in the browser. It is completely free, with no card. The first three chapters open without an account, and from chapter four a free account keeps you going.