Showing posts with label ChatGPT. Show all posts
Showing posts with label ChatGPT. Show all posts

Saturday, July 11, 2026

LLM Analysis of College Ratings

Money Magazine recently published research on the best colleges in America. Its methodology includes metrics that are weighted based on graduation rates, costs, financial aid, student debt, and alumni salaries. My daughter is looking at University of California schools (among other options), and I noticed that while UC Berkeley, UC Davis, UC Irvine, and UC San Diego achieved 5-star ratings, UCLA “only” received 4.5 stars.

Beyond what was stated in the methodology, I could not find detailed information about the raw scores that led to the star ratings, so I decided to consult various large language models (LLMs) to see if they could do some digging for me and propose plausible explanations for why UCLA wasn’t given a full 5 stars as I would have expected.

I asked the exact same question to Copilot, ChatGPT, and Claude: “Money Magazine released their 2026 ratings of the best colleges in the US (https://money.com/best-colleges/). Try your best to find out why UC Berkeley, UC Davis, UC Irvine, and UC San Diego achieved 5-star ratings while UCLA only received 4.5 stars. Make sure to examing their scoring methodology and determine what criteria may have caused UCLA to not achieve a full 5 stars.” Note that I copied and pasted the same typo (“examing” should have been “examine”).

Copilot didn’t spend much time researching this topic, and it started spitting out an answer 1-2 seconds after I pressed the “return” key. One of its main explanations for UCLA not achieving 5 stars was due to affordability which didn’t make sense based on the table above, so I followed up with a 2nd prompt: “UC Berkeley has a higher estimated full price and estimated priced with average aid than UCLA, yet UC Berkeley achieved 5 stars. Are you still confident in your assessment? If not sure, then state so.” Copilot then backtracked and revised its explanation, and its revision sounded more plausible to me. Here is Copilot’s full response.

A similar thing happened with ChatGPT which spent only a second or two longer to “think” than Copilot, and it responded almost immediately. It too explained that cost was a factor in it not achieving 5 stars, so I followed up with the exact same 2nd prompt: “UC Berkeley has a higher estimated full price and estimated priced with average aid than UCLA, yet UC Berkeley achieved 5 stars. Are you still confident in your assessment? If not sure, then state so.” Like Copilot, ChatGPT also backtracked and revised its explanation. Here is ChatGPT’s full response.

On the other hand, Claude thought long and hard about its response. It displayed multiple websites that it was consulting to formulate its response, including Money Magazine’s methodology page and the websites from each of the schools mentioned. I can no longer see the pages it was consulting, as they disappeared and were not displaying in the final response. Although I didn’t use a stopwatch, I’d estimate that Claude took 20-30 seconds to think before providing a response. I felt that Claude’s explanation was the most analytical and credible of the 3 LLMs, and unlike Copilot and ChatGPT, I didn’t feel that any of its responses were contradicted by the data. Here is Claude’s full response.

Of course, these are my personal opinions which are completely subjective. Also, my conclusion that Claude performed the best on this specific task are not necessarily generalizable to conversations about other subjects. The take home message is that if you often consult LLMs, I’d encourage you to always think critically about LLM replies and to explore multiple LLMs to compare and contrast their responses and determine which ones are best suited to your needs.

Thursday, May 14, 2026

Image Editing with LLMs

Yes, you read that correctly. Large language models (LLMs) can come in handy for not just text and image generation—they can be versatile image editors too. Perhaps you’ve uploaded a photo of yourself to your favorite LLM and asked it to generate a caricature of you or portrayed you as a celebrity being chased by paparazzi. These are LLM-based examples of image editing in which you start with an image, add your text-based prompt to manipulate the image, and the result will hopefully resemble what you had in mind.

I recently came up with another use case for image editing. I was trying to find a high quality image of the Eagles’ Hotel California album cover. I searched the web and found many photos and scans of the album cover, but they all had one major shortcoming—the dark areas in the bottom half of the image had little to no detail. Here are 2 such examples:

I uploaded the images to both Copilot and ChatGPT and entered the following prompt: “These are 2 photos of the cover of the Eagles "Hotel California" album with different exposures and levels of detail in the shadows. Merge them into 1 photo and recover details from the shadows. Significantly boost the shadows so that it is possible to see the trees and bushes. Preserve the fluorescent "Hotel California" words that are superimposed on what looks like a car's side view mirror. Also boost the fluorescent "Hotel California" words so they are more bold. Preserve the original aspect ratio. Reduce overall contrast by making the sky a warmer golden glow and increase overall brightness, especially in the darker shadows.” I got very similar results with both Copilot and ChatGPT, and here is the result from the latter:

As you can see, ChatGPT did a remarkable job of creating the bushes in the lower half of the album cover. I honestly don’t know if it was able to recover detail from the source images or if it generated the bushes from scratch (or perhaps a combination of both). In any case, I was very happy with the recovery of what was otherwise lost detail. It also slightly sharpened the palm trees and did a very nice job of highlighting the “Hotel California” stylized wording while preserving its original look and feel. This was exactly the kind of image quality I had hoped to find in a Hotel California album cover, and although I was unsuccessful with my search, I was thrilled to learn that LLMs could generate the next best thing.

If you’re wondering why I wanted such an image, it’s because I wanted to 3D print the Hotel California album cover, and I wanted to include some detail in the bottom half rather than have it appear pure black. Here is the resultant 3D print:

If you have a 3D printer, you can access the 3D model and print profiles here. I hope this give you ideas for how you can use LLMs to edit photos.

Monday, March 16, 2026

Essay Grading - Human vs. Machine

For my daughter’s high school, I volunteered to read and score scholarship applications that were submitted by graduating seniors. There were 10 categories of applications including Academic Excellence, Arts, Athletics, Leadership, School Service, and others. All applications consisted of an essay, and some of the categories required the submission of supplemental information such as photos, videos, or other information to support the applicant’s scholarship candidacy. Parent volunteers were placed in groups of 3, with each group asked to review 4 or 5 applications. Parents were provided with a grading rubric and were asked to independently evaluate each student’s submission. To reduce the chance of bias, parents were asked to be reassigned to another group if they knew the student.

The grading rubric consisted of 5 dimensions for a total of 20 points:

Followed Directions
2 points – Followed most or all directions
1 point – Followed some directions
0 points – Followed no directions

Answered Essay Prompt
3 points – Answered the prompt completely
2 points – Mostly answered the prompt
1 point – Somewhat answered the prompt
0 points – Essay has nothing to do with the prompt

Well-Written and Use of Good Grammar
5 points – Essay is well-written and almost all of the grammar is correct
4 points – Essay is somewhat well-written and most of the grammar is correct
3 points – Essay is adequately written and the grammar is somewhat correct
2 points – Essay is sloppily written and has numerous grammatical errors
1 point – Essay is poorly written and has many grammatical errors
0 points – Essay is incomprehensible

Provided Examples of Supporting Evidence
5 points – Completely supported essay with examples of evidence
4 points – Mostly supported essay with examples of evidence
3 points – Somewhat supported essay with examples of evidence
2 points – Provided a few examples to support essay
1 point – Did not provide enough examples to support essay
0 points – Provided no examples to support essay

Impact of Essay
5 points – Essay was outstanding and made the reader feel invested in the student’s essay
4 points – Essay was good and the reader felt connected to the student’s essay
3 points – Essay was okay and the reader understood what the student was trying to express
2 points – Essay had a point and the reader didn’t lose interest while reading the essay
1 point – Essay was poor and the reader had to work to engage with the essay
0 points – Essay was disjointed and the reader was unable to connect with the essay

Up to 2 bonus points were also given for applications that required supplemental information, but I’ve omitted those criteria for brevity.

After submitting my scores, I wondered how my scores compared to those of other parents. Because I was the first volunteer in my group to complete my assignment I did not have visibility into how the other 2 parents scored the students’ applications. However, I was able to externally validate my scores against those of various large language models (LLMs).

METHODS

There are too many LLMs to count nowadays, so I consulted the 7 that I was most familiar with, and I’ve listed the most probable models that each one is likely to have used as of the time of this writing. Some LLMs are more transparent with the identification and versioning of their free and paid models. For all 7 models, I used the free tier.

  • ChatGPT: Default model: GPT-5.2 Instant; Fallback model: GPT-5.2 Mini or similar lightweight version if you exceed limits
  • Claude: Sonnet 4.6
  • Copilot: Copilot model, built by Microsoft
  • DeepSeek: DeepSeek-V3.2
  • Gemini: Gemini 3
  • Grok: Grok 4.20 beta, Auto (Fast or Expert)
  • Perplexity: model not shown or configurable on free plan

I used the exact same prompt for all 7 LLMs and all 4 students:

You are a parent of a high school student who has volunteered to evaluate scholarship applications. Students who apply for a scholarship under the category of SCHOOL SERVICE are given the following essay prompt: “What contributions have you made to our high school as someone who serves this community?” Students who apply for a scholarship under the category of LEADERSHIP are given the following essay prompt: “Would others consider you a leader and why?” OR “What is your definition of a leader and how do you embody those characteristics?”

The grading rubric is provided in the attached “Essay Scoring Guidelines.pdf” file. Provide scores as whole numbers for the following dimensions in accordance with the scoring guidelines:

1. Followed Directions (0-2 points)
2. Answered Essay Prompt (0-3 points)
3. Well-Written and Use of Good Grammar (0-5 points)
4. Provided Examples of Supporting Evidence (0-5 points)
5. Impact of Essay (0-5 points)

Ignore the “Bonus Points” dimension in the scoring guidelines because the scoring of that dimension may involve evaluation of photos or videos. The student’s essay is attached. Provide the score for each of the 5 dimensions along with a brief justification for each score.

For each model, I pasted the prompt and attached the essay scoring guidelines in a PDF file along with a PDF file the essay for student 1. I continued using the same chat thread, so I only attached the PDF files of the essays for students 2-4, as re-attaching the scoring guidelines repeatedly for each student would have been redundant. For privacy reasons, I have de-identified the student names and am not sharing the actual student essays.

RESULTS

My ratings, along with those of the 7 LLMs, are as follows (click the image to enlarge):

Although the LLMs did provide brief justifications for their scores, I’ve included only the numeric results but could easily furnish the complete LLMs responses upon request.

Overall, there was general agreement between my ratings and the average ratings from the 7 LLMs.   In terms of rank order, I gave the highest scores to Student 1 (19 points), followed by Student 4 (17), Student 2 (15), and Student 3 (12). Using the average of all 7 LLMs, the highest score went to Student 1 (19.7), followed by a 2-way tie between Students 2 and 4 (19.1), and then Student 3 (15.4). In other words, the LLMs agreed with my ratings for the best and worst applications, although they did not draw a distinction between the two applications in the middle of the pack.

Across the board, I was equally or more critical of the essays than the LLMs, as the LLMs generally gave the same or higher scores in each of the 5 dimensions of the grading rubric. Upon examining the total number of points allocated across LLMs, the 3 most “lenient” graders were Grok (78 total points awarded), Perplexity (77), and Copilot (76), while the “strictest” graders were Claude (69), ChatGPT (70), and Gemini (70).

DISCUSSION

All 7 LLMs were up to the task of grading the essays in accordance with the grading rubric. I considered the possibility that some LLMs might not completely follow directions, but all of them adhered precisely to the grading criteria and listed scores that were concordant with the criteria. Some LLMs even tallied up the total scores for each student even though I did not specifically request it in my prompt, and when they did so, they performed addition without any errors.

There are several possible explanations for the differences between my ratings and the LLM ratings. First, it is possible that I’m a tough grader. I went into this activity thinking that these were all brilliant students, and it would not be helpful if all the students clustered around near-perfect scores. In fact, this is exactly the outcome that was observed with the LLMs, as students 2 and 4 were deadlocked in a tie. Second, it is possible that the LLMs were lenient graders. After all, sycophancy in LLMs has been well-documented and researched, and many companies have made concerted efforts to tone down the level of sycophancy as they introduced new versions of their models.

This experiment validates that LLMs can be used to assess the quality of written text when evaluated against a custom rubric. This is probably not surprising to many readers who have already engaged with LLMs in similar ways, including myself. However, this is the first time I’ve quantified my findings. Another key takeaway is that LLMs can be used to critically appraise a body of written text so the author has a chance to make revisions based on the feedback. In academic settings, the mere usage of LLMs is not tantamount to cheating. It’s the way in which an LLM is used that constitutes whether the LLM serves as a learning aid or if it is used to cheat. In work settings, I encourage professionals to take full advantage of LLMs to enhance learning, spark creativity, and optimize productivity. As long as LLMs are used in a way that they do not substitute critical thinking, I think we have a lot to gain.

Friday, October 24, 2025

Ultra-Think Mode for LLMs

I use large language models (LLMs) almost every day to ask a variety of questions ranging from simple to complex. By default, LLMs generally provide concise answers. Did you know that you can alter the default prompting style? You can shape how LLMs respond through a set of instructions. For example, try using the following as your next prompt (copy and paste in its entirety):

From now on, use Ultra-Think Mode. That means:
• Analyze problems step-by-step before giving conclusions
• Explain reasoning clearly and in plain English
• Consider alternative explanations or perspectives
• End with a short summary or takeaway
If I say “Ultra-Think off”, go back to your normal concise style.

For both ChatGPT and Google Gemini, I’ve found that Ultra-Think Mode results in a structured response that follows a logical flow. Responses are slightly longer, but they remain topically relevant. I have not tried Ultra-Think Mode for other LLMs, so results may vary.

When I asked ChatGPT to tell me about Ultra-Think Mode, it emphasized that it is not an official ChatGPT feature name or setting. Rather, it is a behavioral configuration achieved through prompting. With some additional inquiries, it said that the following are all appropriate terms to describe Ultra-Think Mode:

  • Prompting style: The most accurate everyday term. It’s a way of shaping how the model responds through instructions.
  • Cognitive scaffolding prompt: A more technical term sometimes used in AI alignment and education research — it means a prompt that guides structured reasoning (like step-by-step analysis).
  • Metacognitive framework: In human reasoning terms, it’s a way of prompting the model to think about how it thinks — examining assumptions, alternatives, and reasoning clarity.
  • Response protocol: A practical phrasing for users who treat it as a repeatable, standardized mode of communication.
  • Instructional schema: In prompt engineering, “schema” refers to a reusable instruction set that defines structure, tone, and reasoning process. Ultra-Think Mode fits that definition well.

Based on these explanations, I would expect Ultra-Think Mode to work similarly across many LLMs. Regardless of how you’d classify Ultra-Think Mode, let me know if you think it is useful. And if you have found other prompting styles that are useful for certain inquiries, leave a comment!

Saturday, September 27, 2025

Detecting AI-Authored Content

I came across an article called “What Counts As Cheating With Ai?”. Ironically, when I read the first sentence, I suspected that it was written using a large language model (LLM; a kind of artificial intelligence or “AI” application). To confirm my hypothesis, I consulted with my favorite LLM, ChatGPT. Check out my full conversation for details.

In summary, ChatGPT drew its own conclusions and also consulted various external sources, with the ultimate conclusion that there is an 85-95% probability that the article was AI-written. Based on ChatGPT’s assessment, the reasons it provided include:

  1. Odd word choices / nonnative phrasing
  2. Inconsistent tense / mismatch / weird connectors
  3. Repetitive structure, formulaic transitions
  4. Errors not typical of human edits
  5. Lack of smooth coherence in some parts
  6. Metadata / site context
  7. References / linking style

Those 7 reasons are based on the article that I referenced above, and there are many other criteria that can be used in general to detect AI-authored content. ChatGPT also also compared the text from the article against published criteria from major AI detectors, and it stated:

  • GPTZero: Looks at perplexity (predictability of text) and burstiness (variation between sentences). AI text tends to have low burstiness and oddly consistent perplexity. The uniform style and repeated sentence shapes here match GPTZero’s “likely AI” profile.
  • CrossPlag AI Detector: Notes that AI often creates unnatural collocations and semantic drift. Examples: “analyzable and confusing” or “students will ne’er ace open” are exactly that.
  • Sapling AI Detector: Flags AI when there’s “inflated use of rare words not fitting context”. Words like “erstwhile” and “conscionable” fit this.

In conclusion, if you need help detecting AI-authored content, consider asking AI for help. I found ChatGPT’s reasoning to make a lot of sense, and the AI detectors listed above also seem to have valid criteria for identifying AI-authored content.