The criteria are the part of a teamwork evaluation that tends to get the least attention and that decides almost everything about the result. A careful weighting model applied to answers from a vague question still gives you a precise number with nothing behind it.
This article gives you a set of criteria you can use as they are, explains the thinking so you can adapt them to your own projects, and covers the two ways these questions most often go wrong.

The instinctive thing to ask is some version of "rate this team member's overall contribution from one to ten." It is quick to write and quick to answer, and it produces data you cannot defend.
The problem is that a student answering it has to compress several unrelated things into one number. Did their teammate do a lot of work, or good work, or was easy to work with, or turned up. Different students weigh those differently, which means a seven from one reviewer and a seven from another are not measuring the same thing. When a grade is later challenged, there is nothing to point at except the number itself.
A second problem is that undirected ratings drift towards the top of the scale. In the absence of anything specific to assess, students give their friends and acquaintances high marks, because marking someone down feels like an accusation. The result is a set of nines and tens with one obvious outlier, which tells you only what everyone already knew.
The fix in both cases is the same. Ask about specific, observable behaviour, one behaviour at a time.
The following five criteria cover most group projects. Each is phrased as something a teammate actually saw happen.
Completed their share of the work. This person took on a fair portion of the tasks the group agreed and delivered them. Scale points might read: delivered everything agreed; delivered most things, with some gaps; delivered some things, with significant gaps; delivered very little of what was agreed.
Met the group's deadlines. This person delivered their parts in time for the group to use them. Scale points: on time or early, consistently; occasionally late but with warning; often late, which held the group up; did not deliver in time to be usable.
Was reachable and responsive. This person replied to messages and attended the meetings the group scheduled. Scale points: consistently present and responsive; mostly present, with occasional absences; frequently unreachable; effectively absent.
Contributed to the group's thinking, not only their own section. This person gave useful input on decisions and on other people's parts of the work. Scale points: regularly improved other people's work and helped decisions; contributed to discussions when asked; mainly focused on their own section; contributed nothing beyond their own tasks.
Worked in a way that helped the group function. This person handled disagreement constructively and helped the group move forward when it was stuck. Scale points: actively helped the group work well; caused no difficulties; occasionally made collaboration harder; frequently made collaboration harder.
Five criteria is about the limit for a reliable rating exercise at undergraduate level. Every additional criterion reduces the care given to the others, and the fifth one above is the first to drop if you want a shorter set.
Notice that none of the scales above are bare numbers. This is deliberate and it is the single highest value change you can make.
A one to five scale with no descriptions means whatever each student decides it means, and students are not consistent with each other or with themselves across a form. A scale with described points forces a comparison against a stated standard, which is the thing that makes ratings from thirty different reviewers comparable.
Described scales also make the result explicable. "Three of your four teammates selected 'often late, which held the group up'" is a statement a student can engage with. "Your teammates gave you an average of 2.3" is not.
Asking for a written comment on every criterion produces long forms that students rush. Asking for nothing produces ratings with no evidence behind them.
The workable rule is to require a short comment whenever a rating sits at the bottom or the top of the scale. This targets the requirement at exactly the ratings that carry consequences. A student who is about to select the lowest option has to write one sentence about what happened, which prompts them to check whether they are remembering a real event or expressing a general irritation. A student selecting the highest option gives their teammate something genuinely useful to read.
It also means that if a grade adjustment is later challenged, you have specific statements from several independent reviewers rather than a set of numbers.
Have students rate themselves on the identical criteria before they rate anyone else. This costs two minutes and returns two things.
It makes students think about their own behaviour against named standards, which is a small piece of professional development in itself. And it produces the comparison between self rating and peer rating that tells you far more than either number alone. Large gaps in either direction are worth a short conversation, and outlier detection surfaces them without you reading every form. If you want to know more about how self evaluation and outlier detection are configured, our Group Member Evaluation FAQ covers it.
The five criteria above assume a project where students divide work and reassemble it. Some projects are different, and the criteria should follow.
For team based learning and other courses where groups work together in the room, the division of tasks criterion matters less and preparation matters more. Replace it with something like: came to sessions prepared, having done the reading or pre work the group relied on.
For projects with a long build phase, such as software or design, add a criterion about quality of work handed over, because a teammate who delivers on time but delivers something unusable is a specific and common problem that the deadline criterion misses.
For clinical, laboratory or studio settings, add a criterion about safety and protocol, because these are behaviours students should be learning to notice in each other.
For groups of three or fewer, cut the set to three criteria and lean harder on the comments. In a very small group the ratings are nearly attributable anyway, so the specific statements are what carry the weight.
The first is asking students to estimate percentages of work done. "What percentage of the total effort did each member contribute, totalling one hundred" seems precise and is not. Students have no reliable view of how long other people's tasks took, the numbers rarely sum correctly, and the format invites negotiation within the group before submission.
The second is including criteria about attributes rather than actions. Questions about whether someone was a good leader, was creative, or was committed to the project ask students to judge a person rather than report an event. These produce the least reliable data, generate the most complaints, and are the hardest to defend when challenged.
Run your criteria on a low stakes activity first, ideally a short group task early in the semester with no grade attached. Look at the spread of responses. If almost every student selected the top option on every criterion, your scale points are too generous or the descriptions are not specific enough. If students used the full range and the comments contain concrete events, the set is working.
Then keep it. The value of a good criteria set compounds, because students who meet the same standards in their second year project already know what is being looked for, and start behaving accordingly from the first meeting.
If you would like to see how this works inside the tool itself, the Group Member Evaluation solution page shows how criteria, described scales and self evaluation are set up. The Group Member Evaluation FAQ is a good next stop if you have the practical questions that usually come up before a first run, such as how the activity sits in your LMS, what students see, and what happens when someone misses a deadline.
It is also worth looking at the tools that sit naturally alongside this one, including Group Formation, Team Based Learning and Peer Review. You can see how they all fit together on the FeedbackFruits tool suite page, or read more about what a feedback and assessment solution actually is if you are building the case for your institution.
And if you want to keep reading on this topic, we have How to give individual grades for group work, Eliminating free riding in group work and Make group work work with effective group selection